(Summary generated by AI based on the full job description)
The project focuses on developing and maintaining reliability for a large-scale distributed platform at Allegro scale. Core technologies include Kotlin/Java/Python/Go, Kubernetes, Docker, Prometheus/Grafana/ELK, CI/CD and IaC; Chaos Engineering and performance testing tools (e.g., Gatling) are also used. Responsibilities cover designing and implementing reliability solutions, building an AI incident-response agent, acting as Incident Commander and running post-mortems, developing a performance testing and automated fault-injection platform, managing technical debt, automating infrastructure, mentoring engineers and shaping team architecture.




By clicking "Aplikuj" you confirm that you've read and accepted our Terms and Conditions.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Need more information?