Senior Site Reliability Engineer
Offer summary

(Summary generated by AI based on the full job description)

The project focuses on managing AI hardware infrastructure to ensure high availability and scalability. Required skills include Python, Prometheus, Grafana, OpenTelemetry, Loki, BGP, IPv4/IPv6. Responsibilities cover automation, ticketing tools integration, monitoring, 24/7 incident management and cooperation with vendors and field technicians. A fixed quarterly/annual bonus and extensive benefits like medical and sports care are offered.

newyou can start ASAP

Senior Site Reliability Engineer

Company: Akamai Technologies

from: 19 August 2026
to: 18 September 2026
salary not specifiedcontract of employment (full-time)
Salary details
basic salary
fixed bonus (e.g., quarterly, annual)
Offer parameters
level:senior
working mode:remote
Kraków, Prądnik Biały
Kraków, Prądnik BiałyOpolska 100View on map

Requirements

Expected technologies

Python
Prometheus
Grafana
OpenTelemetry
Loki

Operating system

Windows
macOS
Linux

Our requirements

  • Possess a solid Computer Science foundation, demonstrated through formal education or equivalent practical experience in large-scale SRE or Production Engineering roles.
  • Demonstrate proficient tooling and coding ability in languages like Python to build scalable operational tools, API integrations, and automation frameworks.
  • Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.
  • Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
  • Demonstrate expertise designing service rollouts, establishing operational readiness criteria, telemetry baselines, and defining effective alerting thresholds.
  • Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems.
  • Demonstrate a proven ability to fully own ambiguous technical challenges, coordinate cross-functional teams, and diligently pursue production-grade solutions.

Your responsibilities

  • Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
  • Integrating automated workflows across disparate ticketing platforms like JIRA, Siebel, and PagerDuty to enhance resolution times for hardware and network issues.
  • Leveraging advanced AI utilities and LLM-assisted development paradigms to enhance technical execution, script creation, and system evaluation effectively.
  • Enhancing advanced private cloud and compute technologies to consistently optimize availability, latency, and systemic health in high-density hardware environments.
  • Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.
  • Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.
  • Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities.

About the project

The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings.
This position involves ensuring reliability and operational readiness for advanced hardware and software systems across regional data centers. Collaborate with product teams during development to optimize scalability, performance, and system reliability. Define and monitor key performance indicators while addressing breaches. Analytical abilities, coding expertise, urgent issue resolution, and dedication to maintaining service availability are essential for success in this role.

This is how we organize our work

This is how we work

in houseyou focus on a single project at a time

This is how we work on a project

  • code quality measures
  • code review
  • Continuous Deployment
  • Continuous Integration
  • documentation
  • issue tracking tools
  • testing environments

Work in a way that works for you

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Working for you

At Akamai, we will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life:
  • Your health
  • Your finances
  • Your family
  • Your time at work
  • Your time pursuing other endeavors
  • Our benefit plan options are designed to meet your individual needs and budget, both today and in the future.

    Join our Routing Services Validation team!

    The Cloud Networking division seeks an experienced Senior II Test Engineer for the Routing Services Validation team. Responsibilities include defining test strategies, designing detailed test plans, and creating automation frameworks to validate next-generation routing platforms.
    Company

    Development opportunities we offer

    • assistance in preparation to public speeches
    • conferences in Poland
    • mentoring
    • soft skills training
    • support of IT events
    • technical knowledge exchange within the company

    Benefits

    • sharing the costs of sports activities
    • private medical care
    • life insurance
    • remote work opportunities
    • dental care
    • corporate gym
    • corporate library
    • no dress code
    • coffee / tea
    • parking space for employees
    • leisure zone
    • extra social benefits
    • holiday funds
    • extra leave

    Akamai Technologies

    Akamai powers and protects life online. Leading companies worldwide choose Akamai to build, deliver, and secure their digital experiences helping billions of people live, work, and play every day. With the world's most distributed compute platform from cloud to edge we make it easy for customers to develop and run applications, while we keep experiences closer to users and threats farther away.
    Senior Site Reliability Engineer
    I apply to:
    Akamai Technologies
    Kraków, Prądnik Biały
    Pracodawca zbiera zgłoszenia przez swój system.
    Przejdziesz na zewnętrzny formularz.

    By clicking "Aplikuj" you confirm that you've read and accepted our Terms and Conditions.



    This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

    Need more information?

    • Make sure the body of the offer doesn’t already include what you’re looking for.
    • Ask a question if you need more information you’re interested in.
    • We’ll forward your question to the employer and aim to provide a response within 3 business days.

    Share this offer