T-Hub - AIOps Engineer - AI Infrastructure & Orchestration
Offer summary

(Summary generated by AI based on the full job description)

The project focuses on deploying and operating vLLM-based LLM inference services. Key stack includes vLLM, Kubernetes/OpenShift, NVIDIA GPU, Prometheus, Grafana and observability tools (OpenTelemetry, ELK). Responsibilities cover designing and running inference services on bare-metal GPU, managing GPU partitioning, automating model lifecycle from registries (Hugging Face Enterprise, S3), configuring HPA, optimizing vLLM settings, building metrics and dashboards, implementing an API Gateway, ensuring network security, and participating in troubleshooting and incident response. Required experience includes 5+ years in DevOps/SRE/Platform Engineering, at least 2 years in MLOps/AI infra, familiarity with CUDA and NVIDIA Container Toolkit, and Python and Bash skills.

new

T-Hub - AIOps Engineer - AI Infrastructure & Orchestration

Company: T-Mobile

from: 6 October 2026
to: 5 November 2026
salary not specifiedB2B contract (full-time)
Offer parameters
level:mid • senior
working mode:remote • hybrid
Warszawa, Mokotów
Warszawa, MokotówMarynarska 12View on map

Requirements

Operating system

Windows

Our requirements

  • 5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.
  • At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.
  • Strong experience with Kubernetes and OpenShift administration in production environments.
  • Proven experience deploying and operating vLLM-based inference platforms in production.
  • Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.
  • Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.
  • Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.
  • Strong Python programming skills with experience developing automation and operational tooling.
  • Experience with Bash scripting and Linux systems administration.
  • Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.
  • Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.

Your responsibilities

  • Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure.
  • Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants.
  • Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3.
  • Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency.
  • Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing.
  • Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates.
  • Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements.
  • Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics.
  • Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns.
  • Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services.
  • Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices.
  • Maintain audit logging capabilities to support compliance, security investigations, and operational governance.
  • Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services.
  • Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives.
Company

What we offer

Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.

Recruitment stages

  • 1.
    Let’s meet to better understand each other's expectations.
  • 2.
    Hiring Manager will receive your application after meeting.
  • 3.
    Technical meetings with the Hiring Manager and/or team.
  • 4.
    Time to decide!

T-Mobile

Working at T-Mobile’s T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, T-Mobile not only provides a platform to hone your technical skills but also empowers you to be a catalyst for innovation. You'll have the opportunity to work at the forefront of cutting-edge technologies, from 5G to IoT and AI, shaping the future of connectivity. Our commitment to fostering a diverse and inclusive workplace means you'll collaborate with a wide range of talented professionals, learning and growing together. T-Mobile encourages a culture of continuous learning and development, offering mentorship and support to help you thrive in your career.

This is how we work

T-Hub - AIOps Engineer - AI Infrastructure & Orchestration
I apply to:
T-Mobile
Warszawa, Mokotów
Pracodawca zbiera zgłoszenia przez swój system.
Przejdziesz na zewnętrzny formularz.

By clicking "Aplikuj" you confirm that you've read and accepted our Terms and Conditions.



This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Need more information?

  • Make sure the body of the offer doesn’t already include what you’re looking for.
  • Ask a question if you need more information you’re interested in.
  • We’ll forward your question to the employer and aim to provide a response within 3 business days.

Share this offer