T-Hub - AIOps Engineer - AI Infrastructure & Orchestration
Offer summary

(Summary generated by AI based on the full job description)

The project involves vLLM inference platform on OpenShift/Kubernetes with bare-metal GPU. Key technologies include NVIDIA GPU, CUDA, Prometheus, Grafana, Python, Bash, GitLab CI, Jenkins, ArgoCD. Main responsibilities cover deployment and maintenance of AI services, GPU and HPA management, model lifecycle automation, and building monitoring and security solutions. Expertise in MLOps and LLM, especially in production and inference optimization, is required.

T-Hub - AIOps Engineer - AI Infrastructure & Orchestration

Company: T-Mobile

from: 1 September 2026
to: 1 October 2026
salary not specifiedcontract of employment (full-time)
Offer parameters
level:mid
working mode:remote • full office
Warszawa, Mokotów
Warszawa, MokotówMarynarska 12View on map

Requirements

Expected technologies

Kubernetes
OpenShift
vLLM
Paged Attention
Continuous Batching
NVIDIA GPU
CUDA
NVIDIA Container Toolkit
Prometheus
Grafana
OpenTelemetry
ELK Stack
Python
Bash
Linux
GitLab CI
Jenkins
ArgoCD
Infrastructure as Code (IaC)
MLOps
AI Infrastructure
LLM
DevOps
SRE
Platform Engineering

Our requirements

  • 5+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations.
  • At least 2 years of hands-on experience supporting MLOps, AI Infrastructure, or Large Language Model (LLM) platforms.
  • Strong experience with Kubernetes and OpenShift administration in production environments.
  • Proven experience deploying and operating vLLM-based inference platforms in production.
  • Strong understanding of LLM serving concepts, including Paged Attention, continuous batching, and inference optimization techniques.
  • Deep knowledge of NVIDIA GPU technologies, CUDA drivers, NVIDIA Container Toolkit, and GPU troubleshooting.Hands-on experience with Prometheus, Grafana, OpenTelemetry, and ELK Stack.
  • Experience building observability solutions, including custom metrics, exporters, dashboards, and alerting mechanisms.
  • Strong Python programming skills with experience developing automation and operational tooling.
  • Experience with Bash scripting and Linux systems administration.
  • Familiarity with GitLab CI, Jenkins, ArgoCD, and Infrastructure-as-Code practices.
  • Strong analytical and problem-solving skills with the ability to work in complex, distributed environments.Excellent communication and collaboration skills.

Your responsibilities

  • Design, deploy, and maintain vLLM inference services on OpenShift/Kubernetes running on bare-metal GPU infrastructure.
  • Manage NVIDIA GPU resources, including GPU partitioning and allocation, to maximize utilization across multiple models and tenants.
  • Automate model lifecycle management, including model onboarding, versioning, deployment, hot-swapping, and rollback from private registries such as Hugging Face Enterprise and S3.
  • Implement and manage Horizontal Pod Autoscaling (HPA) based on workload demand, queue depth, and GPU resource utilization.Optimize vLLM configurations and serving parameters to maximize performance, throughput, and resource efficiency.
  • Build and maintain observability and monitoring solutions for AI inference services, including metrics collection, logging, and tracing.
  • Instrument vLLM endpoints to expose metrics related to token consumption, latency, throughput, and error rates.
  • Develop usage tracking mechanisms to monitor token consumption by API key, user, team, or department, supporting quota management and chargeback/showback requirements.
  • Create and maintain Grafana dashboards covering infrastructure health, GPU utilization, inference performance, service availability, and business consumption metrics.
  • Configure proactive monitoring and alerting using Prometheus and Alertmanager to detect infrastructure failures, performance degradation, and unusual consumption patterns.
  • Implement and maintain API Gateway solutions to provide authentication, authorization, rate limiting, and intelligent routing to inference services.
  • Ensure secure operation of AI services through network segmentation, ingress and egress controls, and adherence to security best practices.
  • Maintain audit logging capabilities to support compliance, security investigations, and operational governance.
  • Collaborate with AI Engineering, Platform Engineering, Security, and Infrastructure teams to deliver reliable, scalable, and secure AI services.
  • Participate in troubleshooting, incident response, root cause analysis, and continuous platform improvement initiatives.
Company

What we offer

  • Working at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.
  • You'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.

Benefits

  • sharing the costs of sports activities
  • private medical care
  • sharing the costs of professional training & courses
  • life insurance
  • remote work opportunities
  • flexible working time
  • corporate products and services at discounted prices
  • mobile phone available for private use
  • no dress code
  • parking space for employees
  • extra social benefits
  • sharing the costs of tickets to the movies, theater
  • holiday funds
  • birthday celebration
  • sharing the costs of a streaming platform subscription
  • employee referral program
  • charity initiatives
  • extra leave
  • platforma benefitowa

Recruitment stages

  • 1.
    Prześlij swoje CV
  • 2.
    Spotkaj się z przyszłym liderem/ liderką zespołu
  • 3.
    Witaj w T-Mobile :)

T-Mobile

We are a technology company, and our goal is to create innovative solutions for individual and business clients.
At T-Mobile, we all live in a magenta world! This color is close to our hearts and means faith in the success of undertaken actions, self-confidence, and endurance.
That’s who we are as a team.
At #MagentaTeam , we focus on exchanging experiences, agile work, and quick adaptation to changes! #MagentaTeam is, above all, a mix of different competencies, experiences, personalities, temperaments, and views. And this diversity is our greatest strength.

This is how we work

T-Hub - AIOps Engineer - AI Infrastructure & Orchestration
I apply to:
T-Mobile
Warszawa, Mokotów
Pracodawca zbiera zgłoszenia przez swój system.
Przejdziesz na zewnętrzny formularz.

By clicking "Aplikuj" you confirm that you've read and accepted our Terms and Conditions.



This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Need more information?

  • Make sure the body of the offer doesn’t already include what you’re looking for.
  • Ask a question if you need more information you’re interested in.
  • We’ll forward your question to the employer and aim to provide a response within 3 business days.

Share this offer