(Summary generated by AI based on the full job description)
The project focuses on deploying and operating vLLM-based LLM inference services. Key stack includes vLLM, Kubernetes/OpenShift, NVIDIA GPU, Prometheus, Grafana and observability tools (OpenTelemetry, ELK). Responsibilities cover designing and running inference services on bare-metal GPU, managing GPU partitioning, automating model lifecycle from registries (Hugging Face Enterprise, S3), configuring HPA, optimizing vLLM settings, building metrics and dashboards, implementing an API Gateway, ensuring network security, and participating in troubleshooting and incident response. Required experience includes 5+ years in DevOps/SRE/Platform Engineering, at least 2 years in MLOps/AI infra, familiarity with CUDA and NVIDIA Container Toolkit, and Python and Bash skills.




By clicking "Aplikuj" you confirm that you've read and accepted our Terms and Conditions.
This is how the employer processes your data
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Need more information?