(Summary generated by AI based on the full job description)
The project involves vLLM inference platform on OpenShift/Kubernetes with bare-metal GPU. Key technologies include NVIDIA GPU, CUDA, Prometheus, Grafana, Python, Bash, GitLab CI, Jenkins, ArgoCD. Main responsibilities cover deployment and maintenance of AI services, GPU and HPA management, model lifecycle automation, and building monitoring and security solutions. Expertise in MLOps and LLM, especially in production and inference optimization, is required.




By clicking "Aplikuj" you confirm that you've read and accepted our Terms and Conditions.
This is how the employer processes your data
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Need more information?