About the Company:
Our client is a remote-first global AI product company building a proactive smart assistant for everyday users. The product brings intelligence to conversations, everyday tasks, organization, and workflows while requiring minimal prompting.
The platform is designed to reliably execute long-running workflows, retain persistent context, and complete real-world tasks. It must support multi-step reasoning, interact with external tools, and remain dependable despite the non-deterministic nature of modern AI models.
The goal is to make everyday tasks easier, faster, and more intuitive through practical AI experiences used by people around the world.
About the Role:
As a Machine Learning Platform Engineer, you will build the infrastructure and systems that power the company’s AI capabilities.
You will design and operate the systems behind the AI stack—from model training and evaluation to deployment, inference, observability, and continuous improvement.
You will work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities into production with confidence.
Tech Stack:
Python PyTorch / JAX LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM Cloud infrastructure Distributed systems ML and data pipelines Workflow orchestration GPU infrastructure and performance tooling Vector databases and retrieval infrastructure