ML Ops Engineer
Remote
United States only
USD 120k to 160k/yr
Requirements:
- Architect, build, and operate end-to-end ML pipelines for training, validation and deployment on Google Cloud and AWS.
- Define, instrument, and maintain logging, monitoring, and alerting for model performance and data drift.
- Automate CI/CD for ML artifacts and infrastructure using GitHub Actions or equivalent.
- Collaborate with cross-functional teams, including frontend engineers, backend engineers, research engineers, and infrastructure engineers.
- Write clean, well-documented, fast, and maintainable code.
- Help ensure our systems have high availability and performance.
- Experience in computer graphics or physics-based simulation.
- Background in setting up Prometheus/Grafana, ELK, or similar monitoring stacks.
- Experience with Vertex AI.
- Experience working with custom Domain-Specific Languages.
What we're looking for
- BS in Computer Science or a related field.
- 5+ years of experience as a AI/ML Ops, DevOps, Infrastructure Engineer or equivalent.
- Expert-level Python and TypeScripts skills.
- Experience with Docker, Kubernetes, Terraform, Google Cloud and AWS.
- Deep understanding of machine learning models, including LLMs.
- Experience designing and maintaining CI/CD pipelines to fine-tune or train ML models.
- Excellent written and verbal communication skills.
Bonus Points
- Experience in computer graphics or physics-based simulation.
- Background in setting up Prometheus/Grafana, ELK, or similar monitoring stacks.
- Experience with Vertex AI.
- Experience working with custom Domain-Specific Languages.
Our tech stack
Source: the employer's careers page. Last checked 2026-09-30. Posted 2026-05-14.