Back

Senior / Staff ML Ops Engineer

WaabiRemote US & Canada
You will be redirected to the main website
Posted on: 26 Sep 2026

Job Description

You will.. Build and evolve our training infrastructure on Kubernetes with Infrastructure — GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and the operators and workflow engines that keep long-running training reliable. Shape the developer-facing surface — CLIs, SDKs, job submission, templates, paved paths — designed with the teams who'll use them. Make the common case one command and keep the uncommon case possible. Shorten the inner loop. Time to first training run, edit-to-signal latency, local iteration before a job hits the cluster, fast failure over slow mystery. Measure it, publish it, drive it down. Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping, evaluate honestly, and make the case with working prototypes and migration paths — or say plainly when a shiny thing isn't worth the switching cost. Strengthen the data and artifact layer. Dataset versioning, sharding, and high-throughput loading of large multimodal s

Job Overview

Salary:
Not disclosed
JOB TYPE:
Not specified
Experience:
5+ years
Job Location:
Remote US & Canada
Job Level:
Not specified
Education:
Graduation
Senior / Staff ML Ops EngineerWaabi