Status not verified yet. Help us verify after applying!
Job Description
You will.. Build and evolve our training infrastructure on Kubernetes with Infrastructure — GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and the operators and workflow engines that keep long-running training reliable. Shape the developer-facing surface — CLIs, SDKs, job submission, templates, paved paths — designed with the teams who'll use them. Make the common case one command and keep the uncommon case possible. Shorten the inner loop. Time to first training run, edit-to-signal latency, local iteration before a job hits the cluster, fast failure over slow mystery. Measure it, publish it, drive it down. Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping, evaluate honestly, and make the case with working prototypes and migration paths — or say plainly when a shiny thing isn't worth the switching cost. Strengthen the data and artifact layer. Dataset versioning, sharding, and high-throughput loading of large multimodal s