Job Description
Design, build, and operationalize scalable ML and deep learning models using containers and orchestration platforms (e.g., Kubernetes). Develop and refine LLM prompt and fine-tuning strategies, build evaluation pipelines, and continuously optimize model quality, latency, and cost. Architect, implement, and operate multi-tiered distributed services at hyperscale, with high availability, fault tolerance, and low latency. Build model serving and inference infrastructure, including caching, batching, GPU capacity management, and A/B experimentation at scale. Design scalable APIs, data pipelines, and feature/signal stores that ensure efficient, secure, and reliable data flow between ML systems and product surfaces. Drive live-site excellence: instrumentation, monitoring, capacity planning, and incident response for ML-backed services. Collaborate with applied scientists, data scientists, backend engineers, and product teams to translate requirements into production ML systems. Participate in code reviews and architectural discussions, and mentor engineers across both ML and systems disciplines. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Proven experience designing, developing, and operating multi-tiered distributed services at scale. Hands-on experience building, deploying, and operating machine learning systems in production (model training, evaluation, and/or LLM-based applications). Experience working through full product cycles, from initial design to final delivery. Experience with LLM application patterns: prompt engineering, RAG, fine-tuning, and model evaluation frameworks. Experience with ML infrastructure: distributed training, inference optimization, GPU capacity management, and orchestration platforms such as Kubernetes. Experience with large-scale data systems (streaming, caching such as Redis, feature stores) and experimentation platforms. Demonstrated ability to work across the ML/systems boundary and a strong desire to keep doing both.