Job Description
Drive architecture, implementation, and evolution of cloud platform capabilities that improve deployment velocity, reliability, governance, and operational excellence for AI workloads. Partner with engineering, product management, and research teams to translate complex business and technical requirements into scalable platform solutions. Lead the development of highly available, secure, and observable services leveraging Site Reliability Engineering (SRE) principles and best practices. Improve developer productivity by creating self-service experiences, automation, and tooling that simplify onboarding, deployment, monitoring, and management of AI applications. Analyze production systems, identify performance and reliability opportunities, and drive continuous improvements through data-driven engineering decisions. Mentor engineers, contribute to technical strategy, and influence engineering excellence through design reviews, knowledge sharing, and cross-team collaboration. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience designing, building, and operating large-scale distributed systems or cloud services with high availability, reliability, and performance requirements Experience operating production services, including observability, incident management, capacity planning, and service health monitoring. Experience building developer platforms, control planes, orchestration systems, workflow engines, or platform services used by other engineering teams