Job Description
Lead the design and development of Azure Compute Fleet capabilities that orchestrate Virtual Machine Scale Sets (VMSS) across multiple Azure regions, improving capacity resiliency and reducing regional dependencies for customer applications. Design distributed control-plane services, APIs, and orchestration workflows for capacity placement, scaling, failover, and lifecycle management of large-scale compute resources across regions. Drive adoption of AI and Agentic AI across engineering workflows by reimagining existing processes, automating operational tasks, improving diagnostics, accelerating root-cause analysis, and increasing engineering productivity through intelligent agents. Partner with teams across Azure Compute, Capacity, Networking, and Platform Services to define technical strategy, influence architecture decisions, and deliver end-to-end customer solutions. Serve as a Designated Responsible Individual (DRI), participating in on-call rotations, leading live-site mitigation efforts, and driving improvements in reliability, observability, performance, and operational excellence. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. 4+ years of experience of designing, building and shipping high quality system software. Experience building highly available and fault-tolerant distributed systems. Experience leveraging AI, machine learning, large language models, or Agentic AI systems to improve engineering productivity, service operations, diagnostics, or customer experiences. Experience building cloud services, platform infrastructure, control-plane systems, or highly available backend services operating at hyperscale.