Job Description
Lead strategic and technical initiatives spanning Storage Fabric, Performance & Scale, Fleet & Capacity, and LLM API, particularly where solutions require changes across multiple layers of the platform. Take ambiguous, high-impact infrastructure problems from first principles through strategy, architecture, roadmap, execution, and measurable results. Develop a system-level view of how storage, compute, GPU capacity, models, inference infrastructure, routing, caching, reliability, and product policy interact, and identify opportunities across organizational boundaries. Shape how M365 and Copilot scale, including workload portability, infrastructure utilization, capacity allocation, and reducing resources required per customer outcome. Connect demand, supply, workload characteristics, model efficiency, infrastructure efficiency, and economics to drive capacity and investment decisions. Identify architectural simplifications and platform investments that remove structural bottlenecks rather than repeatedly solving individual performance, capacity, or reliability problems. Use telemetry, experimentation, modeling, and AI-driven analysis to quantify opportunities, test hypotheses, and measure outcomes. Anticipate changes in models, hardware, workloads, and AI architectures and translate them into platform requirements and investment priorities. Work across engineering leaders, architects, product teams, and infrastructure organizations to establish direction and drive complex initiatives through execution. Define success metrics and communicate complex technical and economic tradeoffs to senior leaders in a way that enables decisions. Mentor PMs and technical leaders and raise the organization's ability to reason about infrastructure as an interconnected system. Bachelor's Degree AND 10+ years experience in product/service/program management or software development These requirements include but are not limited to the following specialized security screenings: Bachelor's Degree AND 15+ years experience in product/service/program management or software development OR equivalent experience. 6+ years experience taking a product, feature, or experience to market (e.g., design, addressing product market fit, and launch, internal tool/framework). 8+ years experience improving product metrics for a product, feature, or experience in a market (e.g., growing customer base, expanding customer usage, avoiding customer churn). 8+ years experience disrupting a market for a product, feature, or experience (e.g., competitive disruption, taking the place of an established competing product). Experience leading ambiguous, high-stakes technical initiatives spanning multiple engineering organizations and technology layers. Technical depth in cloud infrastructure, distributed systems, storage, compute, AI infrastructure, or large-scale online services, with the ability to reason across the broader stack. Experience defining platform architecture, technical strategy, or multi-year investment direction for large-scale infrastructure. Experience with LLM inference, model serving, GPU infrastructure, routing, caching, capacity management, storage, fleet management, performance engineering, or resource optimization. Demonstrated ability to use telemetry, experiments, workload data, and economic models to guide technical and investment decisions. A record of improving performance, scalability, reliability, infrastructure efficiency, capacity utilization, or cost through product or technical leadership. AI fluency, including using AI agents, automation, and prototyping to accelerate analysis, technical exploration, and product work. Solid systems thinking and the ability to develop a technical point of view from incomplete information, pressure-test it, and drive an organization toward decisions. Exceptional cross-organizational leadership and executive communication skills, with the ability to influence senior technical and business leaders without direct authority.