Job Description
Build the platform behind Azure's fault self-healing and failure-prediction system with telemetry pipelines, prediction services, decision logic, and automated repair workflows. Ship AI agents to production: prompt and tool design, retrieval over diagnostics data, evaluation harnesses, guardrails, and the CI/CD path that deploys new agent skills safely. Close the loop safely: automated remediation with staged rollout, blast-radius limits, and verification. Own the developer experience: SDKs, APIs, and dashboards that let engineers across the org author and deploy new detection and repair logic themselves. Build and monitor measurements: prediction precision/recall, false-repair rate, action success rate, and regression gates that block a bad model from shipping. 8+ years building and operating production software, distributed services, data platforms, or large-scale automation. Python and C# (or C++/Rust), plus cloud-scale data pipelines at high volume. AI/ML in production, not just experimentation: agent or model serving, evaluation, versioning and rollback, drift and regression monitoring. LLM application patterns: agent/tool-calling, RAG, structured output and judgment about when an LLM is the wrong answer. A track record of automation that takes real actions on real infrastructure, with the safety engineering that requires. Bonus: anomaly detection on time-series data; Kusto/SQL; server hardware, firmware, or datacenter operations. Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 5+ years technical engineering experience OR equivalent experience These requirements include but are not limited to the following specialized security screenings: Preferred Qualifications: Experience supporting live-site operations for storage at fleet scale, including on-call ownership and production incident resolution