Job Description
Lead the design, implementation, testing, and delivery of complex backend capabilities for new search and AI retrieval workloads in Azure, driving technical design and execution through production deployment, operation, and continuous improvement. Drive product and service quality through hands-on DevOps leadership, including monitoring, diagnostics, incident response, root-cause analysis, automation, live-site tooling, and proactive identification and elimination of recurring reliability risks. Engage directly with customers and partner teams to understand complex scenarios, translate feedback into technical requirements, and drive resolution of high-impact issues. Develop deep expertise in Azure provisioning and service management, and use that expertise to influence architecture, operational practices, and engineering decisions across the team. Provide technical leadership for complex, ambiguous projects by defining architecture, identifying risks and dependencies, coordinating execution across teams, and helping other engineers make sound design and implementation decisions. Strong problem-solving, troubleshooting and communication skills. Good understanding of systems fundamentals, including operating systems, networking, concurrency, storage, algorithms and data structures, cloud platforms, and distributed systems. Openness to feedback and effectiveness at collaborating with diverse groups of people. Demonstrated independence, bias for action, and tolerance for ambiguity. 5+ years of professional software engineering experience designing, developing, testing, and operating production software. Experience developing production software in C#, C++, Java, or another object-oriented programming language; experience with scripting, automation, or SQL is an advantage. Building and shipping production grade cloud services, including designing and implementing solutions for telemetry and monitoring. Bachelor's degree in computer science or engineering (or equivalent experience). Experience with Lucene, Elasticsearch, Open Search and the full ELK stack is a plus Experience building systems on top of with one of the large cloud platforms is a plus Experience coordinating resources across diverse teams to restore service and maintain SLA's Ability to conceptualize a distributed service, it's dependencies and the transactional flow when troubleshooting across network, application, caching, queuing, load-balancing, storage and distributed services layers. High enthusiasm, integrity, ingenuity, results-orientation, self-motivation, and resourcefulness in a fast-paced environment. Deep desire to work collaboratively, find win/win solutions and celebrate successes Always leading with deep passion and empathy for customers and co-workers.