Job Description
Design, build, test, deploy, and operate reliable backend services, APIs, data pipelines, and automation workflows. Create intelligent capabilities that turn large volumes of operational data into timely, actionable insights for engineering teams. Apply software engineering and AI techniques to improve anomaly detection, alert correlation, incident diagnosis, and safe mitigation. Investigate production issues, identify root causes, and turn lessons from live-service events into durable product and engineering improvements. Partner with service teams, product managers, and engineers to understand reliability challenges and translate them into scalable solutions. Improve the security, observability, performance, quality, and maintainability of the services you own. Contribute to technical designs, documentation, operational readiness, and shared engineering practices. Bachelor's degree in Computer Science, Engineering, or equivalent practical experience. Strong experience designing, building, and operating distributed systems, cloud services, or large-scale backend platforms. Strong technical depth in service architecture, reliability, debugging, performance, security, and operational excellence. Proven ability to lead ambiguous technical work from problem framing through design, execution, deployment, and production support. Strong cross-team collaboration and communication skills, with the ability to influence technical decisions across partner teams. Experience collaborating across teams and communicating technical concepts clearly to engineering and product stakeholders.