Job Description
Lead the design, development, and operation of highly available distributed systems that power critical GitHub platform services. Drive architecture and technical direction for scalable backend services and platform capabilities using modern engineering practices, automation, and observability. Define and deliver improvements to service reliability, performance, latency, capacity, security, and operational efficiency for mission-critical workloads. Own complex features and services throughout their lifecycle, from requirements and design reviews through implementation, rollout, production support, and live-site excellence. Build automation and reusable platform capabilities that reduce operational overhead and improve engineering productivity across teams. Lead investigations of complex production issues across distributed systems, storage, networking, and service infrastructure, and drive durable corrective actions. Mentor engineers, raise engineering standards through design and code reviews, and leverage AI-assisted workflows to improve software delivery, testing, incident response, and operational excellence. Required Bachelor's or Master's degree in Computer Science or a related field, or equivalent practical experience. 7+ years of professional software engineering experience designing, building, and operating production systems. Expert coding skills in one or more programming languages such as C#, Java, Go, C++, Python, Ruby, or similar, with a track record of delivering maintainable production software. Deep understanding of data structures, algorithms, concurrency, API design, and distributed systems fundamentals. Proven experience designing and operating large-scale backend services, cloud systems, or platform infrastructure with high availability and reliability requirements. Strong experience with databases, caching systems, messaging technologies, service-oriented architectures, and observability practices. Demonstrated ability to lead ambiguous technical projects, influence architecture across teams, mentor engineers, and communicate effectively with technical and non-technical stakeholders. Experience with public cloud platforms, Kubernetes, and containerized services in production. Experience architecting high-throughput, highly available internet-scale systems and planning for capacity, resilience, and disaster recovery. Strong knowledge of observability, incident management, reliability engineering, security, and production operations. Experience with GitHub, CI/CD systems, developer platforms, package management ecosystems, or internal platform products. Familiarity with Go, Ruby, MySQL, Kubernetes, distributed storage, service infrastructure, and large-scale system migrations.