Job Description
Lead the design, implementation, and long-term evolution of a scalable infrastructure platform powered by agentic tools. Own the end-to-end development of features, ensuring reliability, performance, and operational stability. Partner with engineering, product, and incident management teams to improve service health, observability, and operational readiness. Participate in on-call rotations and live-site investigations, driving long-term solutions to prevent recurring issues. Elevate engineering quality through robust testing strategies, clear documentation, and disciplined release practices. Mentor engineers through design reviews, code reviews, and technical coaching, while strengthening the team's overall engineering foundation. Bachelor's Degree in Computer Science or related technical field AND technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Rust, Python OR equivalent experience. Demonstrated professional experience in software engineering and/or site reliability engineering including designing, developing, and delivering software and systems solutions. Master's Degree in Computer Science or related technical field AND technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, Rust, Python. Experience with Linux engineering. Experience designing, building, or maintaining large infrastructure systems. Demonstrated technical leadership (e.g., driving design reviews, guiding architecture decisions, improving engineering practices). Experience owning and delivering software systems across the full software development lifecycle.