Job Description
Design and develop features for large-scale distributed systems and backend infrastructure services. Help improve service reliability, efficiency, scalability, and operational excellence through data-driven engineering. Develop telemetry, monitoring, diagnostics, and automation solutions to improve system health and operational visibility. Investigate production issues and participate in live-site mitigation and recovery efforts. Participate in design reviews, code reviews, testing, deployment, and operational support. Continuously learn emerging technologies, distributed systems concepts, and AI-assisted engineering practices. Contribute to a culture of innovation, accountability, and continuous improvement. Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 3+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 5+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Solid fundamentals in data structures, algorithms, distributed systems, and software design. Experience developing features or components as part of a service or distributed system. Ability to debug issues, write maintainable code, and deliver features end-to-end with appropriate guidance. Experience using AI-assisted engineering tools such as GitHub Copilot, ChatGPT, Cursor, Claude, or equivalent tools as part of day-to-day development workflows. Demonstrated ability to review, validate, and refine AI-generated code, tests, documentation, and troubleshooting guidance before production use. Experience contributing to backend services or distributed systems, including understanding of scalability, reliability, and performance tradeoffs. Familiarity with concurrency, resource management, service-to-service communication, and cloud-native architectures. Experience working with telemetry, logging, monitoring, and production diagnostics. Exposure to cloud platforms such as Azure, AWS, or similar. Microservices and service-oriented architecture Event-driven systems and APIs Resource management or infrastructure systems Containerized or cloud-native environments Exposure to AI-powered services, large-scale inference workloads, or rapidly changing traffic patterns. Ability to effectively collaborate with AI coding assistants and engineering agents to accelerate implementation, testing, debugging, and documentation while maintaining quality and ownership. Demonstrated curiosity, continuous learning mindset, and ability to rapidly adapt to new engineering tools and technologies. Solid communication skills and ability to work effectively within a team. Experience collaborating across teams and functions is a plus.