Back

Principal Software Engineer

YammerRedmond, WA, USA
You will be redirected to the main website
Posted on: 27 Sep 2026

Job Description

Design, test and maintain large scale research platforms that drive our compute clusters, optimizing uptime and performance across diverse hardware. Develop end to end systems, libraries, and tools that accelerate the pace of research (including agentic development). Gather data and insights to develop the AI Infrastructure (compute and engineering systems) roadmap, identifying long-term investments and emerging technologies that will supercharge our systems. Provide vision, expertise, mentorship, and technical leadership to other team members. Collaborate closely cross-discipline with service engineers, product managers, and research and science teams to build better solutions together. Embody our culture and values Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience with distributed computing. Demonstrated experience designing/delivering platforms from concept to production. Experience building infrastructure for large-scale machine learning or generative AI workloads (including expertise in one or more of: Kubernetes, Docker, Volcano, SLURM, or similar). Experience building agentic workflows to automate AI developer inner and outer-loops. Experience in leading technical projects and supporting architectural decisions with data. Proven ability to profile, benchmark, and optimize large performance-critical systems. Track record of contributing to high-performance computing or large-scale AI infrastructure projects. Experience with GPU programming (CUDA, NCCL) and frameworks such as PyTorch. Deep expertise in networking (InfiniBand, NVLink), storage systems, or distributed training parallelisms.
Principal Software EngineerYammer