Job Description
Architect and optimize AI systems across hardware, compilers, kernels, frameworks, runtimes, and distributed infrastructure to improve training and inference performance. Partner with hardware architecture teams on hardware-software co-design, influencing accelerator features, memory systems, interconnects, execution models, and future silicon roadmaps. Drive performance optimization across kernels, communication, memory movement, quantization, attention, mixture-of-experts (MoE), and other critical AI workloads. Architect distributed training and inference solutions spanning large accelerator clusters, including parallelism, communication, memory management, and scaling strategies. Drive model enablement and performance improvements across AI frameworks and inference technologies such as PyTorch, Triton, vLLM, SGLang, and related ecosystems. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience partnering with hardware architecture teams on hardware-software co-design, accelerator enablement, or performance optimization. 7+ years of experience developing and optimizing high-performance AI systems, kernels, or accelerator software using CUDA, ROCm, Triton, or similar programming models. Deep experience optimizing large-scale AI workloads, including attention, mixture-of-experts (MoE), quantization, FP8, KV-cache management, memory efficiency, or related techniques. Experience enabling and optimizing large language, reasoning, multimodal, or other foundation models on AI accelerators. Experience designing distributed training or inference systems using techniques such as tensor, pipeline, expert, or sequence parallelism. Deep knowledge of AI frameworks such as PyTorch and experience optimizing production-scale training or inference workloads. Demonstrated experience leading complex technical initiatives across multiple engineering organizations and influencing technical strategy beyond immediate team boundaries. Publications, patents, open-source contributions, or other recognized contributions in AI systems, distributed computing, machine learning infrastructure, or hardware acceleration.