Back

Multimodal Infrastructure

YammerUnited States
You will be redirected to the main website
Posted on: 27 Sep 2026

Job Description

Design, develop and maintain large-scale multimodal data processing pipelines. Design, develop and maintain large-scale multimodal model pretraining and post-training frameworks. Design, develop and maintain large-scale multimodal model inference and serving frameworks. Work with research scientists and product engineers to solve infra-related problems. Find a path to get things done despite roadblocks to get your work into the hands of users quickly and iteratively. Enjoy working in a fast-paced, design-driven, product development cycle. Embody our Culture and Values. Bachelor's Degree in Computer Science, or related technical discipline AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Bachelor's Degree in Computer Science or related technical field AND 10+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience Strong proficiency in distributed data processing infra (resource utilization management, fault tolerance, ray & spark) and CPU/GPU batch processing optimizations Experience with state-of-art model inference and serving frameworks Experience with image/video/audio data processing Experience with common data formats for efficient I/O Strong proficiency in deep learning frameworks such as PyTorch, Megatron and Deepspeed Knowledge of auto-regressive and diffusion transformer models Experience with distributed training techniques such as data parallelism, model parallelism, and pipeline parallelism Proven experiences in at least one of the following areas: image/video generation and editing; efficient architectures (e.g., MoE, window attention); efficient model design; or reinforcement learning training methods (e.g., RLHF, DPO, GRPO) Strong proficiency in serving frameworks such as vLLM, TensorRT-LLM, SGLang, xDiT, Cache-DiT etc. Knowledge of distillation techniques such as Progressive Distillation, DMD, Self forcing etc. Knowledge of quantization and compression techniques like AWQ, GPTQ, and FP8 for multi-modal pipelines Experience in distributed inference scaling across multi-node clusters using Ray Serve and Triton Experience in leading technical projects and supporting architectural decisions with data
Multimodal InfrastructureYammer