Status not verified yet. Help us verify after applying!
Job Description
Design and implement core inference infrastructure for serving frontier AI models in production. Identify and drive improvements to end-to-end inference performance and efficiency of state-of-the-art LLMs and GenAI models from OpenAI, Anthropic and xAI hosted on AI Foundary. Design and implement efficient load scheduling and balancing strategies, by leveraging key insights and features of the model and workload. Scale the platform to support the growing inferencing demand and maintain high availability. Deliver critical capabilities required to serve the latest and greatest Gen AI models such as GPT5, Realtime audio, Sora, and enable fast time to market for them. Collaborate with our partners both internal and external. Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, or Java OR equivalent experience. Technical background and solid foundation in software engineering principles, distributed computing and architecture. Experience working on high scale, reliable online systems and experience in real-time online services with low latency and high throughput. Experience working with L7 network proxies and gateways. Knowledge in Network architecture and concepts (HTTP and TCP Protocols, Authentication and Sessions etc). Knowledge and experience in OSS, Docker, Kubernetes, C++, Golang, or equivalent programming languages. Cross-team collaboration skills and the desire to collaborate in a team of researchers and developers