Back

Software Engineer L5 - AI Observability & Agent Evaluation

NetflixLos Gatos, CA, USA
You will be redirected to the main website
Posted on: 27 Sep 2026

Job Description

Build the observability framework and platform capabilities that give ML and GenAI systems metrics, logs, and distributed traces across online inference, batch scoring, feature pipelines, and agent orchestration, so teams can instrument their systems consistently. Build the primitives that let teams monitor model performance (accuracy, calibration, error rates), data quality, drift, and degradation on their own systems, rather than monitoring individual models yourself. Build and extend evaluation frameworks for LLM and agentic systems that support response quality, grounding and hallucination, task success, tool-use and trajectory correctness, and LLM-as-a-judge and human-in-the-loop scoring, giving teams reusable building blocks to define and run their own evals. Build reusable libraries, SDKs, and templates that make observability and evaluation the default for new systems ("observability-by-default"), lowering the barrier for teams to instrument and evaluate their work. Provide the dashboarding, alerting, and SLO/SLI building blocks (plus sensible out-of-the-box templates) that teams use to track model performance, latency, cost, and reliability. Experience in software, AI/ML, or platform engineering, with hands-on time in production observability, monitoring, or ML/LLM evaluation Ability to work cross-functionally with ML, data, infra, and product teams, and to communicate clearly about system behavior, quality, and risk. Experience with vendor integration and VPC deployment. Hands-on experience with ML/LLM observability and evaluation tools (e.g., Arize, Braintrust, LangFuse, Weights & Biases, Galileo, Vertex AI Model Monitoring, SageMaker Model Monitor). Experience building or shipping LLM/GenAI applications and evaluating them: prompt/result logging, evaluation metrics, LLM-as-a-judge, and human-in-the-loop review. Experience evaluating agentic systems: tool use, multi-step reasoning, and trajectory/task-success measurement.

Job Overview

Salary:
Not disclosed
JOB TYPE:
Not specified
Experience:
Not mentioned
Job Location:
Los Gatos, CA, USA
Job Level:
Not specified
Education:
Graduation
Software Engineer L5 - AI Observability & Agent EvaluationNetflix