Key Takeaways
- DoorDash's GenAI platform team frames evals as one of four pillars, alongside an LLM gateway, an agent gateway, and open-weights model hosting — all aimed at balancing accuracy, latency, and cost.
- Four groups own different parts: strategy & ops set the quality bar, product turns it into rubrics, operations run annotation, engineering supplies APIs, telemetry, datasets, and judges.
- The loop presented on stage: trace → sample down → annotate → review → golden dataset → calibrate judges → monitor → repeat.
- Because the platform was API-first, non-engineers used Codex and Claude Code to vibe-code their own annotation UIs — which the team credits with cutting per-annotation spend on thousands of rows a week.
An eval is a repeatable way to score whether your AI feature is doing its job — and the single biggest mistake teams make is treating it as an engineering harness. That is the argument DoorDash’s GenAI platform team made at the AI Engineer conference: evals started at DoorDash as “another engineering thing,” and only worked once strategy, product, and operations people were writing the rubrics themselves. If you are standing up evals this quarter, the order of operations matters more than the tooling: capture traces, sample them down, get humans who know the domain to annotate, turn that into a golden dataset, then calibrate an LLM judge against it.
Why evals are not an engineering project
DoorDash’s GenAI platform team is horizontal: product teams build on primitives it provides. Its stated value is helping those teams balance three forces — accuracy, latency, and cost — a framing the presenters said applies to agents just as much as to raw models. Its pillars are an LLM gateway for swapping models, an agent gateway that centralizes tool connections, authentication, and agent identity so the security team can sign off once, hosting for open-weights models to control cost, and evals.
The eval pillar broke the pattern. When the team surveyed internal product teams, the needs did not rhyme. A consumer discovery and shopping assistant team wanted session-level quality judgments. Multi-agent systems needed trajectory-based evals — scoring the path an agent took, not only the final answer. Several teams simply needed a way to scale up human judgment.
What all of those have in common is that the person who knows whether an answer is good is not the person who can write the scoring code. At DoorDash that person was a strategy-and-operations analyst, a product manager, or a labeling partner. The presenters called evals a “team sport,” and the platform decisions followed from that: UI-first at the start (guidance the team attributed to co-founder Andy Fang), then API-first so engineers were not blocked waiting on a central team, and now workflow-first, where coding agents let ops and PMs drive the platform directly.
Who owns what
The clearest thing in the talk is the role split. Copy this table onto a whiteboard before you write a line of eval code:
- Strategy and operations set priorities and decide what the quality bar is.
- Product translates that bar into rubrics and workflows.
- Operations run the annotations.
- Engineering provides APIs, telemetry, datasets, and judges.
Notice that engineering owns none of the definitions of quality. That is the point. The presenters also said prompt ownership varies by team at DoorDash — in some teams strategy and ops own the judge prompt, in others the PM does, in others engineering — and the platform is deliberately built to allow all three, because org design is still being figured out.
The loop, stage by stage
The talk closed on a single slide the presenters described as an eight-step continuous loop. The stages named aloud were: capture traces and sessions, sample them down to a set you can actually look at, annotate with domain expertise, review, build golden datasets, calibrate judges and agents against those datasets, monitor over time, and repeat.
Two platform surfaces support it. The telemetry layer holds traces, scores, and observations, reachable over an SDK, the APIs, or an MCP server. The workflow layer is where non-engineers live: setting up annotation tasks, reviewing golden datasets, creating judges, and calibrating them.
Stage 1 — Trace, then sample hard
You cannot score what you did not record. Capture full sessions and traces from your agents and LLM calls, then immediately sample down. The presenters were blunt that the sampled set should be small enough that humans will genuinely look at it. This is the same discipline that shows up in agent-heavy engineering workflows — reviewing evidence instead of raw output only works if the evidence set is human-sized.
Stage 2 — Annotation, the stage everyone underestimates
Annotation is where domain knowledge enters the system, and it is also where platform teams drown. DoorDash found it effectively impossible to build a bespoke annotation UI for every use case — image annotation, manual testing, session review, each wanted something different, even though the underlying patterns were similar.
Their fix is the most transferable idea in the talk. Because the annotation, scores, and dataset APIs were stable and owned by the platform team, they handed the last mile to the operators: strategy-and-ops people used Codex or Claude Code to build their own annotation UIs against those APIs. The example shown on stage was unglamorous — annotating a restaurant menu — and that was the argument. It did the job, and no engineer was in the loop. If your org is already handing coding agents to non-engineers, this is a high-value first target; our guide on how to adopt coding agents at work covers the guardrails that make it safe.
The role split inside annotation is worth repeating: the platform team owns the APIs, a strategy-and-ops person decides what gets annotated, and an annotator does the annotating.
Stage 3 — Calibrate the judge, do not just write one
An LLM-as-a-judge prompt written once and never checked is a vibe, not a metric. The process described was: start with a draft judge prompt and a clear statement of what you want to measure in the output, run it across your traces to get baseline scores, then run an optimization loop using an off-the-shelf prompt-optimization library against your golden dataset. When the partner team is satisfied, that prompt gets promoted to be the team’s official LLM judge.
Two design choices made this stick. First, the team wrapped it in a self-serve UI so a PM or operator can set the configs and run the calibration loop without an engineer — and can pick the model, with Gemini shown in the demo and Claude or OpenAI models equally available. Second, they made the output reviewable: the UI shows the original prompt next to the calibrated prompt, so partners can see what changed and decide whether to trust it. Judge calibration is a black box by default, and trust was the blocker.
What DoorDash says it got
The presenters reported qualitative wins rather than headline numbers. Per-annotation spend dropped — they annotate thousands of rows every week, expensive at DoorDash scale — because self-serve annotation replaced work previously paid out to external annotators. Iteration got faster, teams calibrate their own judges without filing a ticket, and the platform reused existing DoorDash infrastructure. No percentages were given on stage, so treat these as directional.
What happens next
The trajectory the team described — UI-first, then API-first, now workflow-first — is the part to watch. Once coding agents are competent enough for an ops analyst to build their own tooling, the platform team’s job shifts from shipping interfaces to shipping stable APIs and letting everyone else generate the last mile. Expect eval tooling generally to move that way, and expect trajectory-level evaluation to get more attention as multi-agent systems become the default. If you are starting from zero, the honest sequencing is: instrument tracing first, sample small, get a real human rubric before you write a judge prompt, and only then automate.
Quick poll
Who should own your LLM-as-a-judge prompt?
At DoorDash all three happen: the presenters said prompt ownership varies by team, and the platform is built to allow each configuration.
FAQ
What is an AI eval, in one sentence? It is a repeatable scoring process — human annotations, an LLM judge, or both — that tells you whether your AI feature meets a quality bar you defined in advance, run against captured traces rather than ad-hoc prompts.
Do I need an LLM judge to start? No. The loop presented starts with tracing, sampling, and human annotation; the judge comes later and is calibrated against the golden dataset those humans produce. Building the judge first means you have nothing to check it against.
What is a golden dataset? A reviewed set of traces with trusted human labels attached. It is what you calibrate judges against and what you re-measure against over time, which is why the annotation stage is worth the effort.
How is this different from evaluating a coding assistant? Same loop, different unit of judgment. Agent systems need trajectory-based evals that score the steps taken, not just the final output — one of the distinct needs DoorDash’s platform had to support. For tool-level comparisons rather than in-house measurement, see our roundup of the best AI coding assistants.