Key Takeaways
- Uber's time-to-first-review went from 3 hours in 2024 to 9 hours in 2026 as PR volume and size grew, making review the bottleneck.
- uReview posts about 25,000 comments per week; roughly 10% get any developer feedback, only 4% of PRs draw negative feedback, and the overall addressal rate is about 67%.
- Against a naive implementation, Uber's observability and eval work cut costs 60% while raising quality and accuracy roughly 70%.
- Figma's Eyal Blum argues verification is the highest-value investment — push checks into deterministic tooling, reserve humans for "is this the right thing to build."
Uber’s engineers used to get their first code review within three hours. By 2026 that number had grown to nine. That metric — from Will Bond and Ameya Ketkar’s AI Engineer talk — is why Uber built uReview, an in-house multi-agent code review engine now posting roughly 25,000 comments a week. Pair it with Eyal Blum’s talk on Figma’s agent rollout and you get the clearest public picture yet of what breaks when a big org runs coding agents: not the generation, but the review around it.
The bottleneck moved to review
Uber has thousands of engineers across hundreds of teams in 12 sites, working in six language-specific monorepos. Over the past 24 months, Bond said, both the volume and the size of pull requests have grown — the visible symptom being that tripled time-to-first-review. It is the shift showing up across the conference circuit: generation got cheap, so every downstream step that assumed expensive generation is now the constraint. We covered a related version of it in browser agents and the new shape of code review.
Uber evaluated commercial options first. Four constraints pushed the team in-house: most vendors do not support Phabricator, which Uber still uses while migrating to GitHub; agents writing code need the same review a human author gets, so the review layer has to be callable from inside the agent loop; hundreds of teams mean review knowledge cannot be centrally managed, so uReview plugs into Uber’s existing ownership model; and security and compliance reviews must run on everything, because you cannot rely on teams remembering to invoke a review skill.
Many generators, ruthless post-processing
Three surfaces — GitHub, Phabricator, and the agent loop — feed a single uReview service that takes requests, ingests feedback, and routes work to multiple generators tuned for different performance and cost profiles, plus hooks into third-party review systems for benchmarking. The interesting part is what happens after generation. Multiple generators mean duplicate comments, and a lot of them; as Bond noted, anyone who has pointed an LLM at a diff has seen the firehose. So uReview rates, categorizes, filters, and deduplicates, and only the highest-confidence actionable comments reach an engineer. Generation is the easy half. Suppression is the product.
Observability turned a demo into a system
uReview started humble: a single prompt doing per-file logic checks, one agent doing a thorough review, a dispatcher choosing between them. Observability was equally thin — cost tracking, an NPS survey, a Google Form, Slack support. The result, Ketkar said, was a quality-to-cost ratio “all over the place” instead of clustering in the good quadrant. The fix came in stages: sentiment analysis on developer replies to review comments, which surfaced whole classes of fixable issues; then addressal rate, meaning does the developer actually change the code; then agent trajectory profiling — tool calls, reasoning steps — so the team could see why the agent did what it did and tune cost accordingly.
Two learnings stood out. The model does not know when it is wrong, so correction has to come from team-supplied guidance: style guides, patterns, anti-patterns baked into the agent. And guardrails matter as much as instructions, because review has a time budget and an agent burning turns on the wrong things produces a worse review.
Customization without a central bottleneck
The stack layers up. Single-file reviewers do general-purpose “find the logic bugs in this file” passes. Multi-file deep review carries each monorepo’s anti-patterns and style guides. Above that sit AI linters — few-shot setups where developers deterministically gather context and run rules against it to catch mechanical issues. At the top, teams define a custom agent wired to their own knowledge base, past PRs, and a review skill.
Shipping that to hundreds of teams took three unglamorous things: piggybacking on Uber’s ownership model, co-locating customizations next to the code so developers can update them fast, and deterministic routing deciding which team gets which review from which generator on which model. Then surfacing observability back to each team, so a rule author can see their rule underperforming and fix it.
Ketkar’s punchline deserves framing on a wall: writing the skill was easy — teams just asked Claude to read their past PR reviews and generate one. Running those skills at scale with consistent quality and low cost was the hard part.
The numbers: about 25,000 comments a week, roughly 10% drawing any feedback, only 4% of PRs drawing negative feedback. Addressal rate near 67%, with almost three quarters of high-severity issues addressed. Against a naive implementation, costs down 60% and quality and accuracy up around 70%.
Figma’s three acts, and who gets stuck in act two
Blum frames adoption as three acts. One: you pick up an agent, get something simple working, feel 10x. Two: you apply the same practices to a bigger problem, it fails badly, and the trust collapses. Three: you learn the real skill — guardrails, prompting, context. The organizational problem is that adoption is uneven and everyone ships the same product anyway.
He also flagged friction rarely discussed on stage. Reduced developer agency costs job satisfaction: engineers who took pride in writing code now wait on output and burn out. The best engineers — the ones holding undocumented institutional context in their heads — become bottlenecks and are therefore slowest to adopt, because they see every failure firsthand. And design docs, Slack messages, and emails now run three or four times longer and arrive two or three times as often, saying about as much as before.
Verification beats prompting
Blum’s central claim: investing in verification is the highest-value thing you can do to a codebase. Any time you shift a check left — from a human to an agent — that is a win; once Playwright and MCP arrived, agents could explore the app themselves instead of a human navigating for them. Better still, when the agent reliably does something well, encode it into a deterministic flow that repeats cheaply and reserves the LLM for reasoning. He also recommends TDD ordering: have the agent write the failing test first. Reverse it and the agent fits the test to the code rather than the code to the criteria.
The mental model is the testing pyramid. Push what you can into deterministic analysis — linters, the compiler, unit tests. Let agents handle the middle, reviewing against architectural standards you have encoded. Leave humans the top: does this work, and is it the right thing to build? Anthropic’s engineering leadership describes a similar division of labor, covered in how Anthropic builds with Claude.
Plans, not prompts
Blum’s answer to lost agency is to move the craft upstream. It is not uncommon, he said, to spend a week writing a detailed plan — making the decisions, iterating, circulating it for review — and only then handing it to an agent. A good plan starts with a why at the top, which prevents drift on long runs, then breaks into small parts each verifiable independently, with a validation gate per phase so stage five is not built on unvalidated assumptions from stage one. His size heuristic is delightfully human: if you would need a cup of coffee before reading the corresponding PR, the chunk is too big. One example produced roughly 20 PRs of 10 to 100 lines each — about six weeks of coding work compressed into one, a 5x speedup including review.
Two cultural moves round it out. Bring the skeptics in and hand them the roadmap: their objections map exactly where your tooling fails. And practice attention-aware communication — human attention is the scarce resource now, so mark what an agent wrote versus what you wrote. Blum’s team opens every PR description with a short hand-written summary before the AI-generated section, a habit he adopted after sending an unmarked AI analysis to a respected senior engineer who read it as sloppy and said so.
What happens next: expanding the outer loop, not killing it
As engineers author less code directly, Uber sees a short path to some percentage of code landing with automatic approval. That raises the accuracy bar inside the inner loop: a low-quality comment sends the agent fixing backwards in a loop Bond called cavitation. It also changes what a nit costs — an agent will cheerfully fix 100 nits that would infuriate a human reviewer.
The hard question is feedback. Uber’s quality gains came from human replies to review comments; take humans out and that signal disappears. Bond’s conclusion is that the industry is not killing the outer loop but expanding it — moving human responsibility up a layer, from implementation details and API compatibility to architecture, domain expertise, and product thinking. Whichever assistant your org standardizes on — see our roundup of the best AI coding assistants in 2026 — the review layer is where the organizational work lives.
Quick poll
Would you let AI-reviewed code merge without a human approval?
At Uber humans still approve, but Bond said the company sees a short path to a percentage of code landing with automatic approval.
FAQ
What is uReview? Uber’s in-house multi-agent automated code review engine, presented at AI Engineer by Will Bond and Ameya Ketkar. It serves GitHub, Phabricator, and Uber’s agent loop from one service, posting roughly 25,000 comments a week.
Why did Uber build its own instead of buying? Per Bond: most vendors do not support Phabricator, Uber wants the same review running inside the agent inner loop, hundreds of teams need decentralized rule ownership, and security and compliance reviews must run on everything.
Does automated review actually get acted on? At Uber, yes — the overall addressal rate is about 67%, and almost three quarters of high-severity issues get addressed. Only 4% of PRs draw negative feedback on uReview’s comments.
What is the biggest lesson for a team starting out? From Figma’s Eyal Blum: invest in verification first. Push what you can into deterministic checks, let agents review against encoded architectural standards, and save human attention for whether the thing is worth building at all.