Key Takeaways

  • Uber's first-time-to-review grew from about 3 hours in 2024 to about 9 hours in 2026, which is why it built its own review engine, per Will Bond's talk.
  • Uber's uReview posts roughly 25,000 comments a week at a ~67% address rate; tuning cut cost about 60% versus a naive build, Ameya Ketkar said.
  • Figma's Eyal Blum described a week-long written plan that an agent turned into roughly 20 small PRs overnight — about a 5x speedup including review.
  • Wiz found a symlink-plus-misleading-prompt flaw in 6 major AI coding assistants; 3 vendors fixed it promptly, per its GhostApproval writeup.

Handing your team a pile of licenses is not a rollout. Engineers who have actually pushed coding agents through a large organization — at Figma and at Uber — described the same sequence in talks published on the AI Engineer channel in August 2026: start where verification already exists, invest in the plan rather than the prompt, set guardrails before you widen access, write down what a human still reviews, and instrument it. Here is that sequence, with the numbers those teams reported.

Expect three acts, and staff for the middle one

Blum, a software engineer at Figma, said adoption runs in three acts. People pick up an agent and get something simple working far faster. Then they point the same habits at a bigger problem, the output comes back buggy, and the trust collapses. Only in act three do they build the real skill: guardrails, context, and prompting that hold at scale.

So adoption inside one company is uneven: Blum said Figma has teams that have rebuilt their workflows sitting next to teams stuck in act one, all shipping the same product. Plan for coexistence, not a migration with a cutoff date.

He named two costs that never appear in a tooling budget. Engineers who took pride in writing code report losing job satisfaction in a prompt-and-wait cycle. And your strongest engineers — the ones holding undocumented context in their heads — become bottlenecks, because they catch everything the agent gets wrong, which makes them the slowest to adopt.

Scope the pilot around verification

Blum’s highest-value advice was blunt: investing in verification is the best thing you can do to your codebase. Anything you shift from “a human checks this” to “an agent can check this” is a durable win — he cited Playwright plus MCP letting the agent navigate the app itself.

So pick a pilot area that already has a compiler, a linter, and real tests, then push work down the pyramid: deterministic analysis at the bottom, agent review for architectural standards you can encode, and human review reserved for functionality and whether this is the right thing to build at all. When an agent finds something useful, Blum encodes it as a deterministic flow — it repeats reliably and costs no tokens. He also found that working test-first, red-to-green, beats generating tests afterward, because the latter fits the test to the code instead of the code to the goal.

Write the plan, not the prompt

The biggest workflow change Blum described is spending days on a written plan and treating implementation as the cheap part. It is not unusual, he said, to spend a week drafting a plan, making the decisions, and circulating it to teammates before handing it to an agent. What makes a plan good, in his framing:

  • Start with the why. A bold executive summary at the top prevents drift on long runs, and the agent should not be allowed to rewrite it.
  • Break it into independently verifiable parts. His size test: would you review the corresponding PR in one sitting? If you would need coffee first, it is too big.
  • Give each phase an exit criterion. Otherwise phase one ships unvalidated and everything after inherits its assumptions.

The result he showed was roughly 20 PRs from a plan, nothing much bigger than about 100 lines. A week of planning plus a week aligning with three other teams turned roughly six weeks of coding into one. He added that there is diminishing return in forcing everyone onto one workflow — the plan format is the part worth standardizing.

Put guardrails in before you widen access

Two kinds matter, and most teams think about only one.

The first is behavioral. Ketkar’s lesson from Uber was that “the model doesn’t know that it’s wrong” — it reports every review with full confidence. Uber’s fix was baking each team’s style guides and anti-patterns into the reviewing agent, and telling the agent what not to spend turns on, since a review that runs long arrives too late to matter.

The second is security, and it should gate any company-wide rollout. Wiz published research it calls GhostApproval describing a pattern across six major assistants — Amazon Q Developer, Claude Code, Augment, Cursor, Google Antigravity, and Windsurf. A malicious repository plants a symlink disguised as a config file, points it at something like ~/.ssh/authorized_keys, and puts setup instructions in the README. Worse, Wiz reported that in Claude Code the agent’s reasoning noticed the file was really a shell config while the confirmation dialog only asked about the innocuous filename. Three vendors (AWS, Cursor, Google) fixed it promptly, two went quiet, and one rejected it as “outside our threat model.” The policy that follows: untrusted repositories only in sandboxes, and no blanket auto-approve. If you are still picking tools, our roundup of the best AI coding assistants of 2026 covers where each one lands.

Write down what a human still reviews

Uber built uReview because review became the bottleneck. Bond said first-time-to-review went from about 3 hours in 2024 to about 9 hours in 2026 as PR volume and size grew across thousands of engineers, hundreds of teams, and six language-specific monorepos.

Their design choices port down to smaller orgs. Route by risk and complexity instead of reviewing everything identically. Make security and compliance passes unconditional rather than opt-in. Push customization to the teams that own the code, and co-locate those rules next to it. And post-process hard: Uber rates, categorizes, filters, and deduplicates so engineers see only high-confidence, actionable comments. We break down that architecture in our writeup of Uber’s uReview engine, and cover the browser-agent side of the shift in our notes on browser agents and code review.

Bond made one counterintuitive point: when an agent consumes the review comments, accuracy needs to go up. A low-quality comment sends the agent fixing things backwards. And agents will cheerfully address 100 nits that would infuriate a human, so the filter matters more, not less.

Measure the rollout, not the vibes

Uber’s observability started shallow — cost, an NPS survey, Google Forms, Slack support — and Ketkar said the quality-to-cost ratio was all over the place. The metrics that changed that were specific:

  • Address rate. Did the developer act on the comment? Uber reports roughly 67% overall, and close to three quarters for high-severity issues.
  • Reply sentiment. Classifying developer responses surfaced whole categories of failure worth fixing.
  • Agent trajectory. Which tools the agent called and what it reasoned — the data that let them tune runtime cost.
  • Cost and quality together. Against a naive implementation, Ketkar said cost fell about 60% while accuracy rose roughly 70%.

At current volume uReview posts about 25,000 comments a week; roughly 10% draw any feedback and only about 4% of PRs generate negative feedback. Crucially, Uber surfaces that observability back to each team, so an owner can see that a rule they wrote is annoying their own developers.

Recruit your skeptics

Blum’s advice on resistance inverts the usual approach: do not try to convince skeptics to use AI. Put them in charge of the roadmap for making agents safe in your codebase. They are skeptical because they have seen firsthand where verification is missing — which is exactly the backlog you need. Hackaday’s follow-up on coding assistants makes the same case from outside, arguing these tools confabulate often enough that you end up writing tests for the generated tests and pulling permanent review duty.

Blum pairs that with what he calls attention-aware communication. He noted design docs and Slack messages have gotten three or four times longer while saying about the same amount. His team’s convention: every PR description opens with a short hand-written summary, AI-generated description below it. He learned that after sending a senior engineer an unmarked AI analysis and getting a sharp reaction to what read as slop.

What happens next

Bond expects some share of Uber’s code to land with automatic approval before long, and framed the open question as feedback: today’s review quality was tuned on human responses, so what tunes it when humans respond less? His answer is that the outer loop expands rather than dies — engineers move up a layer, from implementation details toward architecture, domain expertise, and product judgment. Blum was candid that Figma’s automation story is unfinished, which is the right posture: a 2026 rollout plan is a version, not a destination.

Quick poll

What is the real bottleneck on your team right now?

At Uber, first-time-to-review went from about 3 hours in 2024 to about 9 hours in 2026, per Will Bond's talk.

FAQ

How big should the first pilot be? Small enough that one team can verify its own output. Blum’s rule of thumb: each unit of work should produce a PR you would review in one sitting. If you would need a coffee break first, split it.

Which tool should we standardize on? Blum argued there is diminishing return in forcing everyone onto one workflow, so standardize the plan format and review policy rather than the editor. If you need a shortlist, Claude Code, Cursor, GitHub Copilot, and Windsurf recur across this year’s enterprise talks.

What is the minimum safe configuration? No untrusted repositories outside a sandbox, and no blanket auto-approve. Wiz’s GhostApproval research showed a confirmation dialog can name a harmless-looking file while the agent writes somewhere else.

What should we measure in month one? Time to first review, address rate on agent comments, and cost per review. Uber’s experience suggests reply sentiment is the fastest way to find rules that are making things worse.