Key Takeaways
- Chrome stamps every input event trusted or untrusted: synthetic JavaScript clicks get dropped silently, while clicks issued over the Chrome DevTools Protocol are indistinguishable from a human's.
- Gallon cited an Arize AI eval where a CLI and an MCP server both hit roughly 83% task success — but the CLI took 7 turns and under a minute versus MCP's 71 round trips and 8 minutes.
- Jain says over 30% of code changes now merge with no review at all, and teams spend 4x the time they used to waiting on reviews.
- His fix: capture intent from agent sessions, codify recurring review comments into an "AI slop registry," and review a generated test plan plus its evidence instead of the diff.
Two talks released on the AI Engineer conference channel this month attack the same question from opposite ends: if agents write the code and drive the browser, what is left for a human to check? Corey Gallon of Rexmore showed an AI agent clearing Cloudflare Turnstile, MTCaptcha, a Lemin jigsaw puzzle and reCAPTCHA v2 with no human in the loop — after opening with the news that OpenAI had threatened to ban his account over the prep work. Ankit Jain, co-founder of Aviator, argued that line-by-line code review is already dead. Here is what each claimed, and where they hedged.
Talk 1: “A CDP browser is just like a meat bag with a mouse”
Gallon’s goal is unglamorous: “I want my agents to be able to use the web exactly the way that I do. I want them to book the things, to send the emails, to fill in the forms so that I don’t have to.” Same ambition as Google letting AI Mode complete hotel bookings — except Google negotiates partner integrations, and Gallon works against sites that agreed to nothing.
The talk rests on one slide: a CDP browser is just like a meat bag with a mouse. Drive Chrome through the Chrome DevTools Protocol — what the F12 panel speaks — and the agent’s clicks and keystrokes travel the exact same path inside Chrome that yours do. “At least as far as Google and Cloudflare and the rest can tell,” he added.
Give the agent a CLI, not an MCP server
Capability between a CLI and an MCP server is a wash, Gallon argued; everything else is not. He cited a study by Arize AI in which both, given the same task, succeeded roughly 83% of the time — but MCP took 71 round trips and 8 minutes for a task the CLI finished in seven turns and under one minute. On cost he cited Anthropic’s own reporting that a CLI can be as much as 75 times cheaper in token terms. The mechanism: a CLI sequence can be programmed and replayed with no model in the loop, while MCP hits the model every turn.
Worth flagging — none of those figures are Gallon’s own benchmark. They are second-hand slide citations, with no task set, model or harness described on stage.
The loop, and the ladder you climb only when forced
CDP is enormous — 57 domains as of the talk, hundreds of methods and events, which Gallon buckets into eight groups. Agents need only a small subset, framed as digital senses: see the page via the DOM, accessibility tree or a screenshot; hear it via network traffic and the console; operate it with clicks, keystrokes and navigation.
Around those runs a three-step loop — sense, act, verify — repeated until the page gives in. The discipline is in verify: sense through a different channel than the one you acted on. “If you’ve clicked something, don’t ask the click if it was successful. Check the network or check the screen.”
When the loop won’t close, you climb the meat bag ladder, only as high as the page forces you. Rung one is a synthetic JavaScript click, free and instant; his Outlook demo — open compose, fill programmatically, send, repeat for 20 or 200 emails — never leaves it. He drives that web UI rather than the Office 365 API because the API needs an app registration and admin approval an employee often can’t get, which makes the UI “a universal API, right, like a permissionless API.” Point the same click at Amazon’s add-to-cart button and nothing happens: no error, because hardened pages drop untrusted events silently. Rung two fires it through Chrome’s CDP Input domain, which stamps it trusted, and the item lands in the cart. Rung three is a real mouse path with dwell and jitter, plus vision.
Four gates, and a final boss on a clock
Cloudflare Turnstile hides its checkbox behind a closed shadow root inside a cross-origin iframe that contains another shadow root — no element to grab. So Gallon stopped grabbing: ask the browser where the iframe sits, compute the checkbox position, fire a trusted click at that spot on the glass. MTCaptcha needed vision — screenshot the challenge, read the characters out of the noise, type them back one trusted keystroke at a time. Lemin’s jigsaw slider samples the whole mouse trail, so the drag eases in, curves, deliberately overshoots and settles back, “just like a meat bag with a mouse.”
reCAPTCHA v2 — “the final boss of the internet” — got the architecture that is really the thesis. A solver, pure code with no model at all, does the trusted click, pierces the challenge iframe, screenshots each round and re-arms when a round expires. An operator, the agent, does the one step code can’t: look at the fuzzy tiles and decide which contain a bus. “Code does the deterministic driving and the agent does the only bits that require eyes and a brain.”
Speed is the reason. Rounds expire on a clock and one challenge can run several back to back, so an agent that round-trips a model on every click loses. Gallon says the solve is now repeatable and reliable, but gave no success rate on stage. His own caveats: on counsel’s advice every demo ran on infrastructure he owns and accounts he operates, and his tool Chrome Agent is a Python-installable CLI he offers with an explicit “I’m not here shilling a product.”
Talk 2: we already stopped reviewing code
Jain’s framing is that “when do we stop reading code line by line” is already answered. His problem slide: 861% code churn, a rising incident-to-PR ratio, a climbing median review time, and 4x the time spent waiting on reviews. Over 30% of changes, he said, now merge with no review at all. He named no source for those figures on stage and didn’t define the baseline window — treat them as his framing, not citable data.
The sharper observation is about AI review itself. If an AI coding assistant wrote the code and another one reviewed it, why is the exchange happening in a GitHub UI a human then skims before merging? “When AI reviews and nobody reads, we have configured the wrong thing.”
He also reframed what review was ever for. Formal code review is only 15 to 20 years old — Google launched Mondrian internally in 2006, and early Windows versions shipped without it. Catching bugs is the visible half; the other is alignment: knowledge sharing, mentorship, architectural feedback, onboarding. That is the piece his earlier five-layer trust model missed, as he said plainly. “For semantic accuracy, we can build better tooling, but alignment must survive.”
Capture the intent, then review the test plan
Spec-driven development doesn’t fix it, in Jain’s read: a spec written up front with no feedback loop is the 1970 waterfall model rebranded, and an LLM isn’t deterministic, so it makes its own calls during implementation. What survives is intent — it lives in the Jira ticket, the PRD, and above all in the prompts, “where the real decisions are being made.” Today we open a pull request and throw the prompts away.
His loop has three parts. Capture the back-and-forth decisions from the agent session as acceptance criteria. Maintain an AI slop registry — the review comments your team writes over and over, codified so they fire automatically; “every recurring comment is now a guardrail that you don’t have to review again.” Then combine criteria and invariants into a test plan, spin up a preview, and verify end-to-end. For a new payment form, an agent browses the app, fills it in, captures screenshots and pairs them with database snapshots to judge whether the criteria were met. That evidence becomes the review surface — reviewers argue intent and architecture, not diffs. His design rule: deterministic where it can be, LLM where you must.
Two caveats he raised himself. The registry is hand-built first — “it does follow a J curve, so pain is real.” And the plan must come from the session, not the code: if the agent that wrote the code also writes the test plan from it, the plan catches nothing. He is also selling — Aviator is piloting a product called Verify that implements this, and he asked for design partners.
What this means for your stack
Side by side, the talks describe one shift. Gallon pushed deterministic work into code and reserved the model for the single step needing eyes and a brain. Jain pushed verification off the diff and reserved humans for intent and architecture. Same line drawn twice.
The practical read: audit where a model sits in your loop and ask whether it needs to be there. Every model call is latency and cost paid on every run — which is why Gallon’s solver beats reCAPTCHA’s clock and a chattier agent doesn’t. Where a model is genuinely required, its speed becomes an architectural constraint, which turns picking between the frontier models into a latency decision as much as a quality one. That last inference is ours, not theirs.
What happens next
Gallon’s closer was methodology, not conquest: explore by hand, climb the ladder, then write the working path down as code or an agent skill. “You figure it out once and you do it forever.” That is also why the detection arms race speeds up from here — durable, reusable bypasses are a different problem for Cloudflare and Google than one-off scripts, and the legal ground is unsettled enough that Gallon fenced every demo behind accounts he owns.
Jain’s ask is testable this week: mine your last 1,000 review comments and build a slop registry from the repeatable ones. If most of that feedback is the same handful of notes, you have your answer about whether line-by-line review was ever the point.
Quick poll
What does a human on your team actually review before a merge?
For the record: Ankit Jain told the summit that over 30% of changes already merge with no review at all.
FAQ
Where are these talks from? Both were published on the official AI Engineer YouTube channel in August 2026 — Corey Gallon’s “The Dark Arts of Web Automation” and Ankit Jain’s “How to Kill the Code Review.” Neither video names its event on screen, so we describe them by the channel that published them rather than guessing an edition.
Why does a CLI beat an MCP server for browser agents? Per the Arize AI numbers Gallon cited, capability is roughly equal at about 83% task success — but the CLI took seven turns and under a minute where MCP took 71 round trips and 8 minutes, because a CLI sequence replays with no model in the loop.
Why do synthetic JavaScript clicks fail on some sites? Chrome stamps every input event trusted or untrusted. A JavaScript-dispatched click is untrusted, and hardened pages drop it silently — no error, nothing happens. A click issued through CDP’s Input domain is stamped trusted.
What is an “AI slop registry”? Jain’s term for the review comments your team writes repeatedly, codified into automatic guardrails so nobody writes them again. Build one by mining your last 1,000 review comments — he warns it follows a J curve, with the payoff after the setup pain.