Key Takeaways

  • Werry defines a context engine as a layer that "delivers organizational context to both your human workers and now increasingly your agents."
  • In the live demo, one Claude Code planning task cost under a dollar and took about a minute with Unblocked attached, versus about two minutes and a higher cost without.
  • He placed most teams at "stage four to five" of an eight-stage AI maturity curve, with stage eight being "software factories."
  • He closed on a customer line: "50% fewer tokens, faster triage, better answers."

Peter Werry of Unblocked ran the same planning task through Claude Code twice on stage — once with his company’s context engine wired in, once without — and pulled up the usage screen for both. With it, the plan cost “sub-dollar to create” and took “about a minute.” Without it, the run took “about 2 minutes” and “costs more to generate all that context.” His point was not the stopwatch. It was that the gap compounds: an agent left to discover context itself “doesn’t discover the right things,” so later steps run on the wrong assumptions and you loop again.

How to Generate Mergeable Code with a Context Engine — Peter Werry, Unblocked How to Generate Mergeable Code with a Context EnginePeter Werry, Unblocked · AI Engineer · Watch on YouTube

Your agent is a new hire who forgets everything overnight

Werry opened by asking the room to travel back to “the before times” — before agents — when, as he put it, “you were the context layer.” Engineers trawled data sources, dug through discussions, and built tribal knowledge the slow way.

Agents inherit all of that and add one problem of their own. “Agents are like new employees,” he said. “They reset their knowledge every time you start a new task.” His framing: an expert software engineer onboarding for the first time, every single time. “Every time they have to rediscover your code base, how your organization builds tests and how they deploy software with each and every task.”

To locate where teams sit, he walked an AI maturity curve slide: autocomplete “back in the GBT35 days,” then Cursor, then organizational wikis, then MCP and skills handed to agents so they can assemble context themselves. “That’s kind of where people are today,” he said. “Most people, they’re at the sort of stage four to five level.” Stage eight is what he called software factories — full automation of agents that “just can’t operate without organizational context. They get lost.”

(If you are stuck at that stage-four plumbing decision, see our comparison of MCP versus CLI for AI agents.)

“Access to information doesn’t equal understanding”

That line was the hinge of the talk, and Werry unpacked it with two failure modes.

The first he borrowed from radiology: satisfaction of search. You scan an X-ray for a region that might indicate cancer, find one indicator, and stop — missing others that would change the diagnosis. Agents do the same, he argued. “They find something that they think is correct and then they stop.” Attaching a wiki does not fix this, because the wiki “still doesn’t tell the agent where the information is that it needs.” The second failure: agents “don’t distill understanding.” They find information, but without the up-front legwork “they don’t understand how all the pieces fit together” — how dependencies interact, how the architecture should scope the work they do next.

He also pre-empted the obvious objection, that you could pour the whole codebase and every architecture document into a long context window. There is more organizational context than fits “even one that’s a million tokens in size,” and even where it fits it backfires: “it causes the agent to get distracted.”

He closed the section on an iceberg. Above the waterline is the code the agent can see and operate on. Below it: “the actual intent, the team conventions, past decisions, things that you’ve discussed in Slack, for example, architecture rationale.” He credited Tariq from Claude Code, in that morning’s keynote, for the term unknown unknowns, and offered his own rephrasing — “finding the things that really matter.”

The demo: two runs, one usage screen

Werry started with the human side, because he does not think it has gone away. He asked his own system about an internal component called the source mark engine and got back an architecture explanation plus a diagram that, he pointed out, “doesn’t exist” — it was generated from the code plus proposed future architecture. Underneath it, sources. “You show your work,” he said. “This is a trust building thing more than anything.” That matters at merge time: “When you hit merge on a PR, you need to understand what it’s doing.” Accountability “stops with us.”

The same answer had flagged optimization opportunities in the source mark calculator, so he asked Claude Code to plan that optimization — first without Unblocked. It searched the code, worked out the algorithm, and reached a conclusion he generously said “does a pretty good job.” Run again with the context engine attached, the plan “really kind of nails the nuances,” picking up PRs where the team had discussed future improvements, plus Slack threads, Notion pages, and architecture docs. Those sources come back into Claude Code itself, so “Claude knows exactly where to jump to next if it needs to elaborate on that context.”

The usage screen was the receipt: sub-dollar and about a minute with context, about two minutes and more cost without. He discounted his own numbers as he showed them — “ignore the wall clock time because I’ve had this open for about an hour” — and was careful about what they prove. “The real value of a context engine is not like the upfront cost on these short tasks. It’s the compounding effect.” He credited a second keynote speaker, Tariq from Sonar, for the same point about loops compounding.

Where organizational context shows up in code review

Unblocked also runs a code review agent, which Werry used to argue that “organizational context” means derived intelligence, not just raw data. The system reads pull request data and generates best practices that align agents to the codebase, then surfaces them during review.

The detail that landed with the room: a review comment appeared, and Richie — one of their senior engineers — replied that it was something he would say. It was something he had said; the system had surfaced his previous comments. Unblocked uses seniority and expertise “as a signal to boost comments that are important.” Uber’s uReview multi-agent code review attacks the same precision problem from the filtering side.

A second Richie interaction carried the talk’s most concrete claim. He noticed the number of surfaced review issues had “dropped precipitously,” debugged it with Unblocked, got roughly to the root cause, and asked the system to fix it. Unblocked can run as a cloud agent — Werry flagged this as internal and experimental — and it opened a PR. The fix mattered less than the write-up: the PR explained why it existed, correlating a model switch (after which, he said, issues “dropped a ton” because the newer Claude model’s behavior “is quite a bit different”) with Richie’s own observation, and linking to the Slack thread where it happened. “It found the Slack conversation, correlated all of that past history back again.”

Two things you can run yourself

The document query engine, covered in his Monday workshop, runs over a GitHub repository, ingests historical pull requests, and “synthesizes a schema based on the documents that it can sample.” You then query it through agent chat.

The engineering social graph is what Unblocked uses internally to pin down expertise and team relationships. On screen it showed clusters of people connected by review relationships — “these lines show like we review each other’s code” — which can be turned into team labels or projected as coverage across the codebase. The payoff: “you can see kind of where the holes are, where you might be lacking expert coverage.” He also demoed a “context engine simulator” that runs a task with and without generated context so you can see the delta on your own work.

What happens next

Werry’s headline numbers come from one live demo on his own codebase, with his own caveat attached. The line he closed on — “50% fewer tokens, faster triage, better answers” — is a customer’s claim, not a benchmark, and he presented it as one.

The argument underneath is more durable. If most teams really are at stage four to five — wikis plus MCP plus skills — the next step is not another retrieval endpoint but a layer that decides what a task needs and hands over only that, with sources attached so the agent can go deeper itself. The open question for the rest of 2026 is whether that layer ships as a product you buy or as plumbing the coding agents absorb; the assistants are moving fast in that direction, as our roundup of the best AI coding assistants tracks. Werry’s bet is that an engine fed by Slack, PRs, docs, and a graph of who knows what is hard to replicate from the repo alone.

Quick poll

What breaks your agent's output most often?

Werry argued both fail: agents stop searching too early, and dumping everything in "causes the agent to get distracted."

FAQ

What is a context engine, exactly? In Werry’s definition, a layer that “delivers organizational context to both your human workers and now increasingly your agents” — indexing pull requests, Slack, and architecture docs, then serving task-specific slices with citations instead of dumping everything into the context window.

Doesn’t a wiki already do this? Werry says no. Attaching a wiki “still doesn’t tell the agent where the information is that it needs” — the agent searches, finds one plausible hit, and stops, the pattern he compared to radiology’s satisfaction of search.

What were the actual numbers? Same planning task in Claude Code: sub-dollar and about a minute with the context engine, about two minutes and a higher cost without. He stressed that the value is compounding across loops, not single-task cost.

Is any of it open source? Two pieces, per the talk: the document query engine and the engineering social graph. He also demoed a context engine simulator that runs a task with and without generated context so you can compare.