Key Takeaways

  • Amazon pilots measured a median 4.5x productivity gain, sometimes over 10x — against the 10-20% Liguori says she felt from completion, chat, and vibe coding.
  • Its Bedrock Mantle team built a new inference data plane with six people in 76 days against an estimate of 30 people over 18 months — but two were distinguished engineers.
  • In a 50-team Amazon Stores pilot measured on deployment velocity, half the teams landed under 3x; 90% used Kiro, so the split tracked habits, not tooling.
  • The new bottleneck is decision-making: frontier teams "spend more time making decisions than they do writing code."

Clare Liguori, a senior principal engineer at AWS who works mostly on Kiro, opened her AI Engineer talk with two numbers that don’t fit together. Across pilots inside Amazon, the company has measured “a median of 4.5x productivity improvement and sometimes more than 10x.” Her own experience of everything that came before — inline completion, chat, vibe coding — was, “completely anecdotally,” worth “maybe 10 to 20% more productive.” What closes that gap isn’t a tool: in a 50-team pilot inside Amazon Stores, “it wasn’t about the tools, it was about the way that they worked.”

From AI-Assisted to AI-Native: Building a Frontier Development Team — Clare Liguori, AWS From AI-Assisted to AI-Native: Building a Frontier Development TeamClare Liguori, AWS · AI Engineer · Watch on YouTube

What Amazon means by a “frontier developer”

Liguori, who has worked on agentic AI for over three years, sketched the progression from inline completion to chat to vibe coding, and now to what Amazon internally calls frontier development. She defines a frontier developer by three behaviors, not by a tool license: hands-off coding, since “frontier developers write maybe 1 to 2% of the code that they produce. The rest is agents”; infrequent interaction, since they “aim to get their coding assistant to run for up to hours at a time without their intervention”; and minimal idle time, meaning “multiple agents in parallel churning through a backlog of tasks.” Everything Amazon changed exists to make those three behaviors survivable.

Three experiments, and the caveat on each

The first proof point was the Bedrock Mantle team, which needed a new inference data plane for Amazon’s model hosting service — estimated at 30 people over 18 months. Instead, “they took six people and they built it in 76 days with Kiro,” measuring on commits and reaching what Liguori called up to 20x improvement. She immediately undercut her own headline: the six were “some of the top engineers literally in the company including two distinguished engineers,” so the story “spread like wildfire across Amazon, but it was also very unachievable for a lot of teams.”

The second was a 10-day Prime Video sprint: six engineers in a room with Kiro, cutting a delivery estimate from 90 weeks to 24 on the strength of their sprint commits. But they had “no on-call duties, limited meetings, very few distractions,” and a senior engineer had spent three weeks pre-writing “very detailed, small, well-scoped tasks” for them. Liguori’s verdict: “this was again not necessarily real life.”

The third experiment is the one worth studying. Amazon Stores ran a structured pilot across 50 teams with a “totally normal distribution” of early-career, mid-career, and senior engineers, all on existing codebases rather than greenfield. Over the better part of a year they measured deployment velocity rather than commits: not how much code came out, but how fast it reached customers.

The result split down the middle. Half the teams saw less than 3x; the other half hit the median 4.5x, in some cases more than 10x. Since 90% of teams used Kiro alongside internal tools, the tool was effectively a constant. What differed was that “the teams that achieved step function improvements intentionally changed the way that they worked,” while the rest “simply kind of sprinkled Kiro… on top of their existing way of working.” That matches what other orgs find when they roll coding agents out to a real team rather than a hand-picked squad.

The five habits

Kiro's official browser interface example showing work across several repositories
Kiro's official browser-interface example provides product context for the workflow discussion. It does not depict Amazon's private pilot or independently substantiate its productivity figures. Unmodified image. Source: Amazon Web Services / Kiro.

Amazon interviewed teams across all three experiments and pulled out five habits. Liguori was deliberate about the word — not practices, not a checklist — because “it’s not about that one sprint. It’s about doing this day-to-day.”

1. Invest in agent context. What engineers normally transfer through Slack, onboarding, code reviews, and standups, frontier teams wrote down. The habit is a reflex after every failure: when the agent does something you wouldn’t have done, ask what’s missing from your skills or steering files. She added a pruning half most teams skip — Sonnet 3.7’s quirks forced “a lot of do nots” into steering files, and with Opus 4.5, she said, “we don’t have to do that as much” — so the recurring question is: “do I still need this in my steering files or is this just bloating context?”

2. Slow down to speed up. “In almost every team that was interviewed, they reported that their productivity actually went down as they intentionally adopted a new way of working.” The investment is ordinary engineering work: better error messages so the model can tell what failed, new tools and MCP servers, restructured codebases agents can navigate. Some teams even changed languages, because untyped code gives “no compiler errors,” so “the model kind of guesses” — Rust “has become very popular inside of Amazon.” She hedged that one: “you don’t have to do that.”

3. Feed agents, don’t babysit them. This was Liguori’s own aha moment. “If you are having a back-and-forth conversation with your agent all day long, of course you’re not going to see four to five x productivity improvements because you are in the loop the entire time.” You wait 30 seconds to a minute for code to review, so you can’t run agents in parallel. The alternative is giving the agent a way to self-validate, “so that agents can self-correct and only come back to you when it meets a certain quality bar” — it compiles, it passes tests, it has coverage.

4. Make intent explicit. Amazon practices a lot of behavior-driven development and has built it into its products; in Kiro, the model can draft the spec for you. The failure mode she contrasts is vibe coding: a high-level prompt, a mountain of code, then “that’s not really what I meant.” Her argument is economic — “it is less productive to iterate with the agent on code when the intent itself was incorrect,” and it’s “a lot easier to iterate with the model… about a document than it is about code changes that are spread across a code base.”

5. Shift testing left. Fast feedback is what lets an agent run for hours unattended. “The agent is going to make mistakes and that’s fine. But if you give it the right signals, it can self-correct.” Teams added linters and unit, integration, performance, and security tests — none of it new, as Liguori conceded: “we all know we should have been doing [this] all along… But now the ROI is, I think, finally high enough.” Her specific highlight: mocking services with deterministic responses that run entirely locally, so the agent never spins up cloud dependencies and gets more loops per hour.

What still hurts

Liguori refused the tidy ending: “I would be remiss if I would tell you if you adopt all of these habits, you will achieve nirvana.” Burnout is the first named risk — she flagged a term for it she said she didn’t coin. Engineers stay up late “trying to get that perfect prompt that’s going to make their agent run for hours overnight.” Cognitive load rises with parallelism: “You’re constantly shifting between terminal tabs.” And review is its own tax: “reviewing AI output is often harder for some than actually writing it,” especially for early-career engineers who haven’t spent years reviewing other people’s code.

She implicated herself in the first organizational failure: leaders see capable models and ask why the team isn’t going faster — “I’ve been guilty of this myself” — while teams need “those two months to invest in your code base,” a pressure worsened by “companies on X saying how they’re shipping 20 PRs a day.” The second is going too broad too fast: roll out to everyone at once and “you have a lot of teams who don’t know what they’re doing” before anyone has found the practices worth copying — a staging argument that rhymes with Anthropic’s account of building with its own tools.

What happens next: the bottleneck moves

The third organizational problem will outlive the current tooling cycle. Writing code manually used to be the constraint; now, inside Amazon, “the speed of decision-making becomes a new bottleneck.” When a product took 9 to 12 months to build, two months to decide and two months to approve a launch vanished into the wash. Now that “the code only takes one to two months to write,” those approvals are the long pole. Frontier teams, she observed, “spend more time making decisions than they do writing code,” so her prescription is speed on reversible calls — “the more that you can make fast decisions, especially ones that are easy to be reversed, the better.”

For Amazon, 2026 is the scaling year: from 50 teams to “the next 2,000 teams.” Her one big takeaway was the same thing that explained the 3x/4.5x split: “frontier engineering is about intentionally changing the way that you work. And that is difficult. That takes time.”

Quick poll

Which is harder on your team right now?

Liguori's observation: Amazon's frontier teams now spend more time making decisions than writing code.

FAQ

What productivity gain did Amazon actually measure? A median of 4.5x, sometimes more than 10x, across internal pilots. In the Amazon Stores pilot the metric was deployment velocity to production, not commit counts, and half of the 50 teams came in under 3x.

What is a “frontier developer” at Amazon? Liguori defines it by three behaviors: writing roughly 1-2% of the code themselves, running agents for hours without intervention, and running multiple agents in parallel against a backlog.

Does adopting this make a team faster right away? No. She says almost every interviewed team reported productivity going down first, while they invested in agent context, error messages, new tools and MCP servers, restructured codebases, and tests. Switching to a typed language, she added, is optional: “you don’t have to do that.”