Key Takeaways
- Garvin prompted a coding agent with a plain-language request to "replicate the Lovable pricing model" and it provisioned a Metronome instance with four scoped credit types, a draft invoice, and synthetic usage flowing against it.
- He drew an explicit boundary: the agent builds a test environment, not a production one — "we're not expecting to ship into production."
- Stripe CLI usage has "exponentially increased over the course of the past five six months," which Garvin framed as agents acting as buyers of infrastructure.
- Metronome's guardrails are two unglamorous things: portable skills files that carry API context, and deliberately verbose error messages "so that the agent can self-correct."
Andrew Garvin, co-founder of Metronome — the usage-billing platform Stripe acquired earlier this year in what he called “the largest deal that Stripe has ever done” — spent his AI Engineer talk typing one sentence into a CLI and letting an agent build a working credit-based billing engine from it. Then he spent the rest of the talk explaining why that agent is not allowed anywhere near production. “The goal that we have from a product development standpoint is not to have a customer operate the entire system without a human in the loop,” he said. The pitch is not autonomy. It is compressing the distance between an idea and a testable sandbox.
One sentence in, a billing engine out
The demo ran through Stripe Projects, which Garvin described as an orchestrator that “provisions a Stripe account for you as well as backend services that you may need” — he named Vercel and Postgres as examples, plus, in this case, a Metronome billing agent. All of it through the CLI. Stripe Projects launched, he noted, “literally the week that Metronome was acquired.”
His prompt was not a spec. He typed a request to create a demo billing engine in Metronome mimicking the Lovable pricing model, and that was the whole input. “The way that we coached the agent to be able to build this was just describing in natural language to replicate Lovable’s pricing model,” he said. “It was nothing more difficult than that.”
Lovable’s model, as Garvin characterized it on stage, is credit-only with a monthly auto-recharge, with multiple credit types scoped to different kinds of usage; if you overspend the balance, you get an invoice at the end of the period. That is more moving parts than it sounds. When he finally opened the Metronome dashboard, the agent had produced a customer with a lifetime spend figure, a first-class credit object for the test period, usage records drawing that balance down, and a draft invoice broken into build credits, plan mode credits, cloud credits, and AI gateway credits. “This is exactly what the Lovable pricing model looks like,” he said.
Worth noting: this was a live demo, and it stalled at the very first step with an initialization failure before a colleague came up to unstick it. Garvin flagged the risk himself before starting. It is a small thing, but it lands next to his argument about error messages rather than against it.
The line he refuses to cross
The most load-bearing minute of the talk is the one where Garvin explains what Metronome will not let the agent do. Billing, he argued, “is a type of system that is both business critical, has deep business logic behind it.” So the recommendation is narrow and specific: use your coding agent as a way to accelerate your work and get into a test mode and test environment. “We’re not expecting to ship into production. We’re not pushing it into production.”
What the agent produces is explicitly a sandbox. And Garvin made a point about what a useful sandbox has to contain that most teams underrate — it is not enough to see that a contract or a customer got provisioned. You need usage flowing through it. So the skills files direct the agent to push synthetic usage into the platform “so that you can see what a live customer would look like.” A billing config that has never metered anything has not been tested; it has been typed.
This is the same instinct that shows up whenever agents touch systems where being approximately right is worthless — the correctness-critical end of the spectrum where teams reach for machine-checked proofs rather than vibes, or where organizations put explicit review gates between agent output and the main branch. Money is in that category. A billing engine that is 95% right is a refund queue.
Skills files and error messages as the actual guardrails
Asked implicitly what stops the agent from wandering, Garvin’s answer was refreshingly unfancy. Metronome, he acknowledged, is “a very complicated and deep product and there’s a lot of different ways to hit foot guns.” The mitigation is “an extensible set of skills files that can provide context to the agent that’s implementing Metronome and working with our API.” They are portable and easy to install, so teams can use them on their own side rather than only inside Stripe Projects.
The second guardrail is developer experience work aimed squarely at a non-human reader. When an error surfaced mid-demo, Garvin used it: “this is nothing new, but our perspective is to have much more verbose and clear errors so that the agent can self-correct.” His DX team’s job, as he described it, is hunting for more failure cases like that, especially around initialization and setup.
That reframing matters. Terse error strings were always a human-ergonomics problem; now they are a retry-loop problem, and a vague 400 costs you a wasted agent turn instead of a confused engineer. It is also part of why the CLI keeps winning as an agent surface: readable failures make the next action obvious to both people and agents.
Agent as product, agent as buyer, agent as user
Garvin’s most portable idea is a three-way split for teams who say they are “building for agents” without decoding what that means.
Agent as your product. If the agent is the thing customers buy and it can run up a token bill on its own, you need to meter it. Metronome, he said, has taken in the API calls to OpenAI and Anthropic and metered them, working with both companies “since before they had any revenue.”
Agent as a buyer. This is what the Stripe Projects CLI demo actually was — an agent procuring a Stripe instance plus backend services. Garvin tied it to that exponential five-to-six-month rise in Stripe CLI usage, and to what he described as an exponential increase in new business formation at Stripe. The strategic implication for vendors: your services need to be discoverable to agents. He said providers including Vercel and Hugging Face are onboarding into the Stripe Projects environment for exactly that reason.
Agent as your user. This is the one with teeth. Garvin’s example was HubSpot, a Metronome customer for the past couple of years, which he said is on a path to transform its entire business from a seat-based model to a credits-based one — starting, per news he referenced rather than announced, in EMEA, where they dramatically lowered the seat price and layered credits on top. The logic: in a world where an agent can operate your entire system, charging per human seat stops describing the value delivered. He called this “headlessness,” noting Salesforce and others have used similar language, and said Metronome is seeing it today.
His anecdotal evidence was a demo day he attended the week before, where he said all five demoing companies were sales-led agents built to operate platforms like SAP or invoicing systems. If one agent absorbs the work of a department, all the value accrues to a single “user.”
What happens next
Garvin’s forward-looking claim is about pricing structure, not tooling. He argued the prepaid-credit auto-recharge model has been dominant since OpenAI launched theirs a couple of years ago through Metronome, and that the next move is enterprise sophistication on top of it. Coding-agent companies in the enterprise — he named Cognition, Cursor, OpenAI, and Anthropic themselves — are “starting to adopt more commit structures like the CSPs have done for the past 10 years”: prepaid commitments, postpaid commitments, and specific offers for specific customer types.
The other thread to watch is spend control. Garvin was direct that failure impact is growing “in particular because agents can run away with spend,” and said Metronome is thinking about controls its customers can pass on to theirs — his example was giving an agent a wallet it alone can spend from. That primitive does not exist as a shipped feature in this talk; it was framed as a direction. For teams currently deciding how much rope to give agents, it pairs with the organizational side of the question in how teams actually adopt coding agents at work.
Quick poll
Would you let a coding agent configure your billing system, even in a sandbox?
Garvin's own line: the agent builds the test environment, and "we're not expecting to ship into production."
FAQ
Did the agent actually deploy a live billing system? No. It built a sandbox in Metronome with a demo customer, credits, synthetic usage, and a draft invoice. Garvin repeatedly said the intent is to test and tweak in that environment before bringing anything into production.
What is Stripe Projects? Garvin described it as a CLI orchestrator that provisions a Stripe account plus backend services you might need — he cited Vercel and Postgres as examples, and a Metronome billing agent in this demo. It launched the same week Stripe acquired Metronome.
What are the “skills files” he mentioned? An extensible set of context files Metronome maintains for agents working against its API, intended to steer around foot guns during setup. He said they are portable and easy to install, so you can use them outside the Stripe Projects flow.
Why does he think seat-based pricing is under pressure? Because if an agent can operate an entire platform, one “user” absorbs the value that used to be spread across many seats. His concrete example was HubSpot moving toward a credits-based model, which he said is starting in EMEA.