Key Takeaways
- Long Lake has raised over $3 billion and acquired 35 services businesses, plus a recently announced $6.3 billion take-private of American Express Global Business Travel.
- His five-rung autonomy ladder runs co-pilot, synchronous agent, asynchronous agent, long-running agent, AI co-worker — and teams must "earn the right to do more" rather than start at the top.
- Traces from agents working beside employees become auto-scored evals with real ground truth, and every week those hill-climbing benchmarks become regression tests.
- Continual learning and enablement are usually two siloed teams; Shenoy argues they are one loop, closed by "extreme software service co-design" done in person.
Varun Shenoy opens his AI Engineer talk with an uncomfortable question: the models keep getting better, so why has nothing changed inside a 200-person property management firm? The Long Lake co-founder argues that stall is not a failure — it is what every general-purpose technology looks like on the way in. “Diffusion of any technology takes a generation,” he says, and AI diffusion is “perhaps the single most important problem for the next 20 years.” His company’s answer is unusual: Long Lake does not sell software to services businesses. It buys them.
The demo is real. The 200-person firm is unchanged.
Shenoy starts where everyone starts: an agent booking a flight, closing a support ticket, producing a block of code ready to commit. “Two years ago any of this would have been complete science fiction,” he says. “The capabilities are real.”
Then he walks into that property management firm — “real people, real properties, real dollars, real customers all across the US” — where nothing has changed. He argues this is normal. Electricity was invented in the 1880s and first demoed at Edison’s Pearl Street Station Dynamo Room in Manhattan, which Shenoy calls “the magic demo of its time.” Adoption still took decades: a Ford factory had to rip out its motors, bring in new equipment, retrain everybody. The photo he shows of Ford’s electrified moving assembly is dated 1924 — four decades after the demo. The models improve on their own schedule; the bottleneck is everything around them.
Why Long Lake buys the companies instead of selling to them
Long Lake has raised over $3 billion in two years from Elad Gil, General Catalyst, and AlphaWave. The strange part, in Shenoy’s words, is that it doesn’t sell software. It acquires and partners with real services businesses — 35 so far, across HOA and property management, architecture, HR services, and more — plus the recently announced $6.3 billion take-private of American Express Global Business Travel, “the world’s largest corporate travel platform.” More than half of Long Lake’s people sit on the technology team, building products and deploying them into the field.
The consequence is the line worth quoting: “We own these businesses. So, when the AI doesn’t work, it’s not their problem. We’re not the vendor. It’s our problem.” That inversion produces the three lessons that follow.
Lesson one: you have to earn the right to do more
Shenoy lays out a spectrum of autonomy with five rungs. The co-pilot is “your simple rag chatbot from two years ago.” The synchronous agent — he names Claude Code and Codex — runs one to five minutes, calling tools and using skills, but still waits on you to ask. The asynchronous agent goes into the background and comes back, and crucially need not be triggered by a user: a task completes, an async job queue picks it up, and the agent proactively offers advice. The long-running agent works for “hours, days, weeks, months,” a core problem he says a lot of the labs are focused on, as is Long Lake. Finally the AI co-worker: “This is what everyone wants to sell you.”
His counter-lesson from owning the outcomes is blunt: “you have to earn the right to do more.” For some tasks the models aren’t there yet. And you have to work with people in the field closely enough that they understand the rungs get climbed over time — the sequencing problem in our guide to adopting coding agents at work.
The jagged frontier: code is solved, services are not
Shenoy overlays the ladder with the jagged frontier. The code column is largely worked out: synchronously, a coding agent on your desktop with file system access; asynchronously, the same agent in a sandbox, left to build and test, returning work as a PR. One reason it works is cultural, not technical. Engineers already parallelize: “It’s very commonplace to launch 10 jobs and be comfortable with the fact that job seven might finish before job three.”
Services has a synchronous equivalent too — a co-working agent with deep context about the enterprise, wired into MCPs, custom tools, and custom integrations. The open square is the bottom right: what does an asynchronous agent look like “in industries where work is traditionally done in a very, very serial manner”? Shenoy leaves questions rather than answers. Since the models are trained on code, “what if we just use that code knowledge and represent knowledge work as code”? How do you parallelize inherently serial work — “people clean out their inbox one email by one email, not 10 emails at once”? And what form factor fits? The launch mechanism for code, he argues, “doesn’t mean that same way is going to work for architecture or property management.”
Lesson two: the most valuable tasks were never on the internet
“Frontier models have learned from everything humanity has written down,” Shenoy says, “but the most valuable tasks are not on the internet.” His examples are deliberately unglamorous: closing the books when you’re missing receipts, scoping a building for construction in a blueprint, coordinating vendors to fix a broken roof. That knowledge lives in people’s heads, in 20-year-old software, and in the way one senior person just knows how to do it.
Long Lake’s flywheel tries to make it explicit. Agents collaborate with employees on real work, producing rich traces — tool calls, hiccups, papercuts, everything that goes wrong. Those traces become evals with genuine ground truth: did the roof get repaired, did the books get closed. And it ratchets: “every week our hill climbing benchmarks become a regression test.”
He names three upshots: evals built and scored automatically, capturing explicit feedback (thumbs up and down, written notes) and implicit feedback — the diff between what the AI generated and what was ultimately submitted, “rich information that almost no one else has”; internal post-training on operating data “completely out of distribution for most frontier labs”; and customization per company, per user, and per client. Building that measurement layer without owning the business is the harder version, covered in our evals guide for teams.
The slide he says captures everything is two panels. The way we talk about LLM tasks is the top one: a slope, a bike, clear sight to success. Real work is the bottom — hills and ravines, death by a thousand paper cuts. “The exceptions are the job,” he says.
Lesson three: continual learning and enablement are one loop
Two trends dominate 2026, Shenoy says: continual learning, in the prompt or in the weights, and enablement. The first belongs to research or platform engineering, the second to growth or customer experience — “usually pretty siloed, not much interaction between the two.”
He argues they are the same snowball: more usage drives continual learning, drives a better agent, drives more usage. Which leaves the elephant. “Everyone assumes the usage just shows up. But as we all know, that’s simply not the case. It never does.” If the person who has closed the books for 20 years keeps doing it the same way, nothing happens.
Extreme software–service co-design
Shenoy’s fix borrows from Jensen Huang, who in his telling dominated the market through “extreme hardware software co-design.” Long Lake’s version is extreme software service co-design: co-designing products with the people and processes inside its businesses, only possible under the same roof.
That means meeting people metaphorically — building into the systems they already live in, so “the energy required for enablement is kept low.” He lists Excel, ERP systems, 3D design software, Outlook, and Gmail.
And it means meeting them physically. Get on a plane. Run a lunch and learn. Go to their conferences and, in one example, make cotton candy and run a stand. Go mountain biking and ask what is actually hard about the day job. “You cannot co-design software with the services business over Zoom or over a support ticket,” he says — his closing line is that “you have to touch some grass.”
What happens next
The talk’s honest gap is that bottom-right quadrant. Shenoy says the async and forking mechanism for code is figured out — spin up sandboxes, work in parallel — while the services equivalent is unsolved. He hedges on the three open questions above — knowledge work as code, parallelizing serial work, the right form factor — framing them as problems he’s working on, not results. If you use tools like Claude Code daily, note he places them at rung two of five. The broader bet — operator-owner beats vendor — is being tested at $6.3 billion scale in corporate travel right now.
Quick poll
What actually blocks AI adoption at your company?
Shenoy's answer: "everyone assumes the usage just shows up... It never does."
FAQ
What are the five rungs of Shenoy’s autonomy ladder? Co-pilot, synchronous agent, asynchronous agent, long-running agent, and AI co-worker. Most vendors sell the last rung, he argues, but teams have to “earn the right to do more.”
Why did coding agents diffuse faster than services agents? Partly because models are trained on code, and partly because engineers already parallelize work. Services work is traditionally serial, and Shenoy treats the async-services pattern as unsolved.
What does “extreme software service co-design” mean? Shenoy’s adaptation of Huang’s “extreme hardware software co-design”: designing products together with the people and processes inside the businesses Long Lake owns. For contrast with a model lab’s own feedback loops, see our writeup on how Anthropic builds.