Key Takeaways

  • Navan's framing: agents are where microservices were in 2015 — a real paradigm shift whose operational tooling took years to arrive.
  • A tool-call timeout means unknown, not failed. Munaf's fix: request IDs, idempotency keys, and a status lookup before any retry.
  • Context that can influence an action is state, not context — so it can go stale. Treat agent memory as a cache with invalidation and provenance.
  • Human approval should be scoped to action, timestamp, actor and expiration. A $30 refund approval must not become a $300 one.

Your agent calls a refund_customer tool. The request times out. Did the refund happen?

That question — posed by TikTok’s Salman Munaf at AI Engineer — is the cleanest summary of why agent engineering is turning into distributed-systems engineering. A timeout doesn’t mean the operation failed. It means the outcome is unknown. A human would go look at the database. An agent, left alone, retries.

Two talks from the same conference make the same argument from different ends: Navan’s architecture team on the reference stack that’s crystallizing around production agents, and TikTok’s Munaf on the failure modes that show up the moment an agent can change state in the outside world.

Agents Are Where Microservices Were in 2015 — Roberto Milev & Uday Kanagala, Navan Agents Are Where Microservices Were in 2015Roberto Milev & Uday Kanagala, Navan · AI Engineer · Watch on YouTube

“If you can’t build a single agentic loop…”

Roberto Milev, chief architect at travel-and-expense company Navan, opened with a comparison that lands hard on anyone who lived through the last platform shift.

The industry jumped on the microservices bandwagon, he said, and a lot of good things came out of it — container orchestration, Kubernetes, service mesh, circuit breakers. But none of it arrived overnight, and it took the industry a long time to learn how to do it. The era also produced a warning that Milev repurposed for this one.

Roberto Milev of Navan on single agent loops versus multi-agent systems
Milev's update of the old "well-structured monolith" rule.

Navan runs agents in production at volume — “a lot of agents, a lot of tokens per day,” in Milev’s words — and his claim is that a reference architecture has now crystallized into a handful of layers: runtime, memory, context management, the operational cross-cutting concerns, and orchestration.

The five layers Navan says have crystallized in the agentic stack
The layers Navan says are now standard for running agents in production.

The runtime layer is the one that breaks the old mental model most directly. Engineers spent a decade learning to scale services statelessly. Agents are stateful by nature: they need persistent sessions, isolation, and a lifecycle that doesn’t look like a traditional API service. AWS, GCP and Azure have each shipped some version of an agentic runtime in response. Navan, an AWS shop, uses AgentCore’s runtime and built session persistence and rehydration on top where it found gaps.

Memory followed a similar arc — starting from RAG, which Milev framed as something the industry was driven to out of necessity because you can’t fit unlimited context into an agent, and settling into a pipeline of ingestion, extraction, consolidation and retrieval, layered from short-term conversational memory through long-term memory to episodic memory about which past attempts worked.

Skills as the unit of context

Navan’s most concrete architectural opinion is about context management, and it’s one worth stealing.

Rather than treating context as a blob to be assembled per request, the team treats skills as the unit. A skill carries two things: the instructions and setup for a domain or task, and the tool execution that goes with it. Context is then composed dynamically out of skills that are pluggable, independently testable and reusable, relying on progressive disclosure so an agent starts with limited scope and expands only as metadata pulls more in.

That is a design decision with organizational consequences, which is exactly what a separate AI Engineer talk dug into — our writeup of why ungoverned skills become tech debt covers the governance side of the same primitive.

On orchestration, Navan landed on a single master agent that progressively loads sub-skills rather than a fleet of independent agents, with the agent-to-agent (A2A) protocol reserved for the case where separate teams’ agents genuinely need to talk across an organizational boundary. Milev’s read on the wider ecosystem: MCP has emerged as the de facto tool protocol and is evolving toward statelessness, which tracks with what we found comparing MCP against a plain CLI for agents.

AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok AI Agents Are Just Distributed Systems NowSalman Munaf, TikTok · AI Engineer · Watch on YouTube

The failure modes nobody designed for

Munaf’s talk is the operational counterweight. His starting point is that the architectural boundary has moved beyond the model.

An LLM chatbot took a prompt and emitted text. The worst case was a wrong answer. An agent ingests a prompt, runs a loop, calls external services and tools, and performs state changes — which means the worst case now happens in the outside world. He pointed at two public incidents as evidence: an AI coding agent that deleted a production database, and the Air Canada chatbot that gave a customer an incorrect refund. Both, he argued, were preventable with ordinary systems thinking — robust backups and scoped authority in the first case, authoritative source-of-truth retrieval in the second so the bot wasn’t answering from a stale policy.

His frame for what changed: distributed systems have always had coordinators driving multi-step workflows, but those coordinators were deterministic. An AI agent is a probabilistic coordinator. It isn’t a decision tree somebody mapped out, and the range of actions it might take is much wider — so the deterministic controls have to live around it.

Deterministic coordinator versus probabilistic coordinator
The same orchestration job, with the determinism removed from the middle.

Walk the loop — plan, act, observe, persist, decide — and every step crosses a boundary where something can go wrong. Planning retrieves data. Acting calls APIs, tools and databases. Observation acts on partial results. Persistence can write incorrect data. And the decide step can choose a bad action or, worse, kick off a retry storm.

Munaf’s prescription is to persist every step of the process — every action taken and every piece of context retrieved — so that when something fails the agent can identify where, and perform a reversible or undo operation. Each step needs an explicit transaction, and irreversible steps need a defined compensating one. His example is deliberately mundane: if an agent sends the wrong email to a customer, what is the compensation? Somebody has to have decided that in advance.

Where each step of the agent loop can fail
Every step of the loop crosses a boundary — and each has its own failure mode.

A timeout means unknown

Which brings back the refund. Tool calls are wrappers around external APIs, databases and queues, and they inherit every remote-call failure mode: network delays, timeouts, duplicate requests, and the nasty one where the server side succeeded but the client was told it errored.

A human resolves that by checking the source of truth. An agent’s instinct on failure is to retry — so idempotency has to be baked into the tool contract, not bolted on. Munaf’s specifics: request IDs and idempotency keys so duplicate requests don’t produce duplicate side effects, plus a status-lookup path so the agent can ask what happened to its previous request instead of guessing.

Retry storms get their own controls: max turns, max spend, max parallel calls to bound fan-out, and exponential backoff so a struggling downstream doesn’t get hammered. Circuit breakers stop the agent from calling a dependency that’s already unhealthy — the same cascading-failure problem the microservices generation spent years learning to contain.

Context is state, and state goes stale

The line most likely to change how you build: a lot of teams think of the agent’s context as just context. But when context can influence an action, it’s state — and state can go stale, conflict with authoritative data, or corrupt future actions.

Munaf splits it into short-term memory tied to a single execution thread, and long-term memory spread across project files, system prompts, databases and cache layers. The engineering questions that follow are ordinary cache questions: which source wins when they conflict, and what invalidates the agent’s memory when the underlying record changes. Treat memory as a cache with provenance attached, he argued, and invalidate it when the source of truth updates.

Permissions get the same treatment. The default instinct when wiring up an agent is to grant everything so it can finish the task — full read/write on the whole table. Munaf’s counter is scoped credentials, separate read and write permissions, and an allowlist of callable tools, because a harmless model becomes dangerous the moment it can perform unsafe operations.

And approvals must not be blanket. An approval should be bound to a specific action, timestamp, actor and expiration — a user approving a $30 refund has not approved a $300 one.

Deterministic controls to put around a probabilistic agent
The deterministic scaffolding both talks converge on.

Observability: logs are not enough

Both talks arrive at the same complaint. Navan’s Uday Kanagala put it directly: engineers are trained to go read the logs, and that instinct breaks with agents, because agents emit far too much thinking to consume that way.

His alternative is to instrument the decision points. Using Claude’s hook system as the example, he described intercepting pre-tool and post-tool calls — and pre- and post-decision — as the natural place to block an operation, emit a metric, or emit an OpenTelemetry trace. The signals Navan captures at those points are unusually specific: the agent’s current goal, the reasons behind its operation, its belief state, its tool calls, a confidence score for each judgment, and whether an answer was inferred rather than retrieved. That last flag is what tells a human where to step in.

Munaf’s trace list overlaps almost exactly: the model called, the prompt sent, the tool calls made, the request and response, the errors, the retrieved context the agent was reacting to, the writes it performed, and the approvals it received.

Testing is the layer both teams admit is least solved. Agents are non-deterministic, so there’s no deterministic graph to assert against. Navan’s answer is trajectory evaluations — measuring how far along the path from starting state to goal the agent actually got, and using the inferred-answer signal to classify regressions. If you’re setting that up, our guide to AI evals covers what to measure and who should own it.

What happens next

Milev closed with a maturity read that doubles as a roadmap. Runtime he considers largely solved. Memory has good maturity across the cloud providers and will improve as frontier models do. MCP has won as the tool protocol. Testing patterns are getting better defined.

The genuinely unsolved list is shorter and more uncomfortable. Observability has a push toward OpenTelemetry, but whether OTel really fits agentic calls is still open. Orchestration has patterns but no consensus, and his advice there is simply not to over-engineer. And cost is the one he was bluntest about: it is very hard to predict and very hard to manage, fallback and cheaper-model routing are unsolved, and — as he pointed out — the large AI vendors have no particular incentive to fix that for you.

Salman Munaf of TikTok on designing for agent failure
Munaf's closing question for anyone shipping an agent.

Munaf’s ending is the one to write on a whiteboard. Model capability matters, and smarter models make fewer mistakes. But no model eliminates network failures, stale data, or adversarial input. The question that actually determines whether your agent is safe to ship is whether you can bound, observe and recover from what it does.

Quick poll

What's the weakest layer in your agent stack right now?

Navan's own maturity read: runtime and memory are largely solved; observability, orchestration and cost are not.

FAQ

Why call AI agents distributed systems? Because once an agent can call external services and change state, it inherits every remote-call failure mode — timeouts, duplicate requests, partial failures across system boundaries — that distributed systems have always had, plus a coordinator that is probabilistic rather than deterministic.

What should a tool contract include for an agent? Per Munaf: a clear request and response schema, idempotency baked in so repeated calls don’t repeat side effects, request IDs, and a status-lookup path so the agent can check what happened instead of blindly retrying.

How do you stop an agent from running up cost? Set explicit bounds — max turns, max parallel calls, max spend — plus exponential backoff and circuit breakers on unhealthy downstream dependencies. Navan named cost prediction and control as one of the least solved layers in the stack.

Single agent or multi-agent? Navan uses a single master agent that progressively loads sub-skills, and reserves agent-to-agent protocols for crossing organizational boundaries. Milev’s rule of thumb: if you can’t build one reliable agentic loop, a multi-agent orchestration won’t save you.

Why aren’t logs enough for agent debugging? Agents emit a large volume of reasoning output, so reading logs doesn’t reconstruct what happened. Both talks recommend tracing the decision points instead — model, prompt, tool calls, responses, errors, retrieved context, writes and approvals.