Key Takeaways

  • Manuja frames every LLM gateway as a four-way tradeoff between availability, latency, guardrails and cost — during a degradation you must pick which to sacrifice.
  • With reasoning models, he says the same prompt can take anywhere from 2 seconds to 60 seconds, and his team saw production P99 pop to 60 seconds "for no good reason."
  • Missing per-model-class, per-route timeouts are "the number one root cause of your silent outage," and gateway-wide aggregate latency is a metric he calls "a lie."
  • He recommends against one central company-wide gateway: most teams want centralized governance, which does not require centralizing traffic through a single point of failure.

Every time an app tells you “Something went wrong. Please try again,” there is usually a fairly complex system behind that message deciding it was the best remaining option. That was the opening move in Kanish Manuja’s AI Engineer talk on productionizing LLM gateways, and his central claim is that a gateway is less a routing layer than a permanent argument between four things you cannot all win. Manuja, a principal engineer at Twilio, put it plainly: at the heart of the gateway is a fight between availability, latency, guardrails and costs, and “in case of a degradation, you cannot maximize all four. You need to pick what you want.”

Productionizing LLM Gateways — Kanish Manuja, Twilio Productionizing LLM GatewaysKanish Manuja, Twilio · AI Engineer · Watch on YouTube

Retries are a habit from cheap APIs

Manuja’s definition is deliberately boring: an LLM gateway is middleware between your apps and the model providers behind them, handling routing, authentication, fallback, rate limits and “all kinds of governance that you can think of.” Start with availability. If you run a single model provider, Manuja says, “their ceiling is your ceiling. Their outage is your outage.” The reflex from ordinary distributed systems is to retry with exponential backoff and jitter, then trip a circuit breaker once you have seen enough failures.

That reflex transfers badly to LLMs. Retrying eats your latency budget fast, because the calls are slow. Tripping a breaker makes little sense when another perfectly fine provider is sitting right there. And because the calls are expensive too, blind retries multiply both cost and tail latency.

His preferred pattern is per-request fallback: try provider A, and if that request fails, try provider B in sequence. Firing at both in parallel is an option, but only if you are, in his phrasing, “highly highly obsessed with latencies,” because it doubles your cost. Circuit breaking still has a role: if the primary has been failing for a while, take it out of the request path, put it in a cool down, and try again after a few minutes.

One design decision he flags as genuinely open is where failure counts live — in memory per instance, or shared fleet-wide. Fleet-wide state gives quicker failovers; instance-local counters get awkward the moment your deployment size changes.

Fallbacks are not transparent

The clean architecture diagram, Manuja notes, hides the gotchas. Chief among them: fallbacks are not transparent. While the industry is converging on an OpenAI-API-compatible format, he says there are still nuances — differences in tool calling schemas, token limits, stop reasons “and what have you.” A gateway can absorb that with a normalization layer, but you have to test the fallback path deliberately rather than assume format compatibility means behavioral compatibility. Checking that provider B produces acceptable output on the same prompts is exactly what a proper evaluation setup is for.

Then there is streaming. Nobody wants to wait 30 seconds for a wall of text, so for many use cases it is non-negotiable. But Manuja is blunt about the price: “you trade away your levers.” Once you have started streaming from provider A you cannot switch mid-stream — whatever reached the client is done. That, he says, is the real origin of the error message he opened with: “It’s not because of laziness. It’s by design.”

His most pointed operational warning is about the second provider. Teams provision and test the primary carefully, and then the fallback “doesn’t necessarily get the same level of love.” He argues headroom on the fallback should be even higher, because it is your last line of defense. If it goes down, your application goes down.

Latency is the quiet failure

Availability failures announce themselves — things fail, alarms fire, someone gets paged. High latency, Manuja says, is the quiet one, and it deserves more attention than availability tuning usually gets. A gateway typically runs mixed workloads: embeddings under a second, classification under a second, chat around three seconds, reasoning requests taking a long time. He asked the room who measures aggregate latency for the whole service, then called it a trick question. You shouldn’t. “It doesn’t make sense. It’s a lie.” Track P99 per model per route, not a gateway-wide figure.

The related instruction is timeouts, set per model class and per route. Without them, “your gateway thinks your request is being happily served while it is not” — the number one root cause of a silent outage. His summary line is worth pinning above the dashboard: “a reasoning model’s normal is actually a chat model’s outage.”

Reasoning and router models, the scariest slide

Manuja called this the slide that has given him the most scars. Reasoning models make latency genuinely unpredictable: you often cannot set temperature to zero, and the same prompt can take anywhere from 2 seconds to 60 seconds. He reported production P99 “suddenly popping to 60 seconds for no good reason.”

He offers no magic fix. He recommends fixing the reasoning level per route, and, with router models — which hide the choice of underlying model behind an abstraction — pushing requests to be as deterministic as you can make them inside a nondeterministic system. The other tool is hedging the tail: if the primary request has already consumed, say, P90 of your latency budget, fire a second one. That is the same instinct as choosing simpler, more predictable plumbing elsewhere in an agent stack, a tradeoff we walked through in MCP vs CLI for AI agents.

Guardrails are just another service that goes down

Guardrails earn their place: prompt injection defense, PII filtering, toxicity filtering, keeping models from swearing at your customers. But Manuja treats them the way he treats model providers — as a dependency that can be unreliable and go down. Which forces the question: fail open, serving the request anyway, or fail closed, blocking it? That is availability versus security in one decision, and he says there is no universal answer — a toxicity filter being unavailable might be survivable. His rule of thumb: “the default choice should be the worst case that you can live with.”

Three things make guardrails less dangerous. First, a time budget — the LLM should be the rate-determining step, never the guardrail, so guardrails run under their own timeouts. Second, fallbacks: secondary providers, secondary checks, cached decisions. Third, placement. A pre-hook on the input is probably safest but adds serial latency. Parallel is his favorite, with a caveat: it does not work well with streaming, so for structured output he suggests not streaming and running checks concurrently. Post-hooks are best for output monitoring and auditing.

The gateway is your newest dependency

Everything above concerns dependencies you already had; the gateway adds one more. Segregate API keys per route and per use case, as granularly as you can imagine — shared limits plus one noisy tenant is among the biggest problems here. Confirm your gateway supports load shedding and exercise it in runbooks and game days, because under a retry storm you cannot simply scale out. Web servers have internal queues; make sure they are bounded. Add traffic prioritization if your most important use cases must survive load.

Centralize governance, not traffic

His closing argument is the one most likely to change a roadmap. A central gateway for an entire company is a single point of failure, and he recommends rethinking it. What he has noticed is that “it’s not the central gateway that they want. They want centralized governance.” Governance — cost tracking, rate limit management — can be delivered through plugins and custom code without funneling all traffic through one deployment. One team can own it; that is different from running it as one deployment company-wide. For organizations still working out who owns AI infrastructure, that pairs with the ownership questions in how to adopt coding agents at work.

What happens next

Nothing here hinges on a model release, which is what makes it durable. The four-way fight is structural, and every capability layered on top — longer reasoning, router abstractions, richer guardrail stacks — widens the latency distribution rather than tightening it. If reasoning-level routing stays opaque, per-route limits and tail hedging start to look less like tuning and more like table stakes.

Manuja ended on a personal note: it was his son’s birthday and, as he put it, “I’m here talking to strangers about circuit breaking.” His ask in return was that the audience go prevent one incident.

Quick poll

When your guardrail service goes down, what should your gateway do?

Manuja says there is no universal answer: "the default choice should be the worst case that you can live with."

FAQ

What is an LLM gateway? Manuja defines it as an entry point or middleware between your applications and the model providers behind them, handling routing, authentication, fallback, rate limits and governance.

Why aren’t normal retries enough for LLM calls? LLM calls are slow and expensive, so retrying eats your latency budget and multiplies cost and tail latency. Tripping a circuit breaker is also the wrong move when a healthy second provider is available; he recommends per-request fallback instead.

Can you fail over to another provider mid-stream? No. Once you have started streaming from provider A, you are committed — whatever reached the client is done. That is why users see generic “something went wrong” messages, which Manuja says is by design rather than laziness.

Should a company run one central LLM gateway? He advises rethinking it, since a central gateway is a single point of failure. In his experience most organizations actually want centralized governance — cost tracking, rate limit management — which plugins and custom code can deliver while the gateway stays decentralized.