Key Takeaways
- Across Arize AI's 500 trials (5 runs × 25 tasks × 4 arms), LLM-judged correctness landed at 0.834 for MCP, 0.833 and 0.826 for two CLI skills, and 0.845 for a bare baseline — a spread of under two hundredths.
- On the hardest tier, Arize reported MCP costing more than six times what the skills cost and taking five times as long; one analysis task ran 71 tool calls, eight minutes, and $2.00.
- Only three of those 71 calls were actually MCP calls — the agent shell-executed bash instead, dropping tool fidelity to 0.33 while the skill arms stayed above 99%.
- MCP still won cleanly on a write task: branch-and-PR creation took 8 calls in 33 seconds for about 16 cents, versus 22 calls and 50 cents for a CLI skill.
Short answer: give your agent a CLI when the tool already runs on the same machine, has years of public documentation behind it, and has auth configured — and reach for an MCP server when the tool is remote, proprietary, needs OAuth, or has real state to manage across steps. That is not a hunch; it is the recommendation Arize AI landed on after running 500 trials pitting the official GitHub MCP server against command-line skills. The surprise in their data is that correctness barely moved. What moved was cost, wall-clock time, and how many turns the agent burned getting there.
What MCP actually is, in one paragraph
Per the official documentation at modelcontextprotocol.io, the Model Context Protocol is an open-source standard for connecting AI applications to external systems — the docs compare it to a USB-C port for AI apps. Architecturally it is client-server: an MCP host (an AI application like Claude Code or Cursor) spins up one MCP client per MCP server, each holding a dedicated connection. The spec splits into a data layer that speaks JSON-RPC 2.0 and a transport layer that handles connection setup and authorization.
The data layer is the part that matters for this decision. MCP defines three primitives a server can expose: tools (executable functions the model can invoke), resources (data sources that supply context), and prompts (reusable interaction templates). Two transports are supported — stdio for local processes, and Streamable HTTP for remote servers, where the docs recommend OAuth for obtaining tokens. Worth noting for anyone building against the current spec: as of protocol version 2026-07-28, elicitation is the client primitive that survives, while sampling and logging are marked deprecated.
The docs are explicit about scope: MCP covers the protocol for exchanging context and does not dictate how an application uses the model or manages what it gets back. That boundary is exactly why a CLI can compete.
What “just give it a CLI” means
The alternative is not exotic. You hand the agent a shell, an already-authenticated command-line tool, and a markdown skill file explaining how to drive it. No protocol, no server process, no schema negotiation. The agent composes commands the way an engineer would — pipe, filter, count — and only the outer loop touches the model.
That composability is the whole argument: deterministic filtering, joins, and aggregation can stay in shell code while the model handles the step that genuinely needs judgment. Arize’s Tier 4 runs make the trade-off visible without needing a second benchmark.
The eval: 25 tasks, four arms, 500 runs
Arize’s setup, written up by Laurie Voss, is worth understanding before borrowing the conclusion. The harness was Claude Opus 4.6 on the Claude Agent SDK — the same SDK behind Claude Code — pointed at a synthetic Python package repo seeded with 12 open issues, three pull requests in different review states, plus milestones, comments, and assignees.
Twenty-five tasks spanned four tiers: simple reads (how many open bugs), complex reads (bugs with linked PRs), writes (create issues, add labels, fix typos), and analysis (milestone completion reports, multi-label patterns). Four arms competed: the official GitHub MCP server; a 2,187-line LobeHub skill covering the full gh surface; a 341-line Vault skill with safety tiers; and a bare baseline with a generic system prompt and no MCP or skill.
Correctness, scored by LLM-as-judge, came back at 0.826 for LobeHub, 0.833 for Vault, 0.834 for MCP, and 0.845 for the baseline. Arize published no significance testing alongside those numbers, so our read is that the ranking is effectively flat rather than a win for going toolless — but it does mean capability was not the differentiator. If you want to run this comparison on your own stack rather than trust anyone’s numbers, the methodology is close to what we describe in our guide to building evals as a team; Arize says the code, tasks, and data are open source.
Where the gap opens: composition
Cost and latency are where the arms separated, and only on the hardest tier. On Tier 4 analysis work, Arize reported MCP costing more than six times what the skills cost and taking roughly five times longer.
The illustrative case is a task computing average time-to-close across milestones. The MCP arm spent 71 tool calls, eight minutes of wall-clock time, and two dollars. The detail that explains everything: only three of those 71 calls were actual MCP calls. The agent, unable to express the aggregation through the server’s endpoints, fell back to shell-executing bash. A skill run on that tier finished in seven tool calls, under a minute, for 19 cents.
Arize’s diagnosis is that a REST-wrapper MCP server is a fixed API surface. When the task maps to an endpoint, it is clean. When it does not, the server cannot compose its way to an answer — but a shell can. That is also why the compact 341-line Vault skill outperformed the 2,187-line encyclopedic one: more instructions is not more capability.
Where MCP won outright
The reverse case is just as instructive. On task 13 — creating a branch and a pull request — MCP finished in 8 tool calls, 33 seconds, and about 16 cents, while the LobeHub CLI skill took 22 calls, close to a minute and a half, and 50 cents. A stateful multi-step write that maps directly onto purpose-built endpoints is precisely MCP’s shape.
There is also a governance argument the benchmark cannot score. Arize flags that CLI auth has no equivalent of the OAuth security model — which matters the moment you ship to customers rather than automate your own laptop. And gh enjoys an advantage no proprietary internal tool will ever have: years of public documentation sitting in the model’s training data.
The fidelity problem to budget for
The number teams underweight is tool fidelity — did the agent stay inside the tools you allowed? The skill arms held above 99%. MCP on Tier 4 scored 0.33, because the agent escaped to bash despite MCP-only instructions.
If your reason for choosing MCP is control — allowlists, audit trails, access boundaries — that finding should reset expectations. Handing an agent an MCP server does not confine it to that server unless the harness actually removes the shell. Whichever AI coding assistant you standardize on, sandbox enforcement is a separate decision from protocol choice.
What happens next
Arize’s own closing framing is that this is not a versus at all: “It’s MCP plus the command line.” Real agents run both, and task shape decides which one handles a given step.
Two things are worth watching. First, MCP servers are still mostly thin REST wrappers; a server designed for aggregation rather than endpoint mirroring would change the Tier 4 result, and Arize says as much. Second, the deprecation of sampling in the 2026-07-28 spec pushes server authors toward integrating LLM providers directly — a nudge toward servers that do more work internally instead of chattering with the client, which is exactly the failure mode the eval exposed.
Quick poll
How do your agents reach external tools today?
For the record: in Arize's Tier 4 runs, only 3 of the MCP arm's 71 tool calls on one analysis task were actually MCP calls.
FAQ
Is MCP slower than a CLI? Not inherently. In Arize’s eval the two were comparable on simple tasks, and MCP was faster on a stateful write. The gap appeared on analysis tasks that required composing operations the server had no endpoint for.
Does a CLI produce worse answers? Not in this eval. LLM-judged correctness spanned 0.826 to 0.845 across all four arms, including a no-tool baseline, so capability was not what separated them.
When should I definitely build an MCP server? Per Arize’s guidance: when the tool is remote or proprietary, absent from training data, needs OAuth, carries real state across steps, or when you want to hide an entire agent behind one simple tool call — particularly for enterprise customers who need access controls.
Does MCP guarantee my agent only uses approved tools? No. Tool fidelity for the MCP arm on the hardest tier was 0.33 because the agent shell-executed bash anyway. Enforcement has to come from the sandbox, not the protocol.