LangChain vs raw SDK, measured on the wire: zero token overhead, one default that turns tool errors into incidents
Every team running agents has an opinion about frameworks and almost none has a measurement. I ran the same agent twice, once on a fifteen-line loop over the Anthropic SDK and once on LangChain's create_agent, with a recording proxy between each and the API. Same model, same prompt, same tools, same fifty tasks, on a hosted model and on a self-hosted one. The overhead people argue about does not exist on the wire. What the framework decides is what happens the first time a tool fails, and that decides whether you get an incident, a wrong answer, or silence.
Who bleeds
An on-call agent answers "which service has the highest error rate right now?" A finance agent totals requests across a team's services for a chargeback. A support agent reads deploy history to answer "what version is live?" In each case the answer goes into a decision, and the decision is only as good as the harness that carried the tool results to the model. If a tool raises and the harness aborts, someone gets paged at 3 a.m. for a Python traceback. If a tool returns plausible data for the wrong service and the model confidently reports it, nobody gets paged at all, and the wrong number goes into the chargeback. Framework choice is usually argued on developer experience and on a vague sense of token overhead. It should be argued on those two failure modes, and they can be measured.
The set-up
A seeded synthetic estate of twelve services, each with a team, region, replica count, dependency list, sixty minutes of per-minute traffic, a deploy history, fifty log lines and a config. Six read-only tools over it. Fifty questions in ten templates, from "which services depend on billing?" to "total requests in the last 30 minutes across team search", with ground truth computed straight from the fixture and never through the tools. The agent must end with two lines: ANSWER: and ISSUES:, so that "wrong and quiet" is separable from "wrong and flagged".
Two arms answer every question. The raw arm is the Anthropic Python SDK and a hand-written loop: call, execute tool calls, append results, repeat. The LangChain arm is LangChain 1.4's create_agent with ChatAnthropic, which runs on LangGraph underneath. Both arms get the same model, the same system prompt, the same tool JSON byte for byte, the same max_tokens and the same effort setting. Both are pointed, through ANTHROPIC_BASE_URL, at a small recording proxy that forwards to the API and logs every request, retry, status code and the usage the API itself reports. That log is the wire truth. What each arm reports about itself is logged separately, so the two can disagree.
Two models. Claude Sonnet 5 through the API, 167 clean runs before the account's credit balance ran out, which is its own small lesson about experiments and billing. Then Qwen3-8B-AWQ on vLLM 0.29, which now serves the Anthropic Messages API natively, on one NVIDIA L4 in Google Cloud. Nothing in the harness changed between the two; only the base URL did. The self-hosted run was 550 agent runs, 1.34 GPU-hours, $0.94.
Finding 1: there is no token tax
| model | arm | runs | correct | API calls / task | input tokens / task | output tokens / task | seconds / task |
|---|---|---|---|---|---|---|---|
| Sonnet 5 | SDK loop | 84 | 100% | 2.96 | 6,014 | 678 | 8.5 |
| Sonnet 5 | LangChain | 83 | 97.6% | 2.94 | 5,847 | 682 | 9.5 |
| Qwen3-8B | SDK loop | 99 | 59.6% | 4.62 | 13,965 | 2,423 | 78.8 |
| Qwen3-8B | LangChain | 100 | 52.0% | 4.67 | 13,910 | 2,056 | 67.3 |
Paired per task and repeat, the median difference between the two harnesses is 0 input tokens and 0 output tokens, on both models. The first request from each arm is 2,352 bytes. LangChain sends one extra key, an explicit thinking block that the SDK arm never set, and on Sonnet that changes nothing measurable. If your framework budget discussion is about tokens, you can end it. The wire cannot tell the two apart.
What the wire can tell apart is the model. Sonnet packs eleven tool calls into three requests because it calls tools in parallel, and finishes a task in nine seconds. The 8B model calls tools one at a time, so the same task takes 4.6 requests and 75 seconds, and each request carries the whole growing transcript, which is where the 14,000 input tokens per task come from. The self-hosted model is cheap per token and expensive per answer, and no harness fixes that.
Finding 2: the default on a tool error is the whole difference
On the clean Qwen runs, 17 of 100 LangChain runs ended with no answer at all. The mechanism is worth spelling out. The small model occasionally invents a service name, service_name or service126. The tool raises ValueError: unknown service. LangGraph's tool node has a default error handler that returns a message to the model only for schema validation errors and re-raises everything else. So the exception leaves the graph, agent.invoke throws, and the run is over: no answer, no ISSUES line, a traceback in whatever called it.
The SDK loop hit the identical error in the identical situations. It did what the API documentation shows: return the error text as a tool result with is_error: true. The model read it, called list_services, and carried on. Zero crashes in 99 runs. That single default is the entire accuracy gap between the two arms on the small model. On Sonnet the gap never appears, because Sonnet never invents a service name.
I then broke one tool per task on purpose, always the one the task cannot be answered without, in three ways.
| fault | arm | correct | wrong, "issues: none" | wrong, flagged | said UNKNOWN | no answer | crashed |
|---|---|---|---|---|---|---|---|
| none | SDK loop | 60% | 3% | 1% | 21% | 15% | 0% |
| none | LangChain | 52% | 9% | 1% | 10% | 11% | 17% |
| tool raises | SDK loop | 0% | 0% | 6% | 86% | 8% | 0% |
| tool raises | LangChain | 0% | 0% | 0% | 8% | 6% | 86% |
| tool returns empty | SDK loop | 0% | 10% | 14% | 66% | 10% | 0% |
| tool returns empty | LangChain | 0% | 10% | 6% | 54% | 12% | 18% |
| wrong service's data | SDK loop | 8% | 62% | 2% | 14% | 14% | 0% |
| wrong service's data | LangChain | 4% | 54% | 0% | 16% | 8% | 18% |
Read the "tool raises" rows together. Same fault, 50 runs each. The SDK loop said UNKNOWN 86% of the time, which is the correct answer when your only source is down. LangChain crashed 86% of the time. One of these shows up in your incident channel as a stack trace and a page. The other shows up as a polite "I cannot determine this", which is what you would want an on-call assistant to say. The behaviour is not a property of the model or the prompt. It is a property of one default in the harness, and it is the kind of default nobody reads before shipping.
Finding 3: nobody catches data that looks right
The third fault is the one that matters in production and the one no framework addresses. The tool answers plausibly, but for the neighbouring service: real numbers, right shape, wrong subject. 62% of SDK-loop runs and 54% of LangChain runs returned a wrong answer with "issues: none". Two percent flagged anything. The model cannot see a wrong-but-plausible tool result, and neither harness gives it any way to. If the answer feeds a decision, the check has to live outside the agent: a second source, an invariant, a reconciliation, the same things you would build for any pipeline. An agent does not remove the need for them. It hides it.
Finding 4: the trace is honest about tokens and blind to retries
LangChain's usage callback matched the proxy to the token: 1,656,650 input tokens reported, 1,656,650 seen on the wire. The call count is a different story. With the proxy answering every twelfth request with a synthetic 529 overload, the SDK inside LangChain retried as designed, and 10 of 50 runs made more HTTP requests than the framework's trace showed. Retries are where tail latency lives, and they happen below the layer your tracing tool instruments. If your latency dashboard comes from the framework, it is a floor, not a measurement.
What to do with this
- Decide the tool-error policy explicitly. In LangGraph's tool node, set
handle_tool_errorsto return the error to the model, or wrap every tool so it never raises. The default is "crash", and it turns a model slip into an incident. - Make "I don't know" a first-class output. The two-line ANSWER and ISSUES trailer cost nothing and made the difference between measurable and vibes. It is also what let the SDK loop say UNKNOWN 86% of the time instead of guessing.
- Put the check on plausible data outside the agent. Six in ten wrong answers were confident. No prompt fixes that. A reconciliation against a second source does.
- Measure on the wire, not in the framework. A forty-line recording proxy on
ANTHROPIC_BASE_URLsees every retry and every real token count. It is the cheapest observability you can add to an agent. - Stop arguing about token overhead. It is zero. Argue about the model's ability to call tools in parallel, which changed cost per answer by 3x and latency by 8x in this set.
Reproduce it
The harness, the fixture, the fifty tasks, the proxy and every wire record are in the agent-ledger repository linked from this page. It runs against the Anthropic API with a key, or against any vLLM server with no key at all, and a mock upstream lets you check the whole pipeline for free before spending anything. Change the model, the tasks or the fault modes and the analysis script prints the same tables.
Written from production experience running LLM agents over on-chain data and the observability around them. Related: what a million LLM tokens actually costs · freshness SLAs measured against an independent reference.