Cutting LLM inference costs: the three levers, in the order they pay
Same rented GPU, same model. $4.18 per million output tokens one way, $0.085 the other. The difference wasn't hardware, quantisation, or a smaller model — it was scheduling. Measured numbers below, reproducible with one script.
The setup
One spot NVIDIA L4 (24 GB) on GCP rented at $0.25/hour. Qwen3-8B served with vLLM. A fixed transaction-labelling prompt (~100 input tokens, 256 max output), fired at the OpenAI-compatible endpoint at four concurrency levels. Cost per million output tokens computed from the actual hourly rate and the measured aggregate throughput — no vendor math, no MFU theory, just the meter and the clock.
Cold start, for the record: 3–4 minutes warm (35 minutes on the very first start, CUDA-wheel fights included) from vllm serve to first token — model download excluded, weight load and graph capture included. Remember that number; it's why "just scale to zero" isn't the free lunch it sounds like.
Lever 1 — utilisation. The one that isn't ML at all.
| Concurrency | TTFT p50 | Aggregate tok/s | $ / 1M output tokens |
|---|---|---|---|
| 1 | 0.07s | 16.6 | $4.18 |
| 4 | 0.14s | 62.7 | $1.11 |
| 16 | 0.18s | 243.3 | $0.29 |
| 64 | 0.33s | 815.0 | $0.085 |
At concurrency 1, self-hosting is a bad deal: $4.18 per million output tokens, against a provider floor of roughly $0.455 per million — from the single provider currently serving it for the same model on OpenRouter. Rent a GPU, serve one request at a time, and you are paying a premium for the privilege of doing a provider's job worse than they do.
At concurrency 64, the same silicon produces 49× the tokens. Nothing about the model changed. vLLM's continuous batching keeps the GPU's compute units fed instead of idle between decode steps, and the cost per token collapses to $0.085 — already five times below the only API price for this model.
Lever 2 — quantisation. Cheaper, faster, and not the same function.
Same benchmark, Qwen3-8B-AWQ (4-bit AWQ) instead of the fp16 build:
| Build | Weights | Aggregate tok/s (c=64) | $ / 1M output tokens |
|---|---|---|---|
| fp16 | ~16 GB | 815.0 | $0.085 |
| AWQ 4-bit | ~5.7 GB | 1,698 | $0.041 |
Two effects stack. Decode is memory-bandwidth-bound, so quarter-size weights move faster. And the ~10 GB of freed VRAM becomes KV cache, which raises the concurrency ceiling — quantisation doesn't just cut cost directly, it amplifies lever 1.
The part the pricing tables skip: a quantised model is a different function wearing the same name. I ran the same labelling prompt 50 times through both builds at temperature 0. All 50 verdicts changed. The fp16 build answered "benign" fifty times out of fifty. The AWQ build answered "suspicious" fifty times out of fifty. Not one identical completion between them. Each build is perfectly consistent with itself — and they contradict each other on the decision. On a borderline input, quantisation didn't add noise; it moved the boundary, and at temperature 0 the flip is total. If you quantise, that delta is part of your product whether you measured it or not — and note that the cheapest API routes for open models are themselves listed as fp8 and int4. You may already be running the other function.
Lever 3 — distillation. The biggest one, and the one you don't pull blind.
The remaining lever is replacing the model itself: fine-tune a 1–3B student on your task using your bigger model's (or reality's) labels, and serve that. Done right it's another 5–20× — a small model doing one job matches a general model on that job at a fraction of the compute.
Done blind, it's how you ship a confidently wrong model with a lower bill. Distillation requires three things: thousands of labeled examples, a training run, and — the part that gets skipped — an evaluation proving the student matches the teacher on your distribution, not a benchmark's. Without the third, you have no idea what you traded for the savings. That's not a hypothetical: the quantisation diff above changed 50 answers in 50, and that was a mild intervention.
My own labels are accumulating right now: an autonomous monitoring agent whose predictions get graded against on-chain reality on a delay. Every resolved prediction is a training pair with ground truth attached. When there are enough of them, the student model gets trained — and evaluated against the same reality, not against the teacher's opinion of itself.
Where the break-even sits — and why it's a market question
Here's where it gets interesting. The spot L4 costs $6 a day and tops out around 147M output tokens a day at full batch. Now compare two markets for the same class of model:
| Model | Providers on OpenRouter | Cheapest API price | Self-host beats it above |
|---|---|---|---|
| Llama-3.1-8B | 5 (floor: fp8 builds) | $0.040 /M | never — that's ≥100% utilisation |
| Qwen3-8B | 1 | $0.455 /M | ~13M tokens/day ≈ 9% utilisation |
My saturated spot L4 lands at $0.041 per million — the crowded market's floor to the third decimal. That is not a coincidence. Five providers competing on the same open weights have already done everything in this article — batching, quantisation, cheap capacity — and bid the price down to the metal. You cannot beat them with the same GPU they're using; the floor is the saturated-GPU price.
But a model with one provider carries a 10× margin over that same metal, and there self-hosting wins from 9% utilisation. So the build-vs-buy answer isn't about your hardware at all: it's about how crowded your model's provider market is. Check the endpoint list before you check the GPU prices.
For the record, my own agent's workload sits below every one of these lines, and it stays on the API. Measuring all of this cost $0.77 in GPU rental, per the billing export. The conclusion — don't self-host at my volume — was worth exactly that much.
Method: vLLM 0.28.0 on a rented spot NVIDIA L4 (24 GB) on GCP ($0.25/hr), Qwen3-8B, fixed ~100-token prompt, 256 max output tokens, streaming; TTFT and aggregate throughput measured per concurrency level; cost = hourly rate ÷ (3600 × aggregate tok/s) × 10⁶. Harness and raw results: a single stdlib-Python file — email hello@staysup.io or see the LinkedIn thread. Related: what the EU AI Act actually requires from your logging stack.