"We're 500 ms from the chain tip." Measured against what?
A four-second-stale price is a trade already lost, a liquidation missed, an oracle reporting a number that moved. If you sell real-time blockchain data, freshness is the product. Yet the standard way to measure it grades your own homework. Here is how to define a freshness SLA you can actually defend, with a measured example and the six questions an honest one answers.
Who bleeds when data is stale
Freshness sounds like a vanity metric until you follow it downstream. A market-making bot quotes off a price that is four seconds old and gets picked off by someone reading a faster feed. A lending protocol's liquidation engine sees the collateral drop late, fires after the position is already underwater, and books bad debt. An oracle publishes a value that the chain moved past two blocks ago, and every contract that trusted it inherits the error. None of these are exotic. They are the normal failure mode of real-time infrastructure, and they all trace to one upstream promise that was never true: that the data was as fresh as the SLA said.
So freshness is not a nice-to-have number on a status page. For anyone selling real-time chain data, or building on someone who does, it is the product's core claim. Which makes it worth measuring properly.
The vocabulary is the discipline
Reliability has a precise grammar, and it is worth using because it forces the right questions. You choose an SLI, the indicator you measure. You commit to an SLO, the internal target you hold yourself to. You sell an SLA, the external promise with money attached when you miss it. The SLA is always looser than the SLO, so you breach your internal target before you ever breach the contract.
"p90 freshness under 500 ms from the head of the chain" is an SLI dressed as an SLA. It sounds precise. But an SLI is a measurement, and a measurement is meaningless without its reference. The question that decides everything is the one the marketing never answers: 500 ms measured against what?
The trap: grading your own homework
You cannot see the true head of the chain. You see it through providers, node pools, your own ingestion, each lagging the real tip by an amount that changes minute to minute, and none of them able to tell you by how much. So the convenient freshness SLI measures your serving time against your own upstream feed. That is grading your own homework with the answer key you wrote. The day your feed slips ten seconds, the gap between "feed saw it" and "we served it" shrinks, and your dashboard reports that you got faster at the exact moment your data went stale.
I measured how large this reference problem is in practice. For one hour I polled ten public Ethereum RPC endpoints on a 2.5-second loop, recording which served each new block first, across 299 blocks. Ranked by median lag behind the fastest independent observer:
| Endpoint | p50 lag | p90 lag | Blocks won |
|---|---|---|---|
| mevblocker | 419 ms | 1,941 ms | 33% |
| flashbots | 659 ms | 2,128 ms | 26% |
| bloxroute | 1,858 ms | 3,870 ms | 6% |
| nodies | 2,205 ms | 4,678 ms | 8% |
| drpc | 2,428 ms | 5,334 ms | 4% |
| tenderly | 4,286 ms | 6,939 ms | 1% |
Same chain, same second, a 10x spread. The median gap between fastest and slowest is 3.9 seconds; a seventh endpoint, dropped for flakiness, ran 13 seconds behind at p90 while reporting itself healthy. Now the trap is concrete: measure your API's freshness against tenderly and you look four seconds fresher than against mevblocker, for identical data. Pick the flaky one and your p90 improves by thirteen seconds precisely when your data is worst. The reference is not a detail of the measurement. The reference is the measurement.
An honest reference: three instruments
A freshness SLI needs a reference that does not depend on the thing being measured. Three instruments build one, and they answer three different questions.
1. The observer mesh. Poll several independent providers at once, from different operators and regions, ideally with a node of your own in the set. The reference for each block is the earliest sighting across all of them, every timestamp taken on your own synced clocks so no external skew enters the math. Now your SLI is measured against the fastest independent view of the chain rather than against yourself, and the spread between providers is your error bar. The table above is exactly this: the reference is "the fastest of ten," recomputed block by block.
2. The protocol clock. The mesh has one blind spot: every provider lagging together, a regional partition or a shared upstream, so the reference drifts while everything agrees. Slotted chains close it for free, because the chain publishes its own clock. Ethereum stamps every block with a slot time, 12 seconds apart from a known genesis, computable from wall clock alone. Compare the mesh against it and you can distinguish "we are behind the tip" from "we have lost the tip." Those are different incidents, and folding the second into a percentile is how an SLO lies. (In my run the mesh tracked slot time to within about two seconds, roughly the poll interval itself, so that comparison was floor-limited by my own sampling. The endpoint-versus-mesh numbers above need no such caveat: they are relative timings on one clock.)
3. Heartbeat transactions. Block timing proves data arrives. It does not prove the data is correct. Send your own transaction on a schedule and require it to come out the far end of your API, findable by hash, having survived ingestion, decoding, and indexing. On a quiet chain the blocks are empty, every block-level metric stays green, and the transaction path can be broken for an hour with nothing to catch it. The heartbeat catches it, on schedule, with a known right answer. The one boundary that keeps this honest: the segment you own starts when the block exists, at its slot time, not when you broadcast. Mempool and inclusion belong to the chain, not your SLA.
Why the number becomes worth trusting
Here is what makes a freshness SLA different from almost every other promise a vendor makes: the blockchain is a public event stream, so the customer can check you without your cooperation. They run two observers of their own, consume your API, recompute your p90, and compare it to the number you published. If you also export the per-block measurement table, an auditor can reconstruct your headline figure and see what your exclusions hid. And the heartbeat transactions are ordinary on-chain objects anyone can verify.
So the vendor who publishes the measurement method first is not giving away an edge. They are converting an audit their customers can already run into a reason to sign. For the buyers who put SLAs through procurement, a number they can verify beats a number they have to trust, and the sophisticated ones have started to ask.
The checklist
A freshness SLA that means something answers all six of these in writing:
| Question | The honest answer looks like |
|---|---|
| Against what reference? | Earliest sighting across N independent observers, cross-checked against the protocol clock |
| Whose clocks? | The measurer's own, synced; producer time only as slot arithmetic |
| Which segment? | Slot time to API availability; chain inclusion excluded and reported separately |
| Blind windows? | Degraded observability detected and reported apart from the p90, never folded in |
| Blocks or transactions? | Heartbeat transactions prove the decode-and-index path, not just block arrival |
| Can I reproduce it? | Measurement table exportable, methodology public, heartbeats visible on-chain |
Most freshness claims on the market answer none of them. That is not usually a lie; it is a number whose meaning nobody pinned down. The teams that pin it down first will have something rarer than low latency: a promise their customers can check.
Written from production experience running multi-chain ingestion and the observability around it. Related: what a million LLM tokens actually costs · EU AI Act logging obligations.