We benchmarked an 84% token reduction. Then we open sourced the protocol.
Building ACP, and sitting with the part we can't solve.
Martina Zrnec
Sep 28, 2026 · 5 min read
I was watching an agent answer a simple question.
The question was small. Three sentences would have covered it. The agent loaded the page, parsed the HTML, and waded through nav bars, footer links, a cookie banner, a sticky subscribe modal, three paragraphs of preamble, before it finally found the part it needed.
Twenty thousand tokens.
For three sentences.
And the thing is, this is happening everywhere right now. Quietly. Constantly. Every agent, every query, every page. We've handed agents a web that was built for human eyeballs and asked them to make it work.
It does. Kind of. Expensively.
The shape is wrong
The web was built for browsers. That's not a complaint, it's just a fact. Humans scroll. We scan. We skip the boilerplate without thinking about it because our eyes know what nav bars look like.
Agents don't get that for free.
They read the whole thing. They have to. There's no shortcut, no "give me the relevant part" channel, scaffolding and all. Every header, every analytics script, every footer link in twelve languages. The cost gets paid in tokens, latency, and the slightly absurd reality that an agent might burn more compute on parsing your nav menu than on actually thinking about your content.
This isn't a performance problem. Performance problems get fixed with caching and faster parsers.
This is a shape problem.
The shape of the web doesn't match how agents read it. And no amount of optimization fixes a shape mismatch.
So we built a shape
It's not a framework. It's not a platform. It's a shape/protocol. A structured envelope you pre-compute and serve first.
The idea is simple enough that it almost feels too simple. You take your content and you do the work once: summary, tags, key entities, classification, provenance. You stamp it with what produced it and when. You persist it. And when an agent shows up asking, you hand it the envelope before the body.
Content gets broken into atoms, discrete units with stable IDs. An agent that needs one specific atom can ask for that atom. Not the page. Not the full body. The atom.
The envelope doesn't replace the source. It sits in front of it. The body is still there if anyone wants it. But most of the time, agents don't need the body. They need the answer. And the envelope is the answer-shaped layer that's been missing.
Built on top of MCP. Designed to complement protocols already in motion, not replace them. Open spec, MIT licensed, npm package out. The point isn't to own the layer. The point is for the layer to exist.
And then we rebuilt Stacklist for it
We didn't just publish the spec.
We rewrote our own product around it.
Stacky - our MCP server, now serves ACO envelopes by default. Every card in Stacklist has an envelope sitting in front of it. The enrichment runs as a background job: a trigger flips a dirty flag, a queue worker picks it up, the envelope gets persisted. Async, out of the write path, no real-time tax. By the time an agent asks, the envelope is waiting.

We did this because we needed to feel it. A spec describes a shape. A product has one. Those are different things, and you only learn the difference when you're staring at a database migration deciding whether the envelope is one column or its own table.
So now Stacky talks to agents the way we wished the web talked to agents. And we can actually measure what that costs, or doesn't.
The numbers
Go ask Stacky about Wikipedia's article on Artificial Intelligence.
Full body read: ~25,000 tokens.
ACO envelope read: ~350 tokens.
Savings: ~99%.

That's not a benchmark we ran in a notebook. That's a real query against a real page through the real product, right now.
The savings aren't marginal. They're the kind of difference where the question stops being “is this worth doing” and starts being “why isn't everything shaped like this already.”
The part I can't solve
Here's where I have to be honest.
Every envelope is stamped. Tool, version, timestamp. You can see what produced an envelope and when. That's the provenance layer, and it's real.
But the envelope claims to faithfully represent the content underneath it. And “faithfully represents” is partly a technical statement and partly a social one.But the direction matters.
What stops someone from publishing an envelope that says one thing while the body says another? What does adversarial enrichment look like? Who watches the enrichers? When an agent reads the envelope and skips the body, which is exactly the efficiency we want, what happens when the envelope is lying?
I don't have a clean answer for this. There are partial ones. Signed envelopes. Verifiable enrichment chains. Reputation layers on top of registries. Each of those is real work, and each shifts the problem rather than solving it.
The honest version is: we built a shape that makes the agent web meaningfully more efficient. We did not solve trust. We made it more visible, which is something, but visible isn't the same as solved.
This is the part I keep sitting with.
The efficiency is real. The shape works. The numbers hold up - in benchmarks and in our own product. And underneath all of it is a question: what does “faithfully represents” mean when the reader has stopped checking? That I think is the actual hard problem of the agent web, and I don't think any of us have answered it yet.
So I'm going to keep building. And keep sitting with it.
Both at the same time.