Numbers and limits
Every published number with its source, what inferred means, and what caveman will not do.
Every figure published on this site, the word that says how it was obtained, and the four rules the code follows when it counts. This is the one page that carries the limits for all of the local tools, so other pages link here instead of repeating them.
Where a number comes from#
Each local tool labels its own output. Run the engine's committed fixture set and the label is in the report:
caveman-engine evals run{
"passed": true,
"tokens_before": 10577,
"tokens_after": 2416,
"token_ratio": 0.7715798430556868,
"probes": { "total": 12, "retained": 11, "recoverable": 1, "lost": 0 }
}{
"name": "json-tool-output",
"content_type": "json",
"tokens_before": 1366,
"tokens_after": 293,
"ratio": 0.7855051244509517,
"basis": "inferred",
"passed": true,
"byte_recoverable": true
}Captured from a caveman-engine build dated 2026-08-24; caveman setup --install today pins bin-v1.1.7. Trimmed to the fields this page is about. Counting is done offline with the o200k_base BPE, which is OpenAI's tokenizer and an approximation of any other vendor's. Ratios between two arms of the same run are meaningful; the absolute counts are estimates.
Vocabulary#
Three words, fixed meanings, never interchangeable. A summariser that blurs them makes every claim downstream wrong.
inferred#
A local, per-run estimate computed on your own machine from an offline token counter. Every public Caveman tool emits this word. It measures size, and it stays inside the run that produced it.
measured#
Traffic that was actually observed, rather than estimated from bytes. Observing a request tells you what was sent, and says nothing about the request you would otherwise have sent, so measured stops one step short of a saving.
verified#
A saving confirmed against a real bill by a hosted rollout system that ran both arms of the comparison. Caveman Cloud is the only source of it; every tool documented on this site emits inferred instead.
recoverable#
A lossy result whose original bytes were written to a content-addressed handle before the result was emitted, so retrieve(handle) returns the original byte for byte. Order is the point: the store happens first, and the smaller view is emitted only once that write succeeded.
Published numbers#
| Claim | The number, with its source |
|---|---|
| Output style, third party | Adobe Research's CAVEWOMAN paper measured caveman-style output across eight models and five datasets and found it cut realised cost 1.4 to 2.4 times per model, up to 3 times. |
| Output style, third party | JetBrains ran 86 real coding tasks as a paired A/B with the skill alone and found 8.5 percent fewer output tokens with no detectable quality change, sign test p = 0.82. |
| Output style, this repo | On the repo's committed ten question eval, the skill cut output tokens 50 percent at the median against a plain "answer concisely" control. Length only, not correctness. |
| Wrapping a coding agent | In a pinned Claude Code benchmark, wrapping used 33.2 percent fewer provider-reported input tokens across 18 paired runs, 95 percent interval 14.6 to 48.5. One case regressed 9.9 percent. All 18 exact-answer checks passed. Pinned counterfactual evidence, not production traffic and not invoice savings. |
| Engine, one fixture | The committed json-tool-output fixture went from 1,366 tokens to 293 in the run above, and its original bytes stayed retrievable. |
| Content-type targets | Targets, not results: search results 80 to 95 percent, logs 85 to 95, JSON 70 to 90, diffs 60 to 80, text and HTML 50 to 80, code 40 to 70. |
The regressed case in the wrap benchmark is an HTML page for which no compression transform applied, so the wrapping overhead was paid and nothing came back. It stays in the table because a table that drops its losses cannot be checked.
The content-type row is a target, not your result. Measure your own payload instead:
caveman-engine compress < yourfile > /dev/nullThe smaller view goes to stdout and is discarded here; the report goes to stderr, and its ratio is what that file compressed by.
Latency#
Compression costs time on the machine doing it. Both figures below were measured on an Apple M3 Pro with binaries dated 2026-08-24, against the example files this site serves.
Each file went through caveman-engine compress ten times with a scratch CAVEMAN_CCR_DB. The column is the median wall clock of those ten runs, process start included. An empty payload takes 18.6 ms at the median on the same machine, so that is the floor in every row and the compression itself is the remainder.
| Example | Bytes in | Median ms |
|---|---|---|
compressors/toon/small.json | 374 | 20.9 |
compressors/tables/results.md | 1,417 | 19.0 |
compressors/tool-schemas/tools.min.json | 1,830 | 24.6 |
compressors/tool-schemas/catalog.min.json | 1,840 | 21.8 |
compressors/config/service.yaml | 2,559 | 20.4 |
compressors/repetition/pytest.log | 2,609 | 19.9 |
compressors/diff/change.diff | 3,312 | 18.6 |
compressors/toon/inventory.json | 4,151 | 22.7 |
engine/app.log | 4,663 | 19.5 |
engine/orders.json | 5,515 | 20.2 |
compressors/search-results/rg.txt | 5,888 | 22.3 |
compressors/text-and-html/postmortem.md | 7,942 | 29.3 |
engine/request.json | 8,039 | 20.0 |
compressors/text-and-html/article.html | 8,294 | 19.8 |
compressors/tables/orders300.csv | 9,009 | 20.5 |
compressors/code/inventory.py | 9,578 | 23.5 |
compressors/terminal/build.log | 9,741 | 20.2 |
engine/SKILL.md | 12,321 | 34.5 |
compressors/accessibility-trees/axtree.json | 18,200 | 23.0 |
engine/conversation.json | 20,905 | 20.0 |
compressors/json/orders.json | 21,875 | 21.9 |
compressors/logs/app.log | 25,289 | 26.8 |
engine/big.json | 62,082 | 30.0 |
The largest file, 62 KB of JSON, spends about 11 ms on compression. Run to run spread is wide at this scale: the same file varied by 6 ms across its ten runs, and engine/SKILL.md by 44 ms.
The proxy hop was measured separately, with caveman-proxy in record mode on loopback in front of a local stub that answers instantly, so nothing in the number is provider time. Fifty identical 223 byte Anthropic-shaped requests went through the proxy and fifty straight to the stub, each on one keep-alive connection. The stub answered in 0.12 ms at the median and the proxied request in 0.31 ms. Three rounds put the added median at 0.12, 0.19 and 0.28 ms.
Record mode forwards bytes unchanged, so that figure is the hop alone. A request the proxy compresses also pays the engine time above, on a payload the size of the request rather than the whole session.
The four rules#
Per run, never per month. A local tool counts one payload, one session you ran, or the history window you pointed it at. The furthest it projects is a per-day rate inside that window, because anything longer would depend on traffic nobody measured.
Tokens first, money only where the provider counted it. Prices depend on the model, the region, the cache state at the moment of the call, and your contract. Compression figures are token counts and stop there. The provider catalog holds the price table, and the hosted plane is where a price is applied to observed traffic.
Ranges beside averages. A median with no spread hides whether the result is solid. The wrap benchmark publishes its 95 percent interval next to its headline for that reason.
Negative results stay in. The regression above is published at the same size as the wins. So is the finding in the Adobe paper that compressing the human's prompt makes models answer longer and worse, which is why nothing here rewrites your prompts.
What caveman will not do#
What the local tools actually do, and what each behaviour means for a number you read.
| Behaviour | Consequence |
|---|---|
Local tools count tokens with an offline tokenizer and label the count inferred. | Every local figure stays inferred. The label is unchanged by crossing a network. |
caveman learn extrapolates a per-turn figure into a tokens/day rate. | That rate is a projection over scanned history, printed with basis: inferred, and it stops at a day. proxy/internal/store/learn_pricing.go projects no window figure forward to a month. |
Where a scanned transcript carries provider-counted usage, caveman learn prices those tokens against the dated provider catalog at your own measured effective input rate rather than list price. | That is arithmetic on tokens the provider counted, not a saving. A model the catalog does not carry is reported separately under an unpriced: version rather than given a sibling model's rate. |
caveman stats reports token counts first and a price only when it can stand behind one. | On a store with no priceable provider usage, savings_usd reads 0 and would_save_usd reads null. A subscription-authenticated session stays tokens-only. |
| The skill copies exact bytes through and answers security warnings and irreversible-action confirmations in full sentences. | Code blocks, symbol names, CLI commands, file paths and error strings survive. What it leaves alone is the list. |
| A lossy compressor runs only once the original has been written to the recovery store. | Without a store the original is forwarded and the ratio is zero. |
| The engine compares sizes before it emits. | A compressed output that is the same size or larger is dropped, the original goes upstream, and the ratio is zero. |
| The proxy binds to loopback. | A non-loopback listen address is refused unless CAVEMAN_AUTH_TOKEN is set, because nothing else authenticates inbound traffic in front of your provider credentials. |
Count it yourself#
caveman-engine compress < your-payload.txt # one payload, before and after, with the basis
caveman stats # the local store, aggregated
caveman trial -- claude # one session with and without, then trial reportcaveman trial needs its own proxy. Add the local tools has the two commands that go around it when persistent routing is already on.
Two fields in caveman stats need their definitions said once. basis is inferred for every row a local proxy wrote. requests_eligible_for_compression counts the recorded requests that reached the compression path as candidates: the proxy was in compress mode, a way back to the elided bytes existed, and the cache epoch allowed the rewrite. It counts candidates, not wins, so a request that was eligible and came out no smaller is still in it. The gap between it and requests is the traffic the proxy never tried to shrink.
What the CLI reports back to Caveman about any of this, and how to turn that off, is on Telemetry.