Research
Measured, not promised
Every claim on this site traces back to a benchmark, a trace study or an invoice we read line by line. The numbers, the method and the limits, in the open.
- Compressing by 20 % is free. Compressing by 50 % is not.−20 %tokens, same answers
- Where a tool-using agent's money goes64 %of cache reads are tool results
- One real invoice: $2.2k a month of LLM, and what compression gets back10 %cache-hit rate found
How we work
- On the invoice, not in tokens
- A saving that does not show on the provider bill is not a saving. Output tokens, reasoning tokens and cache pricing are part of every number we publish.
- Quality with a judge and a reference
- Every compressed answer is scored against a raw answer on the same document. We report the drop, the language and the question type where it happens.
- Nothing read that should not be
- Trace studies use usage metadata only. Audits use aggregates. Customer code and data never enter this repository.
Stop paying for filler. Compress every prompt on one fast, deterministic layer.
Tokeen is in private beta. Request access and see the savings on your own prompts within the week.