Pricing
Pay for what you compress. Keep what you save.
One meter: tokens that go through the compressor. No seats, no minimums, no percentage of savings. Tokeen is in private beta — these are the plans you land on once you are in.
Your own prompts, your own numbers. No card, no NDA — access by request while in beta.
- 1M tokens compressed a month
- All models, all providers
- Python, Node and AI SDK wrappers
- Playground and dashboard
- Community support
Usage-based. You pay for what goes through the compressor, and you keep everything it saves.
- Everything in Free
- Unlimited volume, volume discounts from 1B
- Batch API, 50k items a minute
- Per-route alerts when a ratio drops
- Priority support, 99.9 % SLA
Run it inside your network, with the paperwork a security review asks for.
- On-premise or your VPC
- Zero data retention, account-wide
- SOC 2 report and signed BAA
- Custom aggressiveness per route
- Dedicated engineer, shared Slack
How the meter works
You are billed on tokens sent to the compressor, at a fraction of any model's input price. Whatever it removes, you stop paying your provider for. Here is a month for a product sending 400M input tokens to a model priced at $2.50 per million.
Compare plans
| Feature | Free | Pro | Enterprise |
|---|---|---|---|
| Compression | |||
| Tokens a month | 1M | Unlimited | Unlimited |
| Prompt, context and conversation modes | |||
| Tool-result mode for agents | |||
| Aggressiveness control | Global | Per request | Per route |
| Multilingual model | — | ||
| Platform | |||
| Dashboard and savings report | |||
| Batch API | — | ||
| Alerts on ratio drops | — | ||
| Rate limit | 60 rpm | 3,000 rpm | Custom |
| Security | |||
| Zero data retention | On request | On request | Default |
| EU hosting | |||
| On-premise / VPC | — | — | |
| SOC 2 report, BAA | — | — | |
| Support | |||
| Channel | Community | Email, 1 business day | Dedicated engineer |
| SLA | — | 99.9 % | Custom |
Frequently asked questions
Everything a technical buyer asks in the first call, written down. Missing one? Ask us.
Request accessPrompt compression cuts the number of tokens in an LLM prompt or context window while preserving its meaning, so the same request uses fewer tokens and costs less.
Tokeen finds the most token-efficient way to represent your context: every bit of signal stays, the rest goes. It is fully deterministic and nothing is summarised or rewritten, so the text that reaches the model stays verbatim and in its original order.
Stop paying for filler. Compress every prompt on one fast, deterministic layer.
Tokeen is in private beta. Request access and see the savings on your own prompts within the week.