Pricing

Pay for what you compress. Keep what you save.

One meter: tokens that go through the compressor. No seats, no minimums, no percentage of savings. Tokeen is in private beta — these are the plans you land on once you are in.

Free
$0forever

Your own prompts, your own numbers. No card, no NDA — access by request while in beta.

  • 1M tokens compressed a month
  • All models, all providers
  • Python, Node and AI SDK wrappers
  • Playground and dashboard
  • Community support
Request access
Enterprise
Customannual

Run it inside your network, with the paperwork a security review asks for.

  • On-premise or your VPC
  • Zero data retention, account-wide
  • SOC 2 report and signed BAA
  • Custom aggressiveness per route
  • Dedicated engineer, shared Slack
Talk to sales

How the meter works

You are billed on tokens sent to the compressor, at a fraction of any model's input price. Whatever it removes, you stop paying your provider for. Here is a month for a product sending 400M input tokens to a model priced at $2.50 per million.

Worked example400M tokens · month
Provider bill before Tokeen400M × $2.50 / M$1,000
Tokens removed at default aggressiveness−20 % on uncached prefixes−80M
Provider bill after320M × $2.50 / M$800
Tokeen Pro400M × $0.25 / M$100
Net saving$100 / month

Compare plans

FeatureFreeProEnterprise
Compression
Tokens a month1MUnlimitedUnlimited
Prompt, context and conversation modes
Tool-result mode for agents
Aggressiveness controlGlobalPer requestPer route
Multilingual model
Platform
Dashboard and savings report
Batch API
Alerts on ratio drops
Rate limit60 rpm3,000 rpmCustom
Security
Zero data retentionOn requestOn requestDefault
EU hosting
On-premise / VPC
SOC 2 report, BAA
Support
ChannelCommunityEmail, 1 business dayDedicated engineer
SLA99.9 %Custom

Frequently asked questions

Everything a technical buyer asks in the first call, written down. Missing one? Ask us.

Request access
  • Prompt compression cuts the number of tokens in an LLM prompt or context window while preserving its meaning, so the same request uses fewer tokens and costs less.

    Tokeen finds the most token-efficient way to represent your context: every bit of signal stays, the rest goes. It is fully deterministic and nothing is summarised or rewritten, so the text that reaches the model stays verbatim and in its original order.

Stop paying for filler. Compress every prompt on one fast, deterministic layer.

Tokeen is in private beta. Request access and see the savings on your own prompts within the week.