Context Compression Calculator

Find the session size at which summarising it starts to pay off.

What you pay

List rates at the time of writing, and every price field is editable.

Dollars for 1M fresh tokens.

Dollars for 1M tokens read from the cache.

Dollars for 1M tokens written back.

How your sessions spend tokens
Session type

Most turns resend the same files and instructions, so nearly every input token hits the cache.

Percent of the billed tokens that miss the cache.

Percent of the billed tokens that hit the cache.

Percent of the billed tokens the reply writes.

How far you compress

Percent of the 1M window the summary is capped at.

Compress once the session passes

618k tokens

Summarise and save

Compress once the session passes 618k tokens, where carrying a 10 percent summary costs $0.308, the same as keeping the session.

A summary may keep up to 16.2 percent of the window and still pay off.

Keep the session, full window
$0.497
Carry a 10% summary
$0.308
Difference
$0.190

The summary costs the same at every session size, because it is capped at 100k tokens.

Where the two costs cross

$0.00$0.200$0.400$0.600$0.8000250k500k750k1.0MSession size, in tokens of contextDollars for the next request618k tokens ties
  • Keep the whole session
  • Carry a 10% summary
The climbing line is the session you already hold, which gets dearer with every token you add. The flat line is a summary capped at 10 percent of the window, which costs the same however long the session runs. Where they cross is the size at which compressing starts to pay.
$0.00$0.200$0.400$0.600$0.8000%25%50%75%100%Summary cap, as a share of the windowDollars for the next requestyour cap, 10%16.2% ties
  • Keep the whole session
  • A summary, at the cap on the axis
  • Tie, which is a summary at the break-even cap
The flat line is the session you already hold, priced at a full window. The rising line is a summary, and it costs the full input price on every token it keeps, so a wider cap costs more. They meet at 16.2 percent, which is the widest cap that still pays off, and the dashed line marks the cap you entered. Past a cap of 26 percent the summary line leaves the top of the axis, since the answer there is already settled.

What the rates are made of

Kept rate, 1M session tokens
$0.497
Summary rate, 1M summary tokens
$3.08
Summary size
100k tokens

The shares your mix was read as, scaled so they add to 100 percent.

Miss
4.50%
Cache
95.00%
Output
0.50%

Questions about compressing a session

When does summarising a session save money?

Once the session passes the size named above, and not before that.

Why does session size decide it at all?

Because the summary is capped at a share of the window, so it costs the same at any session size while the session it replaces keeps getting dearer.

Why does the cap matter so much?

A wider cap means more summary tokens, and every one of them is new text billed at the full input price.

What does the token mix change?

It sets how much of your context is a cache hit, which sets how cheaply the kept session grows.

Why is the break-even cap so low on a coding session?

Because a coding session barely misses the cache, so the session you already hold is already the cheap option.

Where do the preset rates come from?

They are published list prices for one model from each major provider, and every price field is editable.