REAP Cost Calculator

This REAP duration calculator estimates how long a Router-weighted Expert Activation Pruning run takes on your hardware, whether one expert block fits in your VRAM, and how much smaller the pruned model gets. The shape is read from a HuggingFace or ModelScope config, a config.json you paste, or numbers you type.

Model source

Read the config from HuggingFace or ModelScope.

Provider

Reading the config

Only needed for gated or private repos. Sent to HuggingFace and nowhere else.

Calibration

The recipe used in the vLLM llm-compressor example. About 1 million tokens.

Published checkpoints use 25, 30, 40, and 50 percent.

Hardware

32 GB of memory, 1,792 GB/s. 32 GB GDDR7. 3352 AI TOPS.

supportedReads at the precision most checkpoints ship in, and runs on every card listed.

The layer-wise observer streams one block at a time, so a slow disk can become the limit.

Assumptions

A share of dense peak, between 0 and 1. A third is realistic.

Framework, observer, and dataloader cost over the raw figure.

Samples in flight at once. Drives the activation buffer.

Off by default. Pruning alone shrinks the model in memory but leaves the arithmetic per token unchanged.

Reading the model config

Questions about REAP

How long would it take to REAP a model?

It is set almost entirely by the calibration token count, which is the number of samples times the sequence length. A quick pass of 512 samples at 2048 tokens is minutes on a modern card. The recipe published with the paper, 24576 samples at 16384 tokens, is about 403 million tokens and runs for tens of hours on a single GPU. Enter a model and a card above to get the figure for your case.

Can I prune a mixture of experts model on one GPU?

Yes. The layer-wise observer added to REAP keeps only one decoder block resident at a time, so the memory requirement is set by the largest single block rather than by the whole model. That is what makes a trillion parameter model tractable on one card. The question is whether one block, plus its activation buffer, fits in your VRAM. The verdict above answers it for the model and card you selected.

How many calibration samples does REAP need?

The paper calibrated on 24576 samples at 16384 tokens. The vLLM llm-compressor example uses 512 samples at 2048 tokens and still ranks experts well. Fewer samples cost proportionally less time but give a noisier saliency ranking, and a dataset that misses a topic can mark the experts for that topic as unimportant.

Does REAP make inference faster?

Not on its own. Pruning removes experts from memory, so the model gets smaller and can fit on less hardware. A token still routes to the same number of experts unless the router top-k is reduced as well, so the arithmetic per token is unchanged. Reducing the top-k is what turns a smaller model into a faster one.

What pruning ratio should I use?

Published checkpoints use 25, 30, 40, and 50 percent. How much quality is lost depends on how redundant the original experts are. Qwen3-30B-A3B recovered 99.8 percent of its baseline score at 50 percent pruning, while Moonlight-16B-A3B recovered only 17 percent at the same ratio. Calibrate once and the saliency scores can be reused to produce any ratio.

Does REAP work on dense models?

No. REAP removes routed experts, so a model with no expert bank has nothing to prune. The calculator says so directly when you enter a dense model.