KV Cache Calculator

Size the cache for a model, a context length, and a sequence count. The config can come from HuggingFace, from ModelScope, from a config.json you paste, or from values you enter yourself.

Model source

Read the config from HuggingFace or ModelScope.

Provider

Reading the config

Only needed for gated or private repos. Sent to HuggingFace and nowhere else.

Tokens per sequence.

Concurrent requests in the batch.

Reading the model config