KV Cache Calculator

Size the cache for a model, a context length, and a sequence count. The config is read from HuggingFace or ModelScope.

Provider

Reading the config

Tokens per sequence.

Concurrent requests in the batch.

Only needed for gated or private repos. Sent to HuggingFace and nowhere else.

Reading the model config