Zero-Shot Classification Reachability Leaderboard

Every model on the clef-evals Decision Model Leaderboard, ranked by the smallest card in the 40 card hardware directory on this site that can hold it. The ranking is recalculated as you set the workload: the weights in the format each checkpoint publishes, the KV cache at your context length and sequence count, an activation buffer, and a framework reserve. Models this site cannot size stay in the table with the reason.

Hardware

Every card in the hardware directory.

Workload

Tokens in each sequence. The KV cache grows with it.

Sequences served at once. The cache and the activation buffer both grow with it.

Published follows each checkpoint. A named format forces one format on every model.

The dtype the cache is held in. A narrower cache lets a smaller card hold a long context.

The most cards a configuration may use before the model is reported as not reachable.

Results

Ordered by the hardware a model needs, so rank 1 needs the smallest card and not the highest Decision Index.

The smallest card that holds the model first, then the models no card holds.

60 of 73 models sized for this workload. The rest are listed with the reason.

Zero-shot classification models ranked by the smallest card that holds each one
#ModelDecision IndexHardware neededSingle cardTokens/sWeightsDetails
1
Lumma-Fev-0.1B
lumma-fev-0.1b · full fine-tune
1.788 GB or larger40 of 401,573450.9 MiB
2
JPT-0.8B
jpt-0.8b · LoRA
19.228 GB or larger40 of 405711.20 GiB
3
Decision 1.0 Eos
decision-eos-0.8b · head / adapter
18.418 GB or larger40 of 405711.20 GiB
4
Kev 0.8B
kev-0.8b-raised · LoRA + head
14.608 GB or larger40 of 405711.20 GiB
5
Tev1-0.8B-experimental
tev1-0.8b · full fine-tune
12.858 GB or larger40 of 405711.20 GiB
6
Intern-Decision-0.8B
intern-decision-0.8b · full fine-tune
11.948 GB or larger40 of 405711.20 GiB
7
MoJev
mojev · head / adapter
11.698 GB or larger40 of 405711.20 GiB
8
LFM2.5-350M-RLCD
lfm350-sdpa · full fine-tune
1.388 GB or larger40 of 40547976.0 MiB
9
Bosun v3.1 0.6B
bosun-v3.1-0.6b · LoRA + head
14.328 GB or larger40 of 403601.11 GiB
10
Lumma-Fev-0.6B
lumma-fev-0.6b · LoRA
2.978 GB or larger40 of 403081.49 GiB
11
Decider 2B [FP8]
decider-2b-fp8-http · full fine-tune
28.978 GB or larger40 of 402043.10 GiB
12
this-that 1.2
this-that-1.2 · full fine-tune
28.148 GB or larger40 of 402043.10 GiB
13
Decision 1.0 Sol
decision-sol-0665a411 · head / adapter
25.328 GB or larger40 of 402043.10 GiB
14
Intern-Decision-2B
intern-decision-2b · full fine-tune
19.388 GB or larger40 of 402043.10 GiB
15
Qwen-2.5-1B-RLCD
harsha · inference technique
3.788 GB or larger40 of 401782.88 GiB
16
Bosun v3.1 1.7B
bosun-v3.1-1.7b · LoRA + head
20.108 GB or larger40 of 401483.20 GiB
17
LFM2.5-2.6B-RLCD
lfm2600 · full fine-tune
6.768 GB or larger40 of 401004.77 GiB
18
JPT-4B
jpt-4b · LoRA
43.0410 GB or larger33 of 40956.97 GiB
19
Jet v6.2
jet-v6.2 · LoRA
42.6010 GB or larger33 of 40956.97 GiB
20
Hopper (G) 1.2
hopper-g-1.2 · LoRA
40.7710 GB or larger33 of 40956.97 GiB
21
Decider 4B
decider-4b · full fine-tune
40.7010 GB or larger33 of 40956.97 GiB
22
JevK5
jevk5 · LoRA
38.8110 GB or larger33 of 40956.97 GiB
23
lev
lev · LoRA + head
38.5410 GB or larger33 of 40956.97 GiB
24
Intern-Decision-4B
intern-decision-4b · full fine-tune
37.8110 GB or larger33 of 40956.97 GiB
25
NeoHorse-Jev-4B
neohorse-jev-4b · head / adapter
36.7510 GB or larger33 of 40956.97 GiB
26
Kev 4B
kev-4b-raised · LoRA + head
34.6410 GB or larger33 of 40956.97 GiB
27
Decision 1.0 Nox
decision-nox-0bb83350 · head / adapter
34.3610 GB or larger33 of 40956.97 GiB
28
Jobe
jobe · inference technique
32.3510 GB or larger33 of 40956.97 GiB
29
open-jev (pngwn)
pngwn-space-uncapped-full · LoRA
29.9110 GB or larger33 of 40956.97 GiB
30
Tev1-4B-experimental
tev1-4b · full fine-tune
29.2410 GB or larger33 of 40956.97 GiB
31
Metask-Jev-4B
metask-jev-4b · LoRA
26.8910 GB or larger33 of 40956.97 GiB
32
SemIf
semif · inference technique
25.9410 GB or larger33 of 40956.97 GiB
33
Winnow-E4B [Q8_0]
winnow-e4b-q8 · LoRA
39.8912 GB or larger32 of 40938.43 GiB
34
openvons
openvons · inference technique
28.4212 GB or larger32 of 40937.49 GiB
35
mini-jev
mini-jev · inference technique
20.9812 GB or larger32 of 40937.49 GiB
36
Clef-flash
clef-flash · Cloudflare · Routing head + LoRA
57.0720 GB or larger12 of 405115.29 GiB
37
JPT-9B
jpt-9b · LoRA
46.8920 GB or larger12 of 405115.29 GiB
38
Decision 1.0 Lux
decision-lux-9b · head / adapter
43.4920 GB or larger12 of 405115.29 GiB
39
Bespoke Nimble 9B v2
nimble-v2 · LoRA
39.5720 GB or larger12 of 405115.29 GiB
40
Kev 9B
kev-9b-raised · LoRA + head
38.4820 GB or larger12 of 405115.29 GiB
41
CLM-v0.1-8B
clm-v0.1-8b · full fine-tune
7.4020 GB or larger12 of 404415.26 GiB
42
Winnow-12B [Q8_0]
winnow-12b-q8 · LoRA
50.0232 GB or larger7 of 406621.91 GiB
43
Jev-Omni
jev-omni · LoRA + head
40.5332 GB or larger7 of 406621.91 GiB
44
Surogate Rune 26B-A4B v3 [bf16]
rune-26b-a4b-v3 · full fine-tune
57.4480 GB or larger6 of 401,20645.87 GiB
45
JoshuaSP diffusiongemma (open-jev)
joshua-diffusion-full · inference technique
49.4780 GB or larger6 of 401,14145.87 GiB
46
djev
djev · inference technique
40.2880 GB or larger6 of 401,14145.87 GiB
47
razorback16 openjev diffusiongemma (NVFP4, vLLM) [one read, NVFP4]
razorback-one-read-3e296f08 · inference technique
37.2580 GB or larger6 of 401,14145.87 GiB
48
mmastrac diffusiongemma (vLLM PR 57250)
vllm-pr57250 · inference technique
32.2480 GB or larger6 of 401,14145.87 GiB
49
Decider 35B-A3B [NVFP4]
decider-35b-nvfp4 · full fine-tune
47.1180 GB or larger6 of 4068363.57 GiB
50
Xor
xor · full fine-tune
41.4880 GB or larger6 of 4068363.57 GiB
51
Clef
clef · Cloudflare · Routing head + LoRA
61.2180 GB or larger6 of 406145.36 GiB
52
AutoJev-27B
autojev-27b · full fine-tune
56.4080 GB or larger6 of 406145.36 GiB
53
simple-jev · Qwen3.8-27B (featherless)
simple-jev-qwen3.8-27b · inference technique
55.7480 GB or larger6 of 406145.36 GiB
54
Jebadiah 27B
jebadiah-27b · LoRA
54.6780 GB or larger6 of 406145.36 GiB
55
Eikos-27B [FP8]
eikos-27b-fp8 · LoRA
53.1380 GB or larger6 of 406145.36 GiB
56
reflex 27B [FP8 · wide choice]
reflex-27b-v2 · inference technique
52.1680 GB or larger6 of 406145.36 GiB
57
Decider chat · Qwen3.6-27B
decider-chat-qwen3.6-27b · inference technique
51.3580 GB or larger6 of 406145.36 GiB
58
Jevfire
jevfire-uncapped · inference technique
49.3780 GB or larger6 of 406145.36 GiB
59
Solomon v1.1
solomon-v11-bf16-encoding · head / adapter
36.4380 GB or larger6 of 406145.36 GiB
60
Decider chat · Gemma-4-31B
decider-chat-gemma4-31b · inference technique
57.3380 GB or larger6 of 404656.15 GiB
61
Jev
jev · TypeSafe · Hosted API (closed)
Closed hosted API
57.91Not sized---
62
GLiNER2.5-Decide
gliner-decide · full fine-tune
Not a decoder transformer (extractor)
11.21Not sized---
63
Lavoir
lavoir · full fine-tune
No cache shape in the config
8.69Not sized---
64
jeff
jeff-uncapped · inference technique
No published config
8.04Not sized---
65
GLiNER 2.5 base
gliner-base · full fine-tune
Not a decoder transformer (extractor)
6.76Not sized---
66
Decision 1.0 Kai
decision-kai-7185f514 · head / adapter
No cache shape in the config
6.52Not sized---
67
Laya
laya-1e28ac20 · full fine-tune
Not a decoder transformer (laya)
6.04Not sized---
68
Julia 1
julia-1 · full fine-tune
No cache shape in the config
5.54Not sized---
69
system-one-gemma
akash-gemma · LoRA + head
No published config
5.07Not sized---
70
Decision 1.0 Lex
decision-lex · head / adapter
No cache shape in the config
4.54Not sized---
71
GLiNER 2.5 multilingual
gliner-multi · full fine-tune
Not a decoder transformer (extractor)
4.26Not sized---
72
GLiNER 2.5 small
gliner-small · full fine-tune
Not a decoder transformer (extractor)
3.82Not sized---
73
Verdict
verdict-8af2496e · full fine-tune
Not a decoder transformer (GLiClass)
1.87Not sized---

Tokens each second is the total across the concurrent sequences you set, for the configuration the ranking recommended.

Where the numbers come from

The model list, the Decision Index, and the measured latency come from the clef-evals Decision Model Leaderboard, which is Decision Index 0.2.1 and was generated on 2026-09-28 on 1 x NVIDIA RTX PRO 6000. The checkpoint configs come from each model base checkpoint on HuggingFace, and the hardware figures come from the same 40 card directory the other calculators here use. This page was last built on 2026-10-06, from 73 models.

Questions about reachability

What does reachability mean on this page?

It is the smallest card in the hardware directory on this site that can hold the whole run, the number of cards here that can hold it, and the decode rate that configuration reaches. The run is the weights, the KV cache at the chosen context length and sequence count, an activation buffer, and a framework reserve.

Where do the models and the scores come from?

The model list, the Decision Index, and the measured latency come from the clef-evals Decision Model Leaderboard. The leaderboard is scored on one RTX PRO 6000, so its latency figure is not the decode rate shown here.

Why is the decode rate different from the measured latency?

The latency on the leaderboard is one end to end call, which includes prefill, the decision head, and the client round trip. The rate here is a bandwidth roofline for decode alone, which is the highest rate the card can reach once the prompt is read.

Why can a model not be sized at all?

Either the checkpoint publishes no config this site can read, or the config describes an architecture that is not a decoder transformer. A closed hosted API publishes no checkpoint. An encoder only classifier publishes one, but it has no KV cache and no decode loop, so the sizing model does not apply to it.

How accurate is the weight footprint?

It is config arithmetic rather than a measurement. The calculator counts the attention blocks, the dense feed forward, the expert bank, the router, and the embedding tables from the published config. It does not count a multimodal vision tower, a multi token prediction block, or the savings of layers that share weights, so a model with any of those can be a few percent off in either direction. The breakdown shows the parameter count the config produced beside the count the leaderboard reports.

Why does forcing a weight format change which cards fit?

The bytes for each weight follow the format. A checkpoint published in BF16 costs two bytes for each weight, while an FP8 or a four bit copy costs one byte or half a byte. Forcing a narrower format stores the same parameter count in fewer bytes, so smaller cards fit.

Why does the context length change the answer?

The KV cache grows with the context length and the sequence count, and it has to be resident beside the weights. A model that fits on a small card at a short context can need a larger card at a long one. The activation buffer follows the sequence count instead, because it holds one token of intermediates for each sequence at once.

What does the card count mean when one card is not enough?

The weights and the KV cache divide across the tensor parallel ranks, while the activation buffer and the framework reserve stay in full on every card. Two cards therefore do not halve the footprint. The page reports the fewest cards that hold the run, and then the smallest card that reaches that count.