Zero-Shot Classification Reachability Leaderboard
Every model on the clef-evals Decision Model Leaderboard, ranked by the smallest card in the 40 card hardware directory on this site that can hold it. The ranking is recalculated as you set the workload: the weights in the format each checkpoint publishes, the KV cache at your context length and sequence count, an activation buffer, and a framework reserve. Models this site cannot size stay in the table with the reason.
Results
Ordered by the hardware a model needs, so rank 1 needs the smallest card and not the highest Decision Index.
The smallest card that holds the model first, then the models no card holds.
60 of 73 models sized for this workload. The rest are listed with the reason.
| # | Model | Decision Index | Hardware needed | Single card | Tokens/s | Weights | Details |
|---|---|---|---|---|---|---|---|
| 1 | Lumma-Fev-0.1B lumma-fev-0.1b · full fine-tune | 1.78 | 8 GB or larger | 40 of 40 | 1,573 | 450.9 MiB | |
| 2 | JPT-0.8B jpt-0.8b · LoRA | 19.22 | 8 GB or larger | 40 of 40 | 571 | 1.20 GiB | |
| 3 | Decision 1.0 Eos decision-eos-0.8b · head / adapter | 18.41 | 8 GB or larger | 40 of 40 | 571 | 1.20 GiB | |
| 4 | Kev 0.8B kev-0.8b-raised · LoRA + head | 14.60 | 8 GB or larger | 40 of 40 | 571 | 1.20 GiB | |
| 5 | Tev1-0.8B-experimental tev1-0.8b · full fine-tune | 12.85 | 8 GB or larger | 40 of 40 | 571 | 1.20 GiB | |
| 6 | Intern-Decision-0.8B intern-decision-0.8b · full fine-tune | 11.94 | 8 GB or larger | 40 of 40 | 571 | 1.20 GiB | |
| 7 | MoJev mojev · head / adapter | 11.69 | 8 GB or larger | 40 of 40 | 571 | 1.20 GiB | |
| 8 | LFM2.5-350M-RLCD lfm350-sdpa · full fine-tune | 1.38 | 8 GB or larger | 40 of 40 | 547 | 976.0 MiB | |
| 9 | Bosun v3.1 0.6B bosun-v3.1-0.6b · LoRA + head | 14.32 | 8 GB or larger | 40 of 40 | 360 | 1.11 GiB | |
| 10 | Lumma-Fev-0.6B lumma-fev-0.6b · LoRA | 2.97 | 8 GB or larger | 40 of 40 | 308 | 1.49 GiB | |
| 11 | Decider 2B [FP8] decider-2b-fp8-http · full fine-tune | 28.97 | 8 GB or larger | 40 of 40 | 204 | 3.10 GiB | |
| 12 | this-that 1.2 this-that-1.2 · full fine-tune | 28.14 | 8 GB or larger | 40 of 40 | 204 | 3.10 GiB | |
| 13 | Decision 1.0 Sol decision-sol-0665a411 · head / adapter | 25.32 | 8 GB or larger | 40 of 40 | 204 | 3.10 GiB | |
| 14 | Intern-Decision-2B intern-decision-2b · full fine-tune | 19.38 | 8 GB or larger | 40 of 40 | 204 | 3.10 GiB | |
| 15 | Qwen-2.5-1B-RLCD harsha · inference technique | 3.78 | 8 GB or larger | 40 of 40 | 178 | 2.88 GiB | |
| 16 | Bosun v3.1 1.7B bosun-v3.1-1.7b · LoRA + head | 20.10 | 8 GB or larger | 40 of 40 | 148 | 3.20 GiB | |
| 17 | LFM2.5-2.6B-RLCD lfm2600 · full fine-tune | 6.76 | 8 GB or larger | 40 of 40 | 100 | 4.77 GiB | |
| 18 | JPT-4B jpt-4b · LoRA | 43.04 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 19 | Jet v6.2 jet-v6.2 · LoRA | 42.60 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 20 | Hopper (G) 1.2 hopper-g-1.2 · LoRA | 40.77 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 21 | Decider 4B decider-4b · full fine-tune | 40.70 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 22 | JevK5 jevk5 · LoRA | 38.81 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 23 | lev lev · LoRA + head | 38.54 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 24 | Intern-Decision-4B intern-decision-4b · full fine-tune | 37.81 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 25 | NeoHorse-Jev-4B neohorse-jev-4b · head / adapter | 36.75 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 26 | Kev 4B kev-4b-raised · LoRA + head | 34.64 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 27 | Decision 1.0 Nox decision-nox-0bb83350 · head / adapter | 34.36 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 28 | Jobe jobe · inference technique | 32.35 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 29 | open-jev (pngwn) pngwn-space-uncapped-full · LoRA | 29.91 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 30 | Tev1-4B-experimental tev1-4b · full fine-tune | 29.24 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 31 | Metask-Jev-4B metask-jev-4b · LoRA | 26.89 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 32 | SemIf semif · inference technique | 25.94 | 10 GB or larger | 33 of 40 | 95 | 6.97 GiB | |
| 33 | Winnow-E4B [Q8_0] winnow-e4b-q8 · LoRA | 39.89 | 12 GB or larger | 32 of 40 | 93 | 8.43 GiB | |
| 34 | openvons openvons · inference technique | 28.42 | 12 GB or larger | 32 of 40 | 93 | 7.49 GiB | |
| 35 | mini-jev mini-jev · inference technique | 20.98 | 12 GB or larger | 32 of 40 | 93 | 7.49 GiB | |
| 36 | Clef-flash clef-flash · Cloudflare · Routing head + LoRA | 57.07 | 20 GB or larger | 12 of 40 | 51 | 15.29 GiB | |
| 37 | JPT-9B jpt-9b · LoRA | 46.89 | 20 GB or larger | 12 of 40 | 51 | 15.29 GiB | |
| 38 | Decision 1.0 Lux decision-lux-9b · head / adapter | 43.49 | 20 GB or larger | 12 of 40 | 51 | 15.29 GiB | |
| 39 | Bespoke Nimble 9B v2 nimble-v2 · LoRA | 39.57 | 20 GB or larger | 12 of 40 | 51 | 15.29 GiB | |
| 40 | Kev 9B kev-9b-raised · LoRA + head | 38.48 | 20 GB or larger | 12 of 40 | 51 | 15.29 GiB | |
| 41 | CLM-v0.1-8B clm-v0.1-8b · full fine-tune | 7.40 | 20 GB or larger | 12 of 40 | 44 | 15.26 GiB | |
| 42 | Winnow-12B [Q8_0] winnow-12b-q8 · LoRA | 50.02 | 32 GB or larger | 7 of 40 | 66 | 21.91 GiB | |
| 43 | Jev-Omni jev-omni · LoRA + head | 40.53 | 32 GB or larger | 7 of 40 | 66 | 21.91 GiB | |
| 44 | Surogate Rune 26B-A4B v3 [bf16] rune-26b-a4b-v3 · full fine-tune | 57.44 | 80 GB or larger | 6 of 40 | 1,206 | 45.87 GiB | |
| 45 | JoshuaSP diffusiongemma (open-jev) joshua-diffusion-full · inference technique | 49.47 | 80 GB or larger | 6 of 40 | 1,141 | 45.87 GiB | |
| 46 | djev djev · inference technique | 40.28 | 80 GB or larger | 6 of 40 | 1,141 | 45.87 GiB | |
| 47 | razorback16 openjev diffusiongemma (NVFP4, vLLM) [one read, NVFP4] razorback-one-read-3e296f08 · inference technique | 37.25 | 80 GB or larger | 6 of 40 | 1,141 | 45.87 GiB | |
| 48 | mmastrac diffusiongemma (vLLM PR 57250) vllm-pr57250 · inference technique | 32.24 | 80 GB or larger | 6 of 40 | 1,141 | 45.87 GiB | |
| 49 | Decider 35B-A3B [NVFP4] decider-35b-nvfp4 · full fine-tune | 47.11 | 80 GB or larger | 6 of 40 | 683 | 63.57 GiB | |
| 50 | Xor xor · full fine-tune | 41.48 | 80 GB or larger | 6 of 40 | 683 | 63.57 GiB | |
| 51 | Clef clef · Cloudflare · Routing head + LoRA | 61.21 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 52 | AutoJev-27B autojev-27b · full fine-tune | 56.40 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 53 | simple-jev · Qwen3.8-27B (featherless) simple-jev-qwen3.8-27b · inference technique | 55.74 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 54 | Jebadiah 27B jebadiah-27b · LoRA | 54.67 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 55 | Eikos-27B [FP8] eikos-27b-fp8 · LoRA | 53.13 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 56 | reflex 27B [FP8 · wide choice] reflex-27b-v2 · inference technique | 52.16 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 57 | Decider chat · Qwen3.6-27B decider-chat-qwen3.6-27b · inference technique | 51.35 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 58 | Jevfire jevfire-uncapped · inference technique | 49.37 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 59 | Solomon v1.1 solomon-v11-bf16-encoding · head / adapter | 36.43 | 80 GB or larger | 6 of 40 | 61 | 45.36 GiB | |
| 60 | Decider chat · Gemma-4-31B decider-chat-gemma4-31b · inference technique | 57.33 | 80 GB or larger | 6 of 40 | 46 | 56.15 GiB | |
| 61 | Jev jev · TypeSafe · Hosted API (closed) Closed hosted API | 57.91 | Not sized | - | - | - | |
| 62 | GLiNER2.5-Decide gliner-decide · full fine-tune Not a decoder transformer (extractor) | 11.21 | Not sized | - | - | - | |
| 63 | Lavoir lavoir · full fine-tune No cache shape in the config | 8.69 | Not sized | - | - | - | |
| 64 | jeff jeff-uncapped · inference technique No published config | 8.04 | Not sized | - | - | - | |
| 65 | GLiNER 2.5 base gliner-base · full fine-tune Not a decoder transformer (extractor) | 6.76 | Not sized | - | - | - | |
| 66 | Decision 1.0 Kai decision-kai-7185f514 · head / adapter No cache shape in the config | 6.52 | Not sized | - | - | - | |
| 67 | Laya laya-1e28ac20 · full fine-tune Not a decoder transformer (laya) | 6.04 | Not sized | - | - | - | |
| 68 | Julia 1 julia-1 · full fine-tune No cache shape in the config | 5.54 | Not sized | - | - | - | |
| 69 | system-one-gemma akash-gemma · LoRA + head No published config | 5.07 | Not sized | - | - | - | |
| 70 | Decision 1.0 Lex decision-lex · head / adapter No cache shape in the config | 4.54 | Not sized | - | - | - | |
| 71 | GLiNER 2.5 multilingual gliner-multi · full fine-tune Not a decoder transformer (extractor) | 4.26 | Not sized | - | - | - | |
| 72 | GLiNER 2.5 small gliner-small · full fine-tune Not a decoder transformer (extractor) | 3.82 | Not sized | - | - | - | |
| 73 | Verdict verdict-8af2496e · full fine-tune Not a decoder transformer (GLiClass) | 1.87 | Not sized | - | - | - |
Tokens each second is the total across the concurrent sequences you set, for the configuration the ranking recommended.
Where the numbers come from
The model list, the Decision Index, and the measured latency come from the clef-evals Decision Model Leaderboard, which is Decision Index 0.2.1 and was generated on 2026-09-28 on 1 x NVIDIA RTX PRO 6000. The checkpoint configs come from each model base checkpoint on HuggingFace, and the hardware figures come from the same 40 card directory the other calculators here use. This page was last built on 2026-10-06, from 73 models.
Questions about reachability
What does reachability mean on this page?
It is the smallest card in the hardware directory on this site that can hold the whole run, the number of cards here that can hold it, and the decode rate that configuration reaches. The run is the weights, the KV cache at the chosen context length and sequence count, an activation buffer, and a framework reserve.
Where do the models and the scores come from?
The model list, the Decision Index, and the measured latency come from the clef-evals Decision Model Leaderboard. The leaderboard is scored on one RTX PRO 6000, so its latency figure is not the decode rate shown here.
Why is the decode rate different from the measured latency?
The latency on the leaderboard is one end to end call, which includes prefill, the decision head, and the client round trip. The rate here is a bandwidth roofline for decode alone, which is the highest rate the card can reach once the prompt is read.
Why can a model not be sized at all?
Either the checkpoint publishes no config this site can read, or the config describes an architecture that is not a decoder transformer. A closed hosted API publishes no checkpoint. An encoder only classifier publishes one, but it has no KV cache and no decode loop, so the sizing model does not apply to it.
How accurate is the weight footprint?
It is config arithmetic rather than a measurement. The calculator counts the attention blocks, the dense feed forward, the expert bank, the router, and the embedding tables from the published config. It does not count a multimodal vision tower, a multi token prediction block, or the savings of layers that share weights, so a model with any of those can be a few percent off in either direction. The breakdown shows the parameter count the config produced beside the count the leaderboard reports.
Why does forcing a weight format change which cards fit?
The bytes for each weight follow the format. A checkpoint published in BF16 costs two bytes for each weight, while an FP8 or a four bit copy costs one byte or half a byte. Forcing a narrower format stores the same parameter count in fewer bytes, so smaller cards fit.
Why does the context length change the answer?
The KV cache grows with the context length and the sequence count, and it has to be resident beside the weights. A model that fits on a small card at a short context can need a larger card at a long one. The activation buffer follows the sequence count instead, because it holds one token of intermediates for each sequence at once.
What does the card count mean when one card is not enough?
The weights and the KV cache divide across the tensor parallel ranks, while the activation buffer and the framework reserve stay in full on every card. Two cards therefore do not halve the footprint. The page reports the fewest cards that hold the run, and then the smallest card that reaches that count.