Compare models by hardware and workload¶
Every retained model configuration measured on this hardware, in one place: what it ran on, how fast it was, whether reasoning was on or off, and where the working recipe lives.
If you already know your hardware and workload, start with measured recipe results. They show a compact strength and limitation summary beside metrics read directly from native artifacts. Return here for the broader historical and cross-model record.
Choose the comparison that answers your question¶
Interactive response
Compare TTFT, E2E, and TPOT/ITL at the same prompt depth and concurrency. Start with the measured recipe results.
Sustained concurrency
Compare aggregate output and p95/p99 request latency only within one matched workload. Start with the Qwen PRO campaign.
Long context
Check actual prompt tokens, output reserve, and measured concurrency. Use the RTX PRO 6000 or RTX 5090 view.
Tools, agents, and media
Speed is not enough. Use the model dossiers for structured output, tool use, agentic/SWE, image, OCR, video, and client evidence.
How to read this table¶
Rates are not interchangeable. Three different instruments appear in the Output rate
column, and mixing them produces nonsense. Each cell says which one it is:
| Tag | Instrument | What it means |
|---|---|---|
long-gen |
Controlled long-generation decode | Sustained decode rate on a long answer. This is the number to compare across models. |
agg |
Short-output aggregate capacity | Total output tokens ÷ wall time on a short-answer batch. Dominated by scheduling and prefill; not a decode rate. |
c128 |
Continuous-batching aggregate | Throughput under 128 concurrent requests. Answers a serving-density question, not a single-user one. |
A row with a small agg number is usually not a slow model — it is a short workload. The GPT-OSS
Puzzle rows below produced only 20 to 86 output tokens, so their agg figures are explicitly
unusable as decode rates.
TTFT only means something with its depth and concurrency. Cells name them wherever a row mixes several, and a column names them once in its header where the whole column shares one shape — note that a row's TTFT depth is often not its served context, because latency was usually probed at 8K on a much larger window. A bare cell was measured at c1. A near-ceiling TTFT (tens of seconds) is prefill-bound and is a prefill-latency proxy, not a latency regression.
Reasoning is part of the recipe, not a preference. Some configurations are only qualified with thinking disabled; forcing it on changes the result and, in several cases, exhausts the completion budget and returns no visible answer.
Engine spread is wide — NGC vLLM 0.19 through 0.25.1, pinned nightlies, a custom Anvil fork, and llama.cpp/q36. Rows measured on different engines are not clean comparisons even at the same context and concurrency.
A row is not always one run. Where a model was measured across several campaigns, the cells here may come from different dates and engine windows — the Qwen3.6 27B MTP row pairs a 2026-07-12 TTFT with a 2026-07-10/11 decode A/B, and the Gemma 12B QAT row pairs the 07-16 bakeoff capacity with the follow-up's equal-length diagnostic (that follow-up recorded 20/68 for the same capacity point, so the runs are close but not identical). Follow the dossier and dated finding before treating a row as a single experiment. Full rules: Methodology and evidence.
Published decision snapshot as of 2026-09-14¶
This section synthesizes the latest dated public benchmark decisions; it does not report active deployment, routing, placement, or availability. In the 2026-09-14 decision, GLM Flash EXL3 r7 no-spec is the selected text-only reference at 327,680 total tokens with a 65,536-token output reserve and C4. It retained 90/100 one-pass MMLU-Pro, agentic 30/30, SWE 4/5, strict120 120/120, and 9/9 post-promotion context. The former SGLang W4A16/NVFP4 profile and 524K EXL3 K3 plus corrected DFlash2 K5 profile remain historical evidence. RadixArk Qwen3.8 Flash Next NVFP4 remains the video-capable rollback. The DeepSeek Infernal Invocation profiles and the Qwen3.8 27B single-service/two-service split remain retained historical recipes. Separately, on one RTX 5090, NInfer NVFP4 MTP3 is the preferred measured direct Qwen3.8 27B text/tools performance challenger; it was not promoted and does not replace the broader GGUF capability evidence.
The September 4–5 Qwen3.8 27B PRO 6000 topology and runtime results are
available in the measured recipe results and the
dated campaign.
They remain no-promotion and do not replace the published current or rollback
historical profiles below. The September 18 vision follow-up records the selected r9 eight-image profile and r8 one-image rollback.
| Model / config | Status | Recipe / config | Quant · KV | 4K median TTFT / E2E | Decode | Multimodal acceptance | Evidence |
|---|---|---|---|---|---|---|---|
| GLM Flash EXL3 r7 no-spec | September 14 text-only reference, C4 | Dated finding | EXL3 4-bpw · FP8 DS-MLA KV · no speculation | not comparable; see C4 finding | 38.908 tok/s median in one matched C4 population | text-only; Pi/Hermes/OpenClaw tool checks pass | finding and evidence |
| Qwen3.8 27B FP8 | retained September 14 rollback | declared original FP8 path | FP8 weights · KV not specified here | not exercised after promotion | not comparable | text-only rollback | selected-current finding |
| GLM-5.3-Flash ormandj W4A16/NVFP4, SGLang | historical 2026-09-02 text/tools/image/OCR; exclusive TP=2, C1 | Pinned recipe | ModelOpt W4A16/NVFP4 · FP8 KV · adaptive EAGLE | 0.177 / 0.500 s | 112.07 tok/s | image/OCR 12/12 plus routed and real-client pass; video disabled | promotion |
| GLM-5.3-Flash EXL3 K3 plus corrected DFlash2 K5 | historical 2026-08-31 exact rollback; exclusive TP=2 | K5 + control | EXL3 K3 · FP8 DS-MLA target KV · BF16 draft KV · fixed K5 | 0.974 / 1.269 s | 83.08 tok/s | image/OCR pass; up to 16 images; no video | 524K xgrammar qualification |
| Qwen3.8 Flash Next RadixArk NVFP4 | historical 2026-08-26 text/image/OCR/video rollback; exclusive TP=2 | MTP3 + control | ModelOpt NVFP4 · BF16 KV · MTP3 | 0.141 / 0.340 s | 155.9 tok/s | direct 30/30; live repeats 57/60 strict; four images or one video | vision promotion |
| GLM-5.3-Flash TR3/EXL3 4 bpw, vision fixed K5 | historical GLM rollback evidence; exclusive TP=2 | Recipe source | EXL3 4 bpw · NVFP4 DS-MLA KV · fixed MTP5 | 1.052 / 1.568 s | 72.8 tok/s | image/OCR pass; video disabled | qualification |
| DeepSeek V4 Flash 0731, Infernal Invocation r18 DSpark K5 | former text Primary; retained TP=2 evidence | K5 + control | B12X W4A8 · FP8 compressed MLA KV | aggregate TTFT/E2E not published | 142.1 tok/s | text-only contract; routed and real-client acceptance passed | promotion |
| Qwen3.8 27B official FP8, SGLang TP=1 MTP=3 | former Primary/general-vision/OCR/video service | MTP3 + control | Official FP8 · FP8 E4M3 KV | 0.577 / 0.962 s | 111.4 tok/s | CPU transport; direct 30/30, live admitted 28/28; two images, one video | video expansion |
| Qwen3.8 27B official FP8, vLLM TP=1 MTP=3 | split rollback, text | Recipe source | Official FP8 · FP8 KV | 0.834 / 1.295 s | 93.6 tok/s | text only; routed tools 20/20 | historical promotion |
| Qwen3.8 27B official BF16, vLLM TP=1 MTP=3 | split rollback, vision/OCR | Recipe source | BF16 · FP8 KV | 0.884 / 1.584 s | 62.0 tok/s | historical routed 30/30; 32-image request 1/1 | historical promotion |
| Qwen3.8 27B Inferact NVFP4, SGLang TP=1 MTP=3 | challenger, no-promotion |
Recipe source | ModelOpt NVFP4 · FP8 KV | 0.448 / 0.914 s | 98.1 tok/s | CPU transport: image/OCR pass; broader media open | qualification |
The BF16 32-image result is one request at concurrency one. It is not a claim of 32 concurrent requests.
GLM scheduler concurrency versus KV capacity¶
Three values must remain separate: the configured scheduler ceiling, the shared or reported KV-token pool, and the deepest concurrency actually measured. The retained GLM campaigns make the distinction concrete:
| Profile | Scheduler | Shared/reported pool | Measured long concurrency | Measured short concurrency | Interpretation |
|---|---|---|---|---|---|
| Current EXL3 r7 no-spec, 327,680 total / 65,536 output | C4 | 1,638,400 reported KV tokens; allocation only | 9/9 context at 255,647–255,672 prompt tokens plus reserve; concurrent full-window capacity unmeasured | strict120 120/120 at C4 | selected text-only lane; no full-window concurrency or soak claim |
| Historical ormandj SGLang adaptive EAGLE, 393K | 1 | 393,216 configured shared tokens | C1 at 304,491 prompt tokens | C1 only | retained historical selected lane; higher admission is unqualified |
| Historical EXL3 K3 + corrected DFlash2 K5 rollback, 524K | 16 | 2,493,817 reported KV tokens / 4.76 windows | C2 at 206,630 prompt tokens/request, 2/2 | no corrected-524K C16 artifact | proven long-context C2 headroom; C16 remains only a scheduler ceiling |
| BrandonMusic vision fixed K5, 262K | 16 | 560,866 / 2.14 windows | C1 at 206,296 prompt tokens | C16 at 4K, 16/16, 28.3 agg tok/s |
successful short batching, not sixteen 262K windows |
| BrandonMusic text fixed K5, 524K | 16 | 565,898 / 1.08 windows | C1 at 495,045 prompt tokens | C16 at 4K, 16/16, 23.85 agg tok/s |
effectively one full 524K window despite maxseq16 |
| BrandonMusic text no spec, 524K | 16 | 1,603,111 / 3.06 windows | C1 at 495,045 prompt tokens | C16 was measured on the 262K companion only | removing speculation preserves more KV headroom; 524K C16 remains unmeasured |
The historical 524K rollback's EXL3 K3 target, FP8 DS-MLA KV, TP=2/DCP=2 layout, and different runtime/state caches plausibly contribute to its larger reported pool. They were not isolated one at a time, so the cross-run comparison proves the resulting capacity, not a per-setting causal breakdown. In that dated comparison, SGLang measured 112.07 tok/s decode at 4K versus 83.08 for the 524K rollback. Neither figure is a matched comparison with the current C4 lane.
For that historical fixed 393,216-token pool, an equal C2 share is 196,608 total tokens/request before practical headroom. A first text-only qualification at 180,000 prompt + 4,096 output per request would leave 25,024 shared tokens; C4 at 80,000 + 4,096 would leave 56,832. Both are planning bounds, not results, and require scheduler/graph-state changes plus full managed requalification. See the cross-run interpretation and exact artifact links.
Remote coding-agent comparison¶
Both campaigns used AI-MBP25 as the isolated worker and reached Primary Node through the router. They did not use the same model recipe, GPU count, harness revision, or complete workload, so this is a bounded capability comparison—not a statistically controlled model ranking.
| Subject | Serving topology | Agentic evidence | SWE-bench Verified evidence | What the comparison supports |
|---|---|---|---|---|
| Qwen3.8 27B official FP8, SGLang TP=1 MTP=3 | 1x RTX PRO 6000 Max-Q active; second equal card idle | smoke 2/2; scout 16/18; tool-recovery-error 2/2; both failures were debug-loop protocol failures |
5/5 fixed scout tasks officially resolved; 19-57 model requests/task; 42m29s overall | strongest local bounded repository-agent evidence; broader SWE sample and efficiency-aware debug follow-up still required |
| DeepSeek V4 Flash 0731 r16 DSpark K5 | exclusive TP=2 on 2x RTX PRO 6000 Max-Q | tool-recovery-error 0/1: retry protocol passed, required final answer failed |
django__django-11099 1/1 officially resolved |
qualified the then-new worker/grader path; too small for an intelligence ranking |
The exact SWE overlap is only django__django-11099, which both models
resolved. Qwen additionally resolved the fixed pytest, scikit-learn, requests,
and SymPy tasks. The shared tool-recovery-error scenario favors Qwen in the
retained evidence, but the benchmark source and request controls changed
between campaigns. Treat that as a useful directional result, not a matched
head-to-head score. See the Qwen scout
and DeepSeek smoke.
Official-FP8 MTP-depth controls¶
MTP=4 and MTP=5 used the then-current FP8 TP=1/393K/maxseq1/batch4096 recipe and were swapped across the equal cards. Values are means of each run's median; three repetitions were taken in the first placement and two after the swap.
| Setting | Card role | TTFT | Prefill | Decode | E2E | Decision |
|---|---|---|---|---|---|---|
| MTP=4 | Compute A | 0.836 s | 4,319 tok/s | 97.9 tok/s | 1.276 s | no-promotion |
| MTP=5 | Compute A | 0.844 s | 4,280 tok/s | 98.3 tok/s | 1.281 s | no-promotion |
| MTP=4 | Compute B | 0.841 s | 4,288 tok/s | 90.4 tok/s | 1.313 s | no-promotion |
| MTP=5 | Compute B | 0.845 s | 4,274 tok/s | 91.6 tok/s | 1.315 s | no-promotion |
The apparent first-placement winner reversed after the card swap. On a fixed card, MTP=5 adds only 0.4-1.3% decode and makes E2E slightly worse. Both settings passed 388,979-token retrieval and bounded behavioral gates. The then-current MTP=3 Compute B control remained ahead at 93.6 tok/s and 1.295 seconds E2E. See the MTP-depth qualification.
SGLang no-speculation controls¶
The digest-pinned SGLang A/B held TP=1, 393,216 context, one request, FP8 E4M3 KV, FlashInfer, 2K prefill chunks, disabled prefix cache, text-only mode, and no speculation fixed. Each model ran three repetitions on its first card and two after the model placements were swapped.
| Model | Runs | TTFT | Prefill | Decode | E2E | Near-limit result | Decision |
|---|---|---|---|---|---|---|---|
| Official FP8 | 5 | 0.554 s | 6,512 tok/s | 48.0 tok/s | 1.451 s | 388,979 tok pass; 258.13 s TTFT | qualified control, no-promotion |
| Inferact NVFP4 | 5 | 0.429 s | 8,409 tok/s | 57.9 tok/s | 1.244 s | 388,979 tok pass; 248.75 s TTFT | text challenger, no-promotion |
NVFP4 reduced matched TTFT 22.6% and raised decode 20.6%; the ranking held after the card swap. It still trailed the then-current vLLM MTP=3 service's 93.6 tok/s decode. SGLang multimodal is not qualified on this WSL2 host because its default CUDA-IPC feature warmup failed; these rows are text-only. See the SGLang/NVFP4 qualification.
Dual RTX PRO 6000 Blackwell Max-Q — 192 GB aggregate, exclusive TP=2¶
These rows used both 96 GB cards over PCIe without NVLink. Each candidate was the sole inference owner; ordinary split-mode, Omni, voice, router, and other model workloads were offline. The original campaign changed no production alias; later rows record separately approved promotions.
| Model / config | Status | Quant · KV | Served · validated context | Reasoning contract | First-output latency | Effective prefill | Completion rate · decode | Recipe |
|---|---|---|---|---|---|---|---|---|
| GLM-5.3-Flash ormandj W4A16/NVFP4, SGLang | historical 2026-09-02 text/tools/image/OCR selection; exclusive TP=2 | ModelOpt W4A16/NVFP4 K32 experts · FP8 KV · adaptive EAGLE | 393,216 · 304,491 actual prompt tokens at the 380K target | explicit off/on request control; both qualified | 0.177 s @4K; 16.748 s @120K; 38.736 s @262K; 55.801 s @380K | 16,729 / 5,749 / 5,579 / 5,457 tok/s | tools 20/20; coding 15/15; media 12/12; endurance 60/60; 112.07/96.17/102.42/99.79 tok/s decode; model-only reserve waiver; routed and real-client pass | pinned recipe |
| GLM-5.3-Flash EXL3 K3 plus corrected DFlash2 K5 | historical same-model rollback; exclusive TP=2 | EXL3 K3 · FP8 DS-MLA target KV · BF16 draft KV · fixed K5 | 524,288 · 206,296-actual-token retrieval; measured C2 nominal 250K | low performance; high bounded quality | 0.974 s @4K; 59.67-59.80 s run medians @240K | retained artifacts | tools 20/20; image/OCR pass; 83.08 tok/s at 4K and pooled 69.99 at 240K; 2,493,817 KV tokens; C2 2/2 | pinned recipe |
| Qwen3.8 Flash Next RadixArk NVFP4 | historical text/image/OCR/video rollback; exclusive TP=2 | ModelOpt NVFP4 · BF16 KV · MTP3 | 262,144 · 253,703 actual prompt tokens with 8,192 output request | thinking disabled | 0.141 s @4K; 6.932 s @128K target; 29.214 s at full reserve | 18,096 tok/s @128K target; 8,684 at full reserve | tools 20/20; media direct 30/30/live 57/60; 155.9/114.7/112.9/102.0 tok/s at 4K/128K/254K/full reserve | pinned recipe |
| GLM-5.3-Flash vision fixed K5, 262K | preferred interactive GLM challenger, no-promotion |
TR3/EXL3 4 bpw · NVFP4 DS-MLA KV · fixed MTP5 · visual tower | 262,144 · 250K target / 206,296 actual retrieval | low measured; high bounded quality | 1.052 s @4K; 32.838 s @128K | 3,138 tok/s @128K | image/OCR, tools 20/20, coding 15/15 high; 72.8/55.7 tok/s decode; 560,866 KV tokens / 2.14 windows | pinned recipe |
| GLM-5.3-Flash text fixed K5, 262K | matched text control, no-promotion |
TR3/EXL3 4 bpw · NVFP4 DS-MLA KV · fixed MTP5 | 262,144 · 128K performance; 524K companion validates 495,045 actual | low measured; high bounded quality on companion | 1.055 s @4K; 32.898 s @128K | 3,133 tok/s @128K | tools 20/20; 69.8/61.9 tok/s decode at 4K/128K | pinned recipe |
| GLM-5.3-Flash no spec, 524K | preferred maximum-context/headroom GLM challenger, no-promotion |
TR3/EXL3 4 bpw · NVFP4 DS-MLA KV | 524,288 · 495,045 actual retrieval and 497,976 actual long-tool prompt | low measured | 144.3 s retrieval E2E | near-limit rate retained in artifact | exact needle and valid tool pass; 1,603,111 reported KV tokens / 3.06 full windows | pinned recipe |
| GLM-5.3-Flash fixed K5, 524K | single-user maximum-context experiment, no-promotion |
TR3/EXL3 4 bpw · NVFP4 DS-MLA KV · fixed MTP5 | 524,288 · 495,045 actual retrieval and 497,976 actual long-tool prompt | low capacity; high bounded quality | 166.0 s retrieval E2E | near-limit rate retained in artifact | exact needle, valid tool, coding 15/15 high; 565,898 KV tokens / 1.08 full windows | pinned recipe |
| GLM-5.3-Flash no spec, 262K | matched reliability control, no-promotion |
TR3/EXL3 4 bpw · NVFP4 DS-MLA KV | 262,144 · 128K performance | low measured | 1.127 s @4K; 30.891 s @128K | 3,422 tok/s @128K | tools 20/20; 42.7/43.5 tok/s decode; c16 short 16/16 at 33.69 aggregate output tok/s | pinned recipe |
| DeepSeek V4 Flash 0731, Infernal Invocation r18 DSpark K5, 1M/maxseq8 | former text Primary; retained TP=2 evidence | B12X W4A8 · FP8 compressed MLA KV | 1,048,576 · 1,040,063 actual prompt tokens | default thinking | matched capacity recorded in finding | matched K5/no-spec evidence retained | 142.1/129.5 tok/s decode at 4K/32K; repeated quality and real-client acceptance pass | promotion |
| DeepSeek V4 Flash 0731, r33 DSpark K5, 393K/maxseq16 | managed TP=2 rollback; historical Primary | B12X W4A8 NVFP4 MoE / FP8 dense · FP8 DS-MLA KV | 393,216 · 359,900 actual tokens direct | high-reasoning quality prior; routed functional pass with retained reasoning-evidence caveat | 65.2 s TTFT @359,900 tok | 5,599 tok/s | direct capacity pass; OpenClaw/Hermes 393K/32K/high client paths pass; client >300K open | pinned recipe |
| Qwen3.8 27B official FP8, TP=2 control | challenger, no-promotion; text-only |
Official FP8 weights · FP8 KV | 393,216 / 600,000 / 1,010,000 · 388,979 / 598,729 / 985,107 actual tokens | thinking disabled | 154.8 / 321.2 / 784.1 s cold TTFT at the three limits; 0.71-0.72 s p50 at 4K | 2,512 / 1,864 / 1,256 tok/s at the limits | all functional/context gates pass; 48.8-49.2 tok/s 4K decode | matrix recipes and evidence |
| Qwen3.8 27B official FP8, TP=2 MTP=3 | interactive-text challenger, no-promotion |
Official FP8 weights · FP8 KV · MTP=3 | 393,216 / 600,000 / 1,010,000 · 388,979 / 598,729 / 985,107 actual tokens | thinking disabled | 158.8 / 324.2 / 780.1 s cold TTFT at the three limits; 0.74-0.76 s p50 at 4K | 2,449 / 1,847 / 1,263 tok/s at the limits | all functional/context gates pass; 85.9-91.6 tok/s 4K decode; 7-9% KV-token cost | matrix recipes and evidence |
| Qwen3.8 27B official BF16, TP=2 control/MTP=3 | multimodal-reference challenger, no-promotion |
BF16 weights · FP8 KV · control or MTP=3 | 393,216 / 600,000 / 1,010,000 · 388,979 / 598,729 / 985,107 actual tokens | thinking disabled | control 168.7 / 344.1 / 820.6 s cold TTFT; MTP no consistent limit-TTFT win | control 2,305 / 1,740 / 1,201 tok/s at the limits | all functional/context gates pass; 35.4-36.1 control and 67.4-75.6 MTP 4K decode | matrix recipes and evidence |
| Qwen3.5 122B A10B NVFP4 | no-promotion TP=2; single-card profile remains rollback |
ModelOpt NVFP4 · BF16 KV | 262,144 · 128K | off for matched gates | 2.32 s TTFT @29,804 tok; 14.59 s @125,444 tok | 12,821 / 8,570 tok/s | 12/12 @32K · 67.5 tok/s; 4/4 @128K · 65.0 tok/s | campaign registry |
| Nemotron 3 Super 120B NVFP4 | no-promotion |
NVFP4 · FP8 KV · EP=2 | 65,536 · 60K | off for matched gates | 2.84 s TTFT @28,438 tok; 5.58 s @53,820 tok | 10,025 / 9,646 tok/s | 12/12 @32K · 59.5 tok/s; 4/4 @60K · 60.0 tok/s | campaign registry |
| Laguna S 2.1 NVFP4 | no-promotion TP=2; single-card profile remains rollback |
NVFP4 · FP8 KV | 262,144 · 240K | must be off | 1.97 s TTFT @29,834 tok; 31.85 s @231,457 tok | 15,134 / 7,252 tok/s | 12/12 @32K · 70.9 tok/s; 4/4 @240K · 66.0 tok/s | campaign registry |
| DeepSeek V4 Flash 0731, r16 DSpark K5 | priority challenger, no-promotion; reserve fail |
B12X W4A8 NVFP4 MoE / FP8 dense · FP8 MLA KV | 131,072 · 128K | low/high/max functional; low measured | warmed 19.44 s TTFO; 23.81 s visible TTFT @125,785 tok | 6,469 tok/s from TTFO | 27/27 coding-agent; 128K pass; 130.7 tok/s matched 4K decode | pinned recipe |
| DeepSeek V4 Flash 0731, r16 DSpark K5, 650K/maxseq16 | historical human-approved Primary; reserve waived; retuned 2026-08-06 to 262,144/batched 8192 (functional preflight only); image upgraded 2026-08-07 to r27 (performance figures in this row were measured on r16) | same B12X/FP8 recipe · FP8 MLA KV | 262,144 (was 650,000) · ~640K at the 650K envelope | low measured; high-default client smokes pass | 2.66 s TTFO; 5.46 s visible TTFT @32K; 120.6 s near-limit retrieval | 8,793 tok/s @32K | Pi protocol; Dark/Mini Pi and Mini OpenClaw; 3/3 @32K · 141.6 tok/s median decode | pinned recipe |
| DeepSeek V4 Flash 0731, r16 DSpark K5, 1M/maxseq4 | deep-session Pi experiment, no-promotion; reserve waived for qualification |
same B12X/FP8 recipe · FP8 MLA KV | 1,000,000 · ~985K | low measured | 2.96 s TTFO; 5.00 s visible TTFT @32K; 235.7 s near-limit retrieval | 7,901 tok/s @32K | Pi protocol pass; 3/3 @32K · 119.9 tok/s median decode | pinned recipe |
| DeepSeek V4 Flash 0731, r16 DSpark K5, 1M/maxseq16 | rejected for client-facing Primary; experimental capacity evidence |
same B12X/FP8 recipe · FP8 MLA KV | 1,000,000 · ~985K | low measured | 2.92 s TTFO; 4.58 s visible TTFT @32K; 237.7 s near-limit retrieval | 8,012 tok/s @32K | Qualification passed, then real client shapes fatally exceeded B12X workspace twice; 3/3 @32K · 129.0 tok/s median decode | retained recipe |
| DeepSeek V4 Flash 0731, r16 DSpark K5, 1M/maxseq1 | rejected for Pi agentic use; retained failure |
same B12X/FP8 recipe · FP8 MLA KV | 1,000,000 · ~985K | low measured | 3.25 s TTFO; 10.52 s visible TTFT @32K; 242.3 s near-limit retrieval | 7,195 tok/s @32K | 3/3 @32K · 13.0 tok/s median decode; fatal B12X workspace error on three-tool burst | retained recipe |
| DeepSeek V4 Flash 0731, r16 DSpark K5 + native offload | capacity challenger, no-promotion; reserve unmeasured |
same B12X/FP8 recipe · FP8 MLA KV · 8/16 GiB CPU offload | 262,144 · 250K | low functional/capacity | cold 43.75 s TTFO @249,573 tok; reload 0.825 s TTFO / 1.974 s visible TTFT @113,674 tok | cold 5,705; reload 137,856 effective tok/s | 250K pass; 16 GiB lane: 113,408 external hits, 1.002 GB CPU-to-GPU in 0.344 s | reload recipe |
| DeepSeek V4 Flash 0731 | challenger, no-promotion |
publisher FP4 experts / FP8 · FP8 E4M3 KV | 32,768 · 30K | reasoning_effort=low |
2.70 s TTFO; 29.11 s first-visible TTFT @21,144 tok | 7,818 tok/s from TTFO | 11/12 · 11.5 tok/s combined reasoning/visible | campaign registry |
| Inkling Small NVFP4 | no-promotion |
ModelOpt NVFP4 · BF16 KV/SWA | 32,768 · 30K | reasoning_effort=low |
2.79 s TTFO; 4.63 s first-visible TTFT @21,879 tok | 7,844 tok/s from TTFO | 12/12 · 73.5 tok/s combined reasoning/visible | campaign registry |
The three thinking-disabled rows use capacity-v3. DeepSeek and Inkling use
capacity-v4-reasoning, whose generation interval begins at the first
reasoning or visible delta. Their decode values include reasoning tokens and
are not visible-only rates. Inkling's separate reasoning-off 32K lane completed
12/12 at 2.84-second TTFT, 7,887 tok/s effective prefill, and 74.6 tok/s
visible decode. The one failed DeepSeek request exhausted 2,048 completion
tokens in reasoning without producing a visible answer.
The r16 DeepSeek row is a separate runtime and workload from the earlier SGLang row. Its 130.7 tok/s figure is the median of three successful run-level p50 decode values at 4K/c1 with a 2,048-token output cap. A same-image no-spec control measured 64.9 tok/s: DSpark improved decode by 101.4%, aggregate output by 70.5%, and E2E by 58.8%, while using 1.6-2.3 GiB more VRAM. Both profiles failed the per-card 3 GiB reported-free reserve.
The 650K and 1M Pi rows use a separate matched three-request 32K/c1 capacity
shape with a 1,024-token output cap. Their decode values are run-level medians,
not controlled long-gen rates. Near-limit retrieval is a separate one-request
needle gate. Moving display output to the AMD iGPU enabled the larger graph
envelopes, but reported free VRAM fell to 797/805 MiB after the 650K workload
and 339/335 MiB after the preferred 1M/maxseq16 probe; maxseq4 measured 207/209
MiB in its separate run. After qualification, the 1M/maxseq16 profile also
crashed on two real agent shapes: 703.64 and 687.83 MiB workspaces were required
against 514.25 MiB available. The second used a 19,118-token Pi prompt and a
5,120 output cap, so router output clamping is not a sufficient 1M mitigation.
The r16 650K and r33 393K rows are historical promotion/capacity evidence.
The Infernal Invocation r15 393K row is now the human-approved text Primary
and uses the operator-approved AI-only policy without a separate
graphics/co-resident reserve gate.
The native-offload row uses a narrowly derived WSL2 image and a different 262,144-token admission ceiling. Its cold 250K result is capacity evidence, not a matched speed comparison with the 131K row. The 8 GiB follow-ups reused the GPU prefix cache. The separate 16 GiB run sized CPU above the measured GPU KV tier and proved reload through unchanged GPU-hit counters, 113,408 new external hits, and 1,001,721,600 new CPU-to-GPU bytes.
RTX PRO 6000 Blackwell Max-Q — 96 GB, sm_120, 300 W¶
The card is power-limited to 300 W (Max-Q); treat external reports from higher-TDP cards as advisory only.
Historical serving chain¶
| Model / config | Status | Quant · KV | Context · adm. | Thinking | TTFT | Output rate | Recipe |
|---|---|---|---|---|---|---|---|
| GLM-5.3-Flash ormandj W4A16/NVFP4, SGLang | historical 2026-09-02 text/tools/image/OCR Primary | ModelOpt W4A16/NVFP4 · FP8 KV · adaptive EAGLE | 393,216 · C1; 4,096 output; image/OCR, no video | explicit thinking control | 0.177 s @4K; 55.801 s @380K target | 112.07 tok/s decode @4K; 99.79 @380K target | 393K promotion |
| GLM-5.3-Flash EXL3 K3 plus corrected DFlash2 K5 | historical same-model rollback | EXL3 K3 · FP8 DS-MLA target KV · BF16 draft KV · fixed K5 | 524,288 · router c16; measured C2 nominal 250K; 16 images, no video | low performance; high bounded quality | 0.974 s @4K; 59.67-59.80 s run medians @240K | 83.08 tok/s decode @4K; pooled 69.99 @240K | 524K xgrammar qualification |
| Qwen3.8 Flash Next RadixArk NVFP4 | historical text/image/OCR/video rollback | ModelOpt NVFP4 · BF16 KV · MTP3 | 262,144 · c1; four images or one video; client 253,952+8,192 | thinking disabled in qualified Chat path | 0.141 s @4K; 6.932 s @128K target | 155.9 tok/s decode @4K; 114.7 @128K target; 112.9 @254K target | vision promotion |
| DeepSeek V4 Flash 0731, Infernal Invocation r18 DSpark K5, 1M/maxseq8 | former text Primary | B12X W4A8 · FP8 compressed MLA KV | 1,048,576 · router concurrency 1 | default reasoning | see dated promotion | 142.1 tok/s decode @4K; 129.5 @32K | historical promotion |
| Qwen3.8 27B official FP8, SGLang TP=1 MTP=3 | former Primary/general-vision/OCR/video service | Official FP8 · FP8 E4M3 KV | 393,216 · 1 request; 2 images/request; 1 video/request | server default disabled; chat caller override | 0.577 s median @4K | 111.4 tok/s decode | video expansion |
| Qwen3.8 27B official FP8 + BF16 vLLM split | former managed split | Official FP8/BF16 · FP8 KV | 393,216 · 1 seq each; BF16 32 images/one video | disabled default, caller override | 0.834 / 0.884 s median @4K | 93.6 / 62.0 tok/s decode | historical promotion |
| Qwen3.5 122B A10B NVFP4 | retained qualified recipe; not immediate restore | ModelOpt NVFP4 · BF16 KV | 262,144 · c1 | default on, per-request disable | 0.15 / 0.26 s p50/p95 @8K c1 · 68.91 s @231K, c1 matched lane | 60.3 tok/s decode @231K · 59.45 agg @8K c1 |
registry |
| Agents-A1 official FP8 multimodal plus Omni | managed split restoration | compressed-tensors FP8 · FP8 KV | 262,144 · c1 | must be off | 0.25 / 0.53 s @8K c1 · 32.97 / 33.44 s @231K c1 | 188.1 tok/s decode @8K · 155.8 tok/s decode @231K | promotion-era evidence |
| Laguna S 2.1 NVFP4 | rollback |
NVFP4 · FP8 KV | 262,144 | must be off | 0.07 / 0.55 s @c1 · 3.44 / 4.37 s @c8 · quality ctx 2.26 / 21.15 / 50.64 s @32K/128K/240K | 75.46 agg @c1 · 83.24 agg @c8 |
registry |
| GPT-OSS Puzzle 88B | rollback |
MXFP4 + Marlin MoE · FP8 KV | 131,072 · 8 seqs | reasoning_effort (low for gates) |
0.393 / 0.956 s @8K c1 · 0.766 / 1.075 s @8K c8 · 25.906 s @128K | 3.85 / 17.85 agg — only 20 / 86 output tokens; not a decode rate |
full recipe |
Qwen3.5 122B is the only row here on BF16 KV; preserve that exact rollback, along with c1 admission and its one-image limit. Agents-A1 keeps thinking disabled, four-image/one-video admission, and the rejected MoE tune inactive.
Other evaluated candidates¶
| Model / config | Status | Quant · KV | Context · adm. | Thinking | TTFT | Output rate | Recipe |
|---|---|---|---|---|---|---|---|
| GPT-OSS 120B | no-promotion |
MXFP4 + Marlin · FP8 KV | 131,072 | reasoning_effort |
655.67 / 1257.35 ms @8K · 28.9 s @128K needle | 183.2 long-gen |
registry |
| Nemotron 3 Puzzle 75B NVFP4 | no-promotion |
NVFP4 MoE + MTP 3 · FP8 KV | 131,072 · 2 seqs | off | 458.93 / 492.91 ms @8K · 13.2 s @128K needle | 137.0 long-gen (MTP 1.50× from 91.4) |
registry |
| MiniMax M2.7 REAP 139B | no-promotion |
NVFP4 · FP8 KV | 65,536 · c1 | off (no parser) | 86 ms warm @8K · 14.3 s @64K | 97.2 agg @c1 (2,179 out tok) |
registry |
| Qwen3.6 27B community NVFP4 + MTP | no-promotion |
ModelOpt NVFP4 + MTP 3 · FP8 KV | 262,144 · 5 seqs | off for capacity | 0.63 s @c1 · 3.22 s @c5 · 26.5 s @131K needle | 95.0 long-gen (MTP 1.36× from 69.9) |
registry |
| Ornith 1.0 35B FP8 | no-promotion |
compressed-tensors FP8 · FP8 KV | 131,072 | off | 772 ms warm @8K · 13.1 s full 131K prefill (fastest of its set) | 29.2 agg @c1 (273 out tok over 10 req) |
registry |
| Agents-A1 BF16 multimodal | challenger · no-promotion |
BF16 · FP8 KV | 131,072 · c16 text, media c1 gated | must be off | 0.30 / 0.35 s @8K c1 · 1.50 / 4.82 s @8K c16 · 11.99 / 12.08 s @128K c1 | 89.98 agg @c1 · 162.33 agg @c16 |
multimodal recipe |
| Agents-A1 ProtoLabs NVFP4 text | challenger · no-promotion |
NVFP4 → Marlin W4A16 · FP8 KV | 131,072 · c16; vision excluded | must be off | 0.27 / 0.32 s @8K c1 · 1.08 / 4.32 s @8K c16 | 104.58 agg @c1 · 197.93 agg @c16 compact |
compact recipe |
| Nemotron 3 Super 120B NVFP4 | no-promotion |
NVFP4 · FP8 KV | 131,072 · 5 seqs | both; 1,024 headroom recommended | 0.62 s @c1 · 2.52 s @c5 · 16.76 s @131K | 33.19 agg @c1 · 45.90 agg @c5 |
cand-nemotron3-super-120b |
| Mistral Small 4 119B NVFP4 | no-promotion |
NVFP4 · FP8 KV | 131,072 · 5 seqs | reasoning_effort; 2,048 headroom |
0.30 s @c1 · 1.85 s @c5 · 51.90 s @131K | 57.82 agg @c1 · 67.04 agg @c5 |
cand-mistral-small4-119b-nvfp4 |
| ThinkingCap Qwen3.6 27B FP8 | no-promotion |
compressed-tensors FP8 · FP8 KV | 262,144 · 5 seqs | on by default (256 + 4,096 headroom); rates below captured with it off | 1.01 s @c1 · 4.66 s @c5 · 32.3 s @131K needle | 6.661 agg @c1 (45 out tok) · 7.92 agg @c5 (42 out tok) |
registry |
| Qwen3.6 27B official FP8 | no-promotion |
FP8 + MTP 3 · FP8 KV | 262,144 · 5 seqs | off for capacity | 1.59 s @c1 · 5.68 s @c5 · 32.9 s @131K | 5.627 agg @c1 (55 out tok) · 8.31 agg @c5 (53 out tok) |
cand-qwen36-fp8 |
| Unsloth Qwen3.6 27B NVFP4 | no-promotion |
NVFP4 + MTP 2 · FP8 KV | 262,144 · 5 seqs | off for capacity | 968.07 ms @c1 (1 req) · 3.68 s @c5 | 10.497 agg @c1 — 1 req / 14 out tok; not a decode rate · 15.21 agg @c5 (66 out tok) |
cand-unsloth-qwen36-27b-nvfp4 |
| Qwen3.5 122B NVFP4, NGC 26.04 | no-promotion |
ModelOpt NVFP4 · FP8 KV | 131,072 · c1 | off for gates | 223 ms p50 @8K · ~28 s @100K | 38.8 agg @c1 (10 req × 8K) |
earlier candidate window |
| Qwen3.5 122B MXFP4 / Marlin | no-promotion |
MXFP4 → Marlin W4A16 · FP8 KV | 131,072 · 2 seqs | off | 720.79 / 974.40 ms @8K · 25.8 s @128K needle | 30.57 agg |
registry |
Gemma 4 family — 2026-07-16 template bakeoff¶
All rows: vLLM 0.25.1, FP8 KV, 256K Heavy window, three attempts per check at 100% pass.
Capacity columns are mixed short-generation workloads — agg, not decode.
| Config | Status | Quality | 32K cap. c1 | 32K cap. c2 | Quality-ctx TTFT 32K / 128K / 240K | 1,024-tok diagnostic |
|---|---|---|---|---|---|---|
| Gemma 4 12B QAT W4A16 | no-promotion |
pass | 1.52 s · 21 agg |
0.27 s · 54 agg |
6.96 / 44.61 / 97.33 s | 109.03 long-gen |
Gemma 4 26B-A4B BF16 (gemma-4-26B-A4B-it) |
rejected |
fail (timeout triage 0/3) | 0.73 s · 36 agg |
0.31 s · 77 agg |
capacity 11.93 s @120K · 34.07 s @240K | not measured |
| Gemma 4 31B W4A16 | rejected (latency) |
pass | 4.02 s · 7 agg |
0.41 s · 19 agg |
15.44 / 112.30 / 248.57 s | 57.8 long-gen — 07-17 probe: 1,024 tok on a 128K serve, not this column's 256K |
| Unsloth Gemma 4 12B NVFP4 | no-promotion |
fail (tool 1/3) | 21 agg |
76 agg |
3.23 / 32.70 / 81.47 s | 98.86 long-gen |
| Unsloth Gemma 4 26B-A4B NVFP4 | no-promotion |
fail (timeout triage 1/3) | 45 agg |
122 agg |
1.83 / 18.93 / 48.27 s | 191.46 long-gen |
| Unsloth Gemma 4 31B NVFP4 | no-promotion |
pass | 7 agg |
30 agg |
9.39 / 92.92 / 223.32 s | 51.49 long-gen |
Under 128 concurrent requests the ranking inverts — NVFP4 wins on density where it lost at c1:
| Runner · config | c128 @1K | c128 @8K | 8K p95 TTFT | c1 @1K | c8 @1K |
|---|---|---|---|---|---|
| Gemma 4 12B QAT W4A16 | 2,042 c128 |
1,053 c128 |
1.58 s | 71 | 578 |
| Unsloth 12B NVFP4 | 2,770 c128 |
1,526 c128 |
1.16 s | 65 | 516 |
| Unsloth 26B-A4B NVFP4 | 3,227 c128 |
1,466 c128 |
1.27 s | 91 | 853 |
| Unsloth 31B NVFP4 | 1,720 c128 |
799 c128 |
2.37 s | 36 | 335 |
Rejected or unmeasurable on this card¶
| Config | Outcome |
|---|---|
| GLM-5.3-Flash adaptive K1-K5 plus ReplaySSM | rejected — 12/20 repeated tools and degenerate repeated handle output despite 71.1/59.5 tok/s decode at 4K/128K. Retained only as negative performance evidence. |
| Laguna XS 2.1 NVFP4 | rejected — corrupted text and 0/20 tools with FP8 KV; stalled without it; SGLang path returned empty 131K needle. No trustworthy numbers. |
| DeepSeek V4 Flash NVFP4, 2026-07-10 single-card attempt | historical rejected — NGC vLLM 0.19 rejected the architecture; nightly load aborted at shard 18/46. Nothing measured in that lane; the 0731 TP=2 result above supersedes it for current compatibility. |
| Gemma 4 31B native MTP | Incompatible — assistant projection dims 6400 vs 10752. |
RTX 5090 — 32 GB, sm_120¶
The current qualification lane uses the 5090 exclusively for one candidate. The older Omni rows describe historical Primary Node reservation shapes; their 27,999 MiB usable budget after a 4,608 MiB system/audio reserve is not the current qualification policy.
| Model / config | Status | Quant · KV | Context · adm. | Thinking | TTFT | Output rate | VRAM |
|---|---|---|---|---|---|---|---|
| Qwen3.8 27B NInfer NVFP4 + MTP3 | preferred direct text/tools performance challenger, no-promotion |
NVFP4/row-scaled FP8 target · INT8 KV · MTP3 | 252,928 · c1; 201,746 actual prompt + 8,192 output cap | disabled | 0.430 s short; 70.4 s at 201,746 actual | 165.9 tok/s short decode; bounded quality pass; C1 tools pass; 20-way burst 17/20 | 29,834 MiB used; 2,354 MiB free |
| Qwen3.8 27B Unsloth GGUF Q4_0 + MTP3, llama.cpp | retained GGUF incumbent and broad-capability challenger, no-promotion |
Q4_0 · Q4_0 KV · Q4_0 MTP3 | 262,144 · c1; 253,822 actual prompt + 8,192 reserve | disabled | 0.29 s short; 246.72 s mean at 253,822 actual | 104.1 tok/s short decode; tools 20/20; images 18/18 | 25,408 MiB startup; 25,460 MiB late |
| Qwen3.8 27B RadixArk NVFP4, SGLang TP=1 | preferred 5090 challenger, no-promotion |
ModelOpt NVFP4 · FP8 E4M3 KV | 131,072 · c1 | disabled | decode TTFT not measured; 119,675-token retrieval 29.8 s E2E | decode not measured; media 30/30; eight images / two videos 4/4 | 20.14 GB weights; 3,928 MiB free after startup |
| Nemotron Nano Omni 30B NVFP4 | current topology |
NVFP4 · KV not recorded | 65,536 · 2 seqs | off for text gates | 122 / 164 ms p50/p95 | 224.08 agg @c2 |
27,706 MiB observed; exclusive |
| Qwen2.5-Omni 3B | challenger · no-promotion |
not recorded | 32,768 (recipe) · 2 seqs | n/a | 0.04 / 0.06 s p50/p95 | 243 agg @c2 |
24,576 MiB reserved; co-resident |
| Gemma 4 E4B FP8-Dynamic | no-promotion (historical Fast) |
FP8-Dynamic · FP8 KV | 32,768 | — | 0.46 s @c1 · 0.58 s @c2 | 49 agg @c1 · 79 agg @c2 |
not measured |
| Gemma 4 E2B W4A16 | no-promotion |
W4A16 · FP8 KV | 131,072 | — | 0.43 s @c1 · 0.21 s @c2 | 96 agg @c1 · 204 agg @c2 |
not measured |
| Gemma 4 31B W4A16 (Fast lane) | rejected |
W4A16 · FP8 KV | 64K practical | — | 2.60 s @30K c1 · 8.81 s @c2 | 9 agg · 8 agg |
128K needs 6.35 GiB KV, only 4.28 available |
Gemma 4 26B-A4B BF16 (gemma-4-26B-A4B-it) |
rejected |
BF16 | — | — | not measured | not measured | 48.57 GiB model; negative 5.74 GiB KV headroom |
Both Omni probes used the same shape (6/6 requests, c2, 2,048-token prompt, 128-token cap) but were reported at different precision, so their TTFTs are not comparable at millisecond granularity.
2026-07-10/11 bakeoff — historical¶
Every row here is historical-invalid for exact rerun: no pinned checkpoint revisions. Per
methodology, historical-invalid does not establish promotion evidence, so
do not rank on these numbers — they are retained because a negative or partial result is
still evidence.
Rate cells in this table come from a 10-request, 256-token-cap agg harness. Batch shape is
not uniform across the page: request count, concurrency, and the completion cap all vary by
campaign. Compare agg values only within a table, and only when the row notes agree; follow the
row's dated finding for its exact shape.
Both llama.cpp rows additionally ran with a warm prefix cache (~0.87–0.90 hit rate),
which inflates their aggregates relative to the two measured vLLM rows in this table, neither
of which records prefix reuse.
| Config | Engine | Served context | Output rate | Warm TTFT @8K | Verdict |
|---|---|---|---|---|---|
| Qwen3.5 35B-A3B Q4_K_M | llama.cpp | 64K | 56.279 agg @c1 (424 out tok) |
178 ms | fast-tier candidate; not promoted |
| Gemma 4 E4B QAT UD-Q4_K_XL | llama.cpp | 64K | 96.96 agg @c1 (401 out tok) |
61 ms | low-latency specialist; not promoted |
| Nemotron Nano Omni 30B | vLLM nightly | 65,536 | 27.3 agg @c1 (236 out tok) |
675 ms | keep experimental |
| Nemotron 3 Nano 30B (text) | NGC vLLM 0.19 | 131,072 | 15.0 agg @c1 (322 out tok) |
1.68 s | keep experimental |
| Gemma 4 31B IT NVFP4 | vLLM gemma4-unified | — | not measured | not measured | rejected — all six configs died at KV allocation |
Audio — STT and TTS¶
All audio rows were measured on the RTX 5090; the PRO 6000 is never the measuring device for these numbers. The STT qualification additionally records the PRO 6000 as protected and running during the run; the TTS round-trip source does not mention that card either way.
STT ran a shared 30-case corpus (stt-corpus/v1, 170.4 s): 24 human recordings are the primary
quality set; 6 synthetic phrases are reported separately and must never be merged in.
| Model | Status | Primary-human WER | Sequential p50 / p95 | Sequential req/s | Concurrency-4 p95 | c4 req/s |
|---|---|---|---|---|---|---|
| Parakeet TDT 0.6B v3 | current |
3.343% | 72.35 / 177.87 ms | 12.80 | 240.43 ms | 21.00 |
| Qwen3-ASR 0.6B | challenger · no-promotion |
3.621% | 67.40 / 113.58 ms | 14.86 | 137.36 ms | 47.25 |
| Nemotron 3.5 ASR 0.6B | rejected |
6.685% | 121.60 / 225.45 ms | 7.87 | 747.82 ms | 8.68 |
Qwen3-ASR is 36.1% faster at sequential p95 and within the declared one-point non-inferiority margin — but Parakeet stays routed. Meeting the margin does not authorize a route change. Nemotron 3.5 regressed 3.343 points, exceeding the margin.
| Model | Status | Stage latency | RTF | Round trip |
|---|---|---|---|---|
| Kokoro TTS | current |
289.27 ms | 0.1006 | 710.68 ms total = 421.41 ms STT + 289.27 ms TTS, WER 0.0 |
Neither Parakeet nor Kokoro has a pinned checkpoint revision recorded — identity is pinned at the image level only. STT RTF and isolated per-model VRAM were never measured.
Where the recipes live¶
Start with the measured recipe results for a stable recipe entry, native metrics, strengths, limitations, and direct evidence. Use the reproduction guide to inspect the complete TOML and translate its fields to another container runner.
-
configs/serve-recipes.toml— 32 recorded recipes with quantization, context, flags, and environment. 9 of the 32 pin an image bysha256:digest; 21 pin a mutable tag (vllm/vllm-openai:nightly,lmsysorg/sglang:latest) and 2 record no image at all, so those 23 cannot be reproduced exactly. Inspect one locally: -
examples/primary-node/— the compose files and serve manifests behind the reference topology. - GPT-OSS Puzzle 88B recipe — the one model with a full standalone operator procedure, because it needs a pinned engine fork.
- Model dossiers — per-model status, identity, decision boundary, and dated history.
- Legacy recipe and gotcha page — retained for existing deep links; the dossiers win on any conflict.
What is deliberately absent¶
Rows carrying not measured are gaps in the evidence, not zeros. Notably: standalone prefill
throughput is not published anywhere (long-context TTFT is used as a labelled proxy instead),
controlled long-generation decode was never captured for the Qwen3.5 122B rollback, and
STT real-time factor was never measured for any ASR model.
Several configurations lack a pinned checkpoint revision and therefore cannot be re-run exactly — MiniMax M2.7 REAP, Ornith 1.0 35B, the historical 2026-07-10 DeepSeek V4 Flash attempt, Parakeet, Kokoro, and the whole 2026-07-10/11 bakeoff set. They are kept because a negative or partial result is still evidence, but they cannot ground a new equivalence claim.
Three 2026-07-12 deterministic planning scores (Qwen 1/5, Nemotron 0/5, GPT-OSS 0/5) are retained
as historical-invalid: the harness let hidden reasoning consume the entire completion budget.
The repository forbids using them for ranking, and so should you.
External benchmark data never appears in these tables. It is an advisory prior for choosing what to test, not a local result — see External benchmarks.