Skip to content

Benchmarks

Eight-image follow-up, 2026-09-18: GLM EXL3 r9 passed four eight-image comparison requests, one eight-image high-resolution request. Limit now eight/request; no concurrent eight-image soak.

2026-09-19 current decision: the GLM-5.3-Flash r10 APC promotion follow-up records explicit human approval and fresh bounded acceptance. The historical r9/r10 campaign remains the evidence and failure record; exact r9 remains the documented rollback.

Compare local models by useful context, quality, and serving performance. Start with the latest decision, choose your hardware, then follow the exact configuration and retained evidence. Evidence reviewed: 2026-09-19 UTC.

Current bounded APC reference — September 19

Human-approved GLM-5.3-Flash r10 APC · Dual RTX PRO 6000

The bounded r10 APC configuration retains the 4-bpw GLM checkpoint, v84 runtime, FP8 MLA KV, no speculation, 327,680 configured tokens, C4, and the eight-image contract. On frozen 32-request finalists with about 25K actual prompt tokens, it improved the shared-prefix path while staying within the unique-prefix gate.

Matched outcome r9 r10 APC
Shared-prefix visible TTFT 23.70 s 7.62 s (−67.84%)
Shared-prefix requests/s 0.146 0.462 (3.17×)
Unique-prefix requests/s 0.144 0.141 (−2.39%)

The 75.58% isolated cache-reuse fraction supports the intended repeated-prefix mechanism. This does not prove four simultaneous maximum-context requests, video readiness, a soak result, or broad intelligence/SWE parity. Roll back to the exact r9 recipe if correctness, containment, routed readiness, or the frozen unique-prefix gate fails.

Read the promotion follow-up · Inspect the historical campaign · Open the matched chart

September 14 quality and performance baseline

Historical r7 text and tool qualification · Dual RTX PRO 6000

GLM-5.3-Flash · EXL3 4 bpw · speculation off

The September 13–14 qualification selected this configuration after bounded quality, context, and client tests.

What was tested Result
Context / output capacity 327,680 / 65,536 tokens
Long-context retrieval 9/9 twice, including after cold reload; 255,647–255,672 input tokens plus output reserve
Quality / agentic / repository tasks 90/100 fixed MMLU-Pro sample · 30/30 agentic · 4/5 SWE tasks
Controlled requests / decode 120/120 · 38.91 tok/s median per request at C4

Limits of this baseline: tested text only; no exhaustive intelligence ranking, prolonged soak, reboot, or simultaneous full-window proof. These are dated local results, not a live status display or an automatic model-selection policy.

Read the qualification · Open the model dossier · Inspect the evidence

Choose by hardware and workload

1. Pick your hardware

Start with the measured RTX PRO 6000, RTX 5090, or Apple M4 Max view. GPU count, interconnect, operating system, and runtime are part of the result.

2. Name the workload

For interactive use, prioritize TTFT and E2E. For sustained serving, use aggregate output, TPOT/ITL, and tail latency. For long context, require actual prompt depth rather than configured context.

3. Check the tradeoff

Use measured recipe results for explicit strengths, limitations, and rejection reasons. Use the model and quantization selection guide for a bounded workload choice, then the complete model inventory when quality, tools, media, or client acceptance matters more than one speed number.

4. Reproduce and verify

Open the exact recipe, then follow its dated finding and native JSON artifact. A copied container configuration is not a reproduced result until the same workload and independent gates pass.

Start with comparable numbers

The model comparison table puts the wider measured configuration — TTFT, throughput, context, reasoning mode, and recipe link — in one place, grouped by card and workload. The measured recipe finder is the shorter path when you want a current configuration, its tradeoffs, and raw evidence.

For the newest single-card Qwen measurements, the dedicated Qwen3.8 27B RTX 5090 quant comparison shows every fresh no-speculation/speculation pair, full-context proof, and failure without compressing the operator's TTFT-first ranking into one score. The newer dual-PRO Qwen3.8 27B comprehensive campaign compares SGLang target-only/DFlash2 tuning, TP1/TP2/two-replica DP2, Inferact/RadixArk targets, and kelnei/vLLM MTP2/no-spec. It retains mean, p50, p95, and p99 TTFT, effective prefill, decode, TPOT/ITL, and E2E plus aggregate throughput and correctness failures. Open the campaign graph (SVG) or its plotted data and source hashes.

For the compact publication contract, use the finding format. It defines screenshot-ready result cards, the common campaign artifact-set template, X and Reddit variants, accessible alt text, and a claim ledger back to the full finding and raw artifacts.

Browse by hardware

Hardware Published evidence scope Start here
2× NVIDIA RTX PRO 6000 Blackwell Max-Q, 192 GB aggregate, sm_120 Measured dual-card reference configuration; split workloads or exclusive TP=2 RTX PRO 6000
NVIDIA GeForce RTX 5090, 32 GB, sm_120 Measured single-card qualification configuration, plus retained historical results RTX 5090
Apple M4 Max laptop, 48 GB unified memory Historical same-host voice candidate evidence; no LLM promotion or reference-topology change Apple M4 Max

The published dual-PRO reference configuration uses two equal PRO 6000 cards. Aggregate VRAM is not unified memory, and the cards communicate over PCIe without NVLink. Exclusive TP=2 runs must prove that both devices were selected and all other inference was offline. The published reference topology keeps Companion Node model-free. Historical mixed-card tests still say which card was measured and which was merely protected or co-resident.

Browse by model

The complete hardware-first inventory lists every dossier and explicit grouped variant, including the September 12 efficient Qwen models and the Apple MLX Qwen lanes. Each dossier is a stable, model-centered summary with its dated findings and raw evidence.

Recent evidence

  • 2026-09-15 — media: The media finding records bounded workflow evidence; use its dossier and finding rather than treating it as a text-model selection result.
  • 2026-09-12 — efficient Qwen variants: Signal, Swift, Qwopus, and Minitron all passed functional checks on the RTX 5090, but none cleared the exact incumbent's output and deployment gates. See the grouped dossier.

Run context, agentic, and repository evaluations

The context, agentic, and SWE job guide describes the registered-worker workflow, bounded profiles through 640K, independent scoring, pinned mini-SWE-agent and official SWE-bench grading, cancellation, and the measured-versus-prior evidence boundary. The guide defines methodology and commands; it does not claim a live model result by itself.

Published decisions and recent controls

The collapsed history below preserves earlier decisions through September 11. Labels such as current and rollback describe their status at that date; the September 13–14 selection above supersedes the primary-model guidance. Use the run catalog to compare dated configurations.

Earlier decisions and controls — historical configurations
1. **GLM-5.3-Flash ormandj W4A16/NVFP4 SGLang** — `current` text/tools/image/OCR published profile at 393,216 tokens/C1 in exclusive TP=2 across both RTX PRO 6000 cards, with adaptive EAGLE, explicit thinking control, a 4,096-token output cap, and real Pi/OpenClaw/Hermes acceptance. The standing reserve was explicitly waived for the model-only pair while the post-workload runtime safety gates remained in force. 2. **GLM-5.3-Flash EXL3 K3 plus corrected DFlash2 K5** — immediate exact rollback at 524,288 tokens/router-c16, measured C2 at a nominal 250K target, 8,192 maximum output, and up to 16 images. Video is unsupported and the DFlash2 draft is noncommercial without separate permission. The SGLang 245,760/C1 profile remains a conservative verified fallback. 3. **RadixArk Qwen3.8 Flash Next NVFP4** — immediate retained text/image/OCR/video rollback at 262,144 tokens, exclusive TP=2, router c1, four-image/one-video admission, and the qualified hash-gated SM120 QSA-fast MTP3 profile. 4. **Inferact Qwen3.8 27B NVFP4 plus DFlash2 K12/chunk1K on the PRO pair** — bounded aggregate-throughput winner: two independent TP1 replicas measured 1,401.8–1,423.4 tok/s at C16, versus one TP1 at 764.3 and TP2 at 587.9. All headline canaries passed 100/100, but TP2 failed strict JSON twice and is rejected. RadixArk is the lower-TTFT tradeoff; kelnei/vLLM MTP2 beat its exact no-spec control but not SGLang. Broad quality and a balancing/client path remain open; `no-promotion`, exact GLM service and route restored. 5. **Gittensor Qwen3.8 27B RTX5090 NVFP4 target-only SGLang** — preferred measured direct RTX 5090 TTFT challenger: 50.9 ms warm median TTFT, 79.5 tok/s decode, and a 244,002-token actual prompt. The advertised DSpark pair failed on incompatible shapes and FP8 KV scales need further validation; `no-promotion`, exact GGUF incumbent restored. Unsloth Dynamic V3.0 MTP3 is the strongest clean bounded-64K speculative arm at 137.7 tok/s with tools 20/20. 6. **Qwen3.8 27B NInfer NVFP4 MTP3** — retained direct decode challenger at 252,928 tokens/C1; 0.430-second median TTFT and 165.9 tok/s decode. Its admission, reserve, runtime-image, routed/client, and broader gates remain open. 7. **DeepSeek V4 Flash 0731 Infernal Invocation r18/r15** — former text Primary profiles with retained TP=2 long-context, performance, quality, and real-client evidence. 8. **Qwen3.8 27B official FP8 SGLang single service and FP8/BF16 vLLM split** — former text/image/OCR/video deployments with retained managed recipes. 9. **Qwen3.5 122B A10B NVFP4**, **Agents-A1 plus Omni**, **DeepSeek r16 650K**, **Laguna S 2.1**, and **GPT-OSS Puzzle 88B** — retained qualified or promotion-era recipes, not the immediate selected deployment. 10. **Gemma 4** and **ThinkingCap Qwen3.6 27B** — historical strict-quality controls. 11. **Nemotron 3.5 ASR** and **Qwen3-ASR 0.6B** — historical RTX 5090 measurements. The RTX PRO 6000 was protected, not benchmarked. 12. **RadixArk Qwen3.8 27B NVFP4** — historical RTX 5090 multimodal challenger; direct 128K text/tools/image/OCR/video evidence, no route or promotion. 13. **FLUX.2 Klein 4B FP8** — published as available for the RTX 5090 ComfyUI image workflow through Hermes/MCP at fixed 512/768/1024-pixel profiles; six of eight bounded visual samples passed strict review, with two retained draft fidelity/count failures. **Wan2.2 TI2V 5B** remains unavailable after its decodable sample failed prompt adherence and spatial quality. On 2026-09-02, the exact ormandj GLM-5.3-Flash W4A16/NVFP4 checkpoint and rc14 SGLang image qualified and was human-promoted at 393,216 tokens/C1 on both PRO 6000 cards. The selected adaptive-MTP lane passed tools 20/20, coding 15/15, media 12/12, endurance 60/60, and 4K/120K/262K/380K capacity through 304,491 actual prompt tokens. After a recorded model-only reserve waiver, the managed direct/routed gates and real Pi/OpenClaw/Hermes acceptance passed; 2,543 MiB/card remained after client work. The exact 524K incumbent is retained as rollback. See the [qualification](../findings/2026-09-02-glm53-sglang-sm120-qualification.md) and [promotion](../findings/2026-09-02-glm53-sglang-sm120-393k-promotion.md). A 2026-09-03 cross-run review clarifies that the old `maxseq16` setting is a scheduler ceiling, not proof of sixteen full-context requests. The current SGLang profile remains qualified at C1 with a 393,216-token shared pool. The corrected 524K rollback reports 2,493,817 KV tokens and separately proved C2 at 206,630 prompt tokens/request; the historical BrandonMusic C16 artifacts used 4K prompts. The review records bounded C2/C4 planning math for a future current-profile experiment without promoting it to measured evidence. See the [concurrency and KV-capacity interpretation](../findings/2026-09-03-glm53-concurrency-capacity-interpretation.md). On 2026-08-31, the exact GLM-5.3-Flash K3 target plus DFlash2 K5 draft moved to a digest-pinned xgrammar-corrected 524K profile after a matched no-speculation A/B and full rollback/forward drill. It measured 83.08 tok/s at 4K and a pooled 69.99 at 240K, versus 42.61 and 43.63 without speculation. The engine reported 2,493,817 KV tokens, C2 nominal 250K completed 2/2, and direct structured, quality, image/OCR, exact router, and real OpenClaw/Hermes/Pi gates passed. Video remains disabled and the DFlash2 license limits the recipe to evaluation/noncommercial use without separate permission. The former 1M profile is the first rollback. See the [GLM xgrammar qualification](../findings/2026-08-31-glm53-xgrammar-524k-qualification.md). The 2026-08-29 Cardillo/Purtell TR3/EXL3 262K/524K qualification remains the historical starting point and GLM-specific rollback evidence. Adaptive MTP remains rejected and the unserved 0xSero 3.0-bpw release remains watch-only. On 2026-08-28, exact digest- and revision-pinned ComfyUI v0.33.4 workflows qualified a 512×512 FLUX.2 Klein PNG and a 17-frame, 512×288 Wan2.2 H.264 MP4 on one RTX 5090. Peak GPU memory was 12,919 MiB and 18,263 MiB respectively from a 943 MiB worker baseline. The worker was removed afterward; no workflow, route, or deployment was promoted. See the [media qualification](../findings/2026-08-28-comfyui-media-qualification.md). A later exact-build live pass exercised cold approval, real Hermes MCP image and video jobs, A2A replay, authenticated artifact delivery, cold-backend failure, and managed teardown. The image samples passed bounded visual review; the video sample failed spatial/prompt adherence, so neither workflow was promoted. See the [live validation](../findings/2026-08-28-media-gateway-live-validation.md). A subsequent production gate enabled only the exact FLUX.2 image workflow. Real Hermes completed `draft` 512×512, `standard` 768×768, and `high` 1024×1024 warm requests plus five cold approval/resume requests. Six of eight images passed strict independent review; two retained draft failures expose origami/material and exact-count limits. The gateway exposes latency phases, requires exact same-job resume, rejects arbitrary dimension overrides, and leaves the worker absent after managed teardown. The final `b46f6ce` regression copied a server-issued exact resume bundle unchanged, reattached to the same job, and returned native image bytes matching the authenticated resource. The image workflow is `available=true` and `promoted=false`. Wan2.2 remains `quality_failed`, unavailable, and has no fallback. See the [production-enablement finding](../findings/2026-08-28-hermes-image-quality-production.md). On 2026-08-26, after the explicit human gate, the exact RadixArk Qwen3.8 Flash Next NVFP4 revision became the text Primary at TP=2/262K/c1 and was then fixed forward to the hash-gated PR #36556 SM120 QSA fast path with matched MTP3 `3/1/4`. It subsequently expanded in place to image/OCR/video after direct media 30/30, live routed repeats 57/60 strict, edge cases 8/8, a six-size context curve, and fresh OpenClaw/Hermes/Pi vision acceptance. It measured 155.9 tok/s at 4K, 114.7 at the new 128K target sweep, and 112.9 at the 254K target; the separate full-reserve gate retained 102.0 tok/s. See the [Qwen3.8 Flash Next dossier](models/qwen38-flash-next.md) and [vision promotion record](../findings/2026-08-26-qwen38-flash-next-vision-promotion.md). On 2026-08-16, after the explicit human gate, the exact digest-pinned r15 TP=2/393K profile became the text Primary at that date. A matched control measured K5 at 150.0 versus 76.4 tok/s median decode at 4K/c1 and 119.245 versus 76.767 at 32K/c1. Direct retrieval passed at 351,118 actual tokens and authenticated routed retrieval at 340,119; tools, streaming, Responses, c8, c2 long-context, and repeated agentic checks passed. The fixed-port r33 393K profile is the transactional rollback. See the [DeepSeek dossier](models/deepseek-v4-flash.md) and [r15 promotion record](../findings/2026-08-16-deepseek-v4-flash-0731-infernal-r15-393k-promotion.md). The upstream r15 recipe was authored by Martin Vit (`voipmonitor`) and was qualified upstream at 131,072 tokens on native Linux with two RTX PRO 6000 Blackwell GPUs on direct PCIe root ports. The local 393,216-token WSL2 result is a separate qualification, not a transferred upstream claim. On 2026-08-15, after a separate human gate, the exact official-FP8 SGLang TP=1/393K Qwen profile became current. Its guarded acceptance passed 108K retrieval, tools 20/20, direct and routed media 18/18, the supported Responses subset, and fresh Hermes/OpenClaw Primary turns without fallback. It is now a former deployment with retained evidence. See the [Qwen3.8 dossier](models/qwen38-27b.md) and [single-service promotion record](../findings/2026-08-15-qwen38-27b-sglang-fp8-single-promotion.md). On 2026-08-16, the then-current Qwen model passed 30/30 direct deterministic media attempts, including video 14/14. A managed router-only expansion added `vision.video` and fail-closed admission for one video; the live admitted subset passed 28/28 along with streaming, tool-use, malformed-input, overflow, and Primary regression probes. See the [video-router finding](../findings/2026-08-16-qwen38-27b-video-router.md). On 2026-08-17, the separate single-RTX-5090 RadixArk NVFP4 profile advanced from its 64K baseline to a 131,072-token served window. It returned a retrieval marker at 119,675 actual prompt tokens, passed tools 20/20, direct image/OCR/video, the complete media corpus 30/30, and boundary cases 4/4 at eight images and two videos. It is the preferred 5090 computer-use perception challenger, but remains direct-only and `no-promotion`; GUI action-loop, routed admission, concurrency, and FP8-KV scale follow-up remain open. See the [128K qualification](../findings/2026-08-17-qwen38-27b-radixark-nvfp4-rtx5090-128k.md). The 2026-08-14 matched vLLM split remains retained historical evidence. Its routed FP8 tools 20/20, BF16 media 30/30, and one 32-image request remain historical capability evidence, not a current admission declaration. On 2026-08-15, official-FP8 MTP=4 and MTP=5 both passed functional, near-393K, and repeated deterministic-quality gates. A cross-card swap showed that the apparent first-run speed gap followed the GPU lane: MTP=5 exceeded MTP=4 decode by only 0.4-1.3% on the same card and did not improve E2E. MTP=3 remained selected within the Qwen profile. See the [MTP-depth qualification](../findings/2026-08-15-qwen38-27b-mtp-depth-qualification.md). Also on 2026-08-15, a digest-pinned SGLang no-speculation A/B qualified official FP8 and audited Inferact NVFP4 as TP=1/393K text controls on both card placements. NVFP4 reduced matched 4K TTFT 22.6% and raised decode 20.6%, while both candidates passed 388,979 actual prompt tokens and bounded deterministic quality. See the [SGLang/NVFP4 qualification](../findings/2026-08-15-qwen38-27b-sglang-nvfp4-qualification.md). A matched MTP=3 follow-up then raised official-FP8 SGLang decode from 48.0 to 111.3 tok/s and Inferact NVFP4 from 57.9 to 98.1. Both retained the complete functional and repeated deterministic-quality gate and passed a 389K retrieval probe. The prior multimodal crash was isolated to SGLang's automatic CUDA-IPC feature transport in this exact runtime; forcing CPU transport let both checkpoints pass bounded image understanding and OCR with MTP enabled. The later matched consolidation corpus added two-image ordering and supported the human promotion of official FP8; Inferact NVFP4 remains `no-promotion`. The then-promoted official-FP8 service subsequently qualified one video. The 32-image ceiling, concurrency above one, and host-memory pressure remain open. See the [MTP/multimodal qualification](../findings/2026-08-15-qwen38-27b-sglang-mtp-multimodal-qualification.md).

How to read the evidence

Evidence labels describe what was observed:

Label Meaning
external-prior Official or community research used to choose a recipe; not local proof.
compatibility-only Loaded or answered bounded probes; no qualification claim.
functional Independent behavior gates passed for the stated contract.
capacity Context, concurrency, throughput, latency, or residency was measured.
quality A declared quality workload was measured with retained results.
historical-invalid Retained run is incomplete, incomparable, or missing identity needed for reuse.

Decision labels are separate: current, rollback, challenger, no-promotion, and rejected. In this public documentation, they describe the latest published evidence decision for a reference configuration, not live deployment state. A quality result can still be no-promotion; a failed load can be rejected while remaining useful compatibility evidence.

Full comparison rules, instrument definitions, and artifact requirements: Methodology and evidence rules.

Complete history

  • Run catalog — every retained, decision-relevant run, indexed by hardware and date.
  • Findings index — the dated reports with full commands, raw artifacts, and failure cases.
  • Finding publication format — compact result and copy formats without weakening evidence boundaries.
  • External benchmarks — advisory priors imported from official and community sources, kept separate from local results.
  • Chronological campaign archive — the stable-URL summary of historical rounds, including Fast-tier and voice campaigns.