Eight-image follow-up, 2026-09-18:GLM EXL3 r9 passed four eight-image comparison requests, one eight-image high-resolution request. Limit now eight/request; no concurrent eight-image soak.
2026-09-19 current decision: the GLM-5.3-Flash r10 APC promotion follow-up records explicit human approval and fresh bounded acceptance. The historical r9/r10 campaign remains the evidence and failure record; exact r9 remains the documented rollback.
Compare local models by useful context, quality, and serving performance.
Start with the latest decision, choose your hardware, then follow the exact
configuration and retained evidence. Evidence reviewed: 2026-09-19 UTC.
Human-approved GLM-5.3-Flash r10 APC · Dual RTX PRO 6000
The bounded r10 APC configuration retains the 4-bpw GLM checkpoint, v84 runtime,
FP8 MLA KV, no speculation, 327,680 configured tokens, C4, and the eight-image
contract. On frozen 32-request finalists with about 25K actual prompt tokens,
it improved the shared-prefix path while staying within the unique-prefix gate.
Matched outcome
r9
r10 APC
Shared-prefix visible TTFT
23.70 s
7.62 s (−67.84%)
Shared-prefix requests/s
0.146
0.462 (3.17×)
Unique-prefix requests/s
0.144
0.141 (−2.39%)
The 75.58% isolated cache-reuse fraction supports the intended repeated-prefix
mechanism. This does not prove four simultaneous maximum-context requests,
video readiness, a soak result, or broad intelligence/SWE parity. Roll back to
the exact r9 recipe if correctness, containment, routed readiness, or the
frozen unique-prefix gate fails.
Limits of this baseline: tested text only; no exhaustive intelligence ranking, prolonged soak,
reboot, or simultaneous full-window proof. These are dated local results,
not a live status display or an automatic model-selection policy.
Start with the measured RTX PRO 6000,
RTX 5090, or Apple M4 Max
view. GPU count, interconnect, operating
system, and runtime are part of the result.
2. Name the workload
For interactive use, prioritize TTFT and E2E. For sustained serving, use
aggregate output, TPOT/ITL, and tail latency. For long context, require actual
prompt depth rather than configured context.
Open the exact recipe, then follow its dated finding and native JSON artifact.
A copied container configuration is not a reproduced result until the same
workload and independent gates pass.
The model comparison table puts the wider measured
configuration — TTFT, throughput, context, reasoning mode, and recipe link —
in one place, grouped by card and workload. The
measured recipe finder is the shorter path when you
want a current configuration, its tradeoffs, and raw evidence.
For the newest single-card Qwen measurements, the dedicated
Qwen3.8 27B RTX 5090 quant comparison
shows every fresh no-speculation/speculation pair, full-context proof, and
failure without compressing the operator's TTFT-first ranking into one score.
The newer dual-PRO Qwen3.8 27B comprehensive campaign
compares SGLang target-only/DFlash2 tuning, TP1/TP2/two-replica DP2,
Inferact/RadixArk targets, and kelnei/vLLM MTP2/no-spec. It retains mean,
p50, p95, and p99 TTFT, effective prefill, decode, TPOT/ITL, and E2E plus
aggregate throughput and correctness failures. Open the
campaign graph (SVG)
or its plotted data and source hashes.
For the compact publication contract, use the
finding format. It defines screenshot-ready result
cards, the common campaign artifact-set template, X and Reddit variants,
accessible alt text, and a claim ledger back to the full finding and raw
artifacts.
The published dual-PRO reference configuration uses two equal PRO 6000 cards.
Aggregate VRAM is not unified memory, and the cards communicate over PCIe
without NVLink. Exclusive TP=2 runs must prove that both devices were selected
and all other inference was offline. The published reference topology keeps
Companion Node model-free. Historical mixed-card tests still say which card was
measured and which was merely protected or co-resident.
The complete hardware-first inventory lists every dossier
and explicit grouped variant, including the September 12 efficient Qwen models
and the Apple MLX Qwen lanes. Each dossier is a stable, model-centered summary
with its dated findings and raw evidence.
2026-09-15 — media: The media finding
records bounded workflow evidence; use its dossier and finding rather than
treating it as a text-model selection result.
2026-09-12 — efficient Qwen variants: Signal, Swift, Qwopus, and Minitron
all passed functional checks on the RTX 5090, but none cleared the exact
incumbent's output and deployment gates. See the grouped dossier.
The context, agentic, and SWE job guide describes the
registered-worker workflow, bounded profiles through 640K, independent scoring,
pinned mini-SWE-agent and official SWE-bench grading, cancellation, and the
measured-versus-prior evidence boundary. The guide defines methodology and
commands; it does not claim a live model result by itself.
The collapsed history below preserves earlier decisions through September 11.
Labels such as current and rollback describe their status at that date;
the September 13–14 selection above supersedes the primary-model guidance.
Use the run catalog to compare dated configurations.
Earlier decisions and controls — historical configurations
1. **GLM-5.3-Flash ormandj W4A16/NVFP4 SGLang** — `current`
text/tools/image/OCR published profile at 393,216 tokens/C1 in exclusive
TP=2 across both RTX PRO 6000 cards, with adaptive EAGLE, explicit thinking
control, a 4,096-token output cap, and real Pi/OpenClaw/Hermes acceptance.
The standing reserve was explicitly waived for the model-only pair while
the post-workload runtime safety gates remained in force.
2. **GLM-5.3-Flash EXL3 K3 plus corrected DFlash2 K5** — immediate exact
rollback at 524,288 tokens/router-c16, measured C2 at a nominal 250K target,
8,192 maximum output, and up to 16 images. Video is unsupported and the
DFlash2 draft is noncommercial without separate permission. The SGLang
245,760/C1 profile remains a conservative verified fallback.
3. **RadixArk Qwen3.8 Flash Next NVFP4** — immediate retained
text/image/OCR/video rollback at 262,144 tokens, exclusive TP=2, router c1,
four-image/one-video admission, and the qualified hash-gated SM120 QSA-fast
MTP3 profile.
4. **Inferact Qwen3.8 27B NVFP4 plus DFlash2 K12/chunk1K on the PRO pair** —
bounded aggregate-throughput winner: two independent TP1 replicas measured
1,401.8–1,423.4 tok/s at C16, versus one TP1 at 764.3 and TP2 at 587.9. All
headline canaries passed 100/100, but TP2 failed strict JSON twice and is
rejected. RadixArk is the lower-TTFT tradeoff; kelnei/vLLM MTP2 beat its
exact no-spec control but not SGLang. Broad quality and a balancing/client
path remain open; `no-promotion`, exact GLM service and route restored.
5. **Gittensor Qwen3.8 27B RTX5090 NVFP4 target-only SGLang** — preferred
measured direct RTX 5090 TTFT challenger: 50.9 ms warm median TTFT, 79.5
tok/s decode, and a 244,002-token actual prompt. The advertised DSpark pair
failed on incompatible shapes and FP8 KV scales need further validation;
`no-promotion`, exact GGUF incumbent restored. Unsloth Dynamic V3.0 MTP3
is the strongest clean bounded-64K speculative arm at 137.7 tok/s with
tools 20/20.
6. **Qwen3.8 27B NInfer NVFP4 MTP3** — retained direct decode challenger at
252,928 tokens/C1; 0.430-second median TTFT and 165.9 tok/s decode. Its
admission, reserve, runtime-image, routed/client, and broader gates remain
open.
7. **DeepSeek V4 Flash 0731 Infernal Invocation r18/r15** — former text
Primary profiles with retained TP=2 long-context, performance, quality, and
real-client evidence.
8. **Qwen3.8 27B official FP8 SGLang single service and FP8/BF16 vLLM split** —
former text/image/OCR/video deployments with retained managed recipes.
9. **Qwen3.5 122B A10B NVFP4**, **Agents-A1 plus Omni**, **DeepSeek r16 650K**,
**Laguna S 2.1**, and **GPT-OSS Puzzle 88B** — retained qualified or
promotion-era recipes, not the immediate selected deployment.
10. **Gemma 4** and **ThinkingCap Qwen3.6 27B** — historical strict-quality
controls.
11. **Nemotron 3.5 ASR** and **Qwen3-ASR 0.6B** — historical RTX 5090
measurements. The RTX PRO 6000 was protected, not benchmarked.
12. **RadixArk Qwen3.8 27B NVFP4** — historical RTX 5090 multimodal challenger;
direct 128K text/tools/image/OCR/video evidence, no route or promotion.
13. **FLUX.2 Klein 4B FP8** — published as available for the RTX 5090 ComfyUI image
workflow through Hermes/MCP at fixed 512/768/1024-pixel profiles; six of
eight bounded visual samples passed strict review, with two retained draft
fidelity/count failures. **Wan2.2 TI2V 5B** remains unavailable after
its decodable sample failed prompt adherence and spatial quality.
On 2026-09-02, the exact ormandj GLM-5.3-Flash W4A16/NVFP4 checkpoint and
rc14 SGLang image qualified and was human-promoted at 393,216 tokens/C1 on
both PRO 6000 cards. The selected adaptive-MTP lane passed tools 20/20, coding
15/15, media 12/12, endurance 60/60, and 4K/120K/262K/380K capacity through
304,491 actual prompt tokens. After a recorded model-only reserve waiver, the
managed direct/routed gates and real Pi/OpenClaw/Hermes acceptance passed;
2,543 MiB/card remained after client work. The exact 524K incumbent is retained
as rollback. See the
[qualification](../findings/2026-09-02-glm53-sglang-sm120-qualification.md)
and [promotion](../findings/2026-09-02-glm53-sglang-sm120-393k-promotion.md).
A 2026-09-03 cross-run review clarifies that the old `maxseq16` setting is a
scheduler ceiling, not proof of sixteen full-context requests. The current
SGLang profile remains qualified at C1 with a 393,216-token shared pool. The
corrected 524K rollback reports 2,493,817 KV tokens and separately proved C2
at 206,630 prompt tokens/request; the historical BrandonMusic C16 artifacts
used 4K prompts. The review records bounded C2/C4 planning math for a future
current-profile experiment without promoting it to measured evidence. See the
[concurrency and KV-capacity interpretation](../findings/2026-09-03-glm53-concurrency-capacity-interpretation.md).
On 2026-08-31, the exact GLM-5.3-Flash K3 target plus DFlash2 K5 draft moved to
a digest-pinned xgrammar-corrected 524K profile after a matched no-speculation
A/B and full rollback/forward drill. It measured 83.08 tok/s at 4K and a pooled
69.99 at 240K, versus 42.61 and 43.63 without speculation. The engine reported
2,493,817 KV tokens, C2 nominal 250K completed 2/2, and direct structured,
quality, image/OCR, exact router, and real OpenClaw/Hermes/Pi gates passed.
Video remains disabled and the DFlash2 license limits the recipe to
evaluation/noncommercial use without separate permission. The former 1M
profile is the first rollback. See the
[GLM xgrammar qualification](../findings/2026-08-31-glm53-xgrammar-524k-qualification.md).
The 2026-08-29 Cardillo/Purtell TR3/EXL3 262K/524K qualification remains the
historical starting point and GLM-specific rollback evidence. Adaptive MTP
remains rejected and the unserved 0xSero 3.0-bpw release remains watch-only.
On 2026-08-28, exact digest- and revision-pinned ComfyUI v0.33.4 workflows
qualified a 512×512 FLUX.2 Klein PNG and a 17-frame, 512×288 Wan2.2 H.264 MP4
on one RTX 5090. Peak GPU memory was 12,919 MiB and 18,263 MiB respectively
from a 943 MiB worker baseline. The worker was removed afterward; no workflow,
route, or deployment was promoted. See the
[media qualification](../findings/2026-08-28-comfyui-media-qualification.md).
A later exact-build live pass exercised cold approval, real Hermes MCP image
and video jobs, A2A replay, authenticated artifact delivery, cold-backend
failure, and managed teardown. The image samples passed bounded visual review;
the video sample failed spatial/prompt adherence, so neither workflow was
promoted. See the
[live validation](../findings/2026-08-28-media-gateway-live-validation.md).
A subsequent production gate enabled only the exact FLUX.2 image workflow.
Real Hermes completed `draft` 512×512, `standard` 768×768, and `high`
1024×1024 warm requests plus five cold approval/resume requests. Six of eight
images passed strict independent review; two retained draft failures expose
origami/material and exact-count limits. The gateway exposes latency phases,
requires exact same-job resume, rejects arbitrary dimension overrides, and
leaves the worker absent after managed teardown. The final `b46f6ce` regression
copied a server-issued exact resume bundle unchanged, reattached to the same
job, and returned native image bytes matching the authenticated resource. The
image workflow is
`available=true` and
`promoted=false`. Wan2.2 remains `quality_failed`, unavailable, and has no
fallback. See the
[production-enablement finding](../findings/2026-08-28-hermes-image-quality-production.md).
On 2026-08-26, after the explicit human gate, the exact RadixArk Qwen3.8 Flash
Next NVFP4 revision became the text Primary at TP=2/262K/c1 and was then fixed
forward to the hash-gated PR #36556 SM120 QSA fast path with matched MTP3
`3/1/4`. It subsequently expanded in place to image/OCR/video after direct
media 30/30, live routed repeats 57/60 strict, edge cases 8/8, a six-size
context curve, and fresh OpenClaw/Hermes/Pi vision acceptance. It measured
155.9 tok/s at 4K, 114.7 at the new 128K target sweep, and 112.9 at the 254K
target; the separate full-reserve gate retained 102.0 tok/s. See the
[Qwen3.8 Flash Next dossier](models/qwen38-flash-next.md) and
[vision promotion record](../findings/2026-08-26-qwen38-flash-next-vision-promotion.md).
On 2026-08-16, after the explicit human gate, the exact digest-pinned r15
TP=2/393K profile became the text Primary at that date. A matched control measured K5 at
150.0 versus 76.4 tok/s median decode at 4K/c1 and 119.245 versus 76.767 at
32K/c1. Direct retrieval passed at 351,118 actual tokens and authenticated
routed retrieval at 340,119; tools, streaming, Responses, c8, c2 long-context,
and repeated agentic checks passed. The fixed-port r33 393K profile is the
transactional rollback. See the [DeepSeek dossier](models/deepseek-v4-flash.md)
and [r15 promotion record](../findings/2026-08-16-deepseek-v4-flash-0731-infernal-r15-393k-promotion.md).
The upstream r15 recipe was authored by Martin Vit (`voipmonitor`) and was
qualified upstream at 131,072 tokens on native Linux with two RTX PRO 6000
Blackwell GPUs on direct PCIe root ports. The local 393,216-token WSL2 result
is a separate qualification, not a transferred upstream claim.
On 2026-08-15, after a separate human gate, the exact official-FP8 SGLang
TP=1/393K Qwen profile became current. Its guarded acceptance passed 108K
retrieval, tools 20/20, direct and routed media 18/18, the supported Responses
subset, and fresh Hermes/OpenClaw Primary turns without fallback. It is now a
former deployment with retained evidence. See the
[Qwen3.8 dossier](models/qwen38-27b.md) and
[single-service promotion record](../findings/2026-08-15-qwen38-27b-sglang-fp8-single-promotion.md).
On 2026-08-16, the then-current Qwen model passed 30/30 direct deterministic
media attempts, including video 14/14. A managed router-only expansion added
`vision.video` and fail-closed admission for one video; the live admitted
subset passed 28/28 along with streaming, tool-use, malformed-input, overflow,
and Primary regression probes. See the
[video-router finding](../findings/2026-08-16-qwen38-27b-video-router.md).
On 2026-08-17, the separate single-RTX-5090 RadixArk NVFP4 profile advanced
from its 64K baseline to a 131,072-token served window. It returned a retrieval
marker at 119,675 actual prompt tokens, passed tools 20/20, direct
image/OCR/video, the complete media corpus 30/30, and boundary cases 4/4 at
eight images and two videos. It is the preferred 5090 computer-use perception
challenger, but remains direct-only and `no-promotion`; GUI action-loop, routed
admission, concurrency, and FP8-KV scale follow-up remain open. See the
[128K qualification](../findings/2026-08-17-qwen38-27b-radixark-nvfp4-rtx5090-128k.md).
The 2026-08-14 matched vLLM split remains retained historical evidence. Its routed FP8
tools 20/20, BF16 media 30/30, and one 32-image request remain historical
capability evidence, not a current admission declaration.
On 2026-08-15, official-FP8 MTP=4 and MTP=5 both passed functional,
near-393K, and repeated deterministic-quality gates. A cross-card swap showed
that the apparent first-run speed gap followed the GPU lane: MTP=5 exceeded
MTP=4 decode by only 0.4-1.3% on the same card and did not improve E2E. MTP=3
remained selected within the Qwen profile. See the
[MTP-depth qualification](../findings/2026-08-15-qwen38-27b-mtp-depth-qualification.md).
Also on 2026-08-15, a digest-pinned SGLang no-speculation A/B qualified
official FP8 and audited Inferact NVFP4 as TP=1/393K text controls on both card
placements. NVFP4 reduced matched 4K TTFT 22.6% and raised decode 20.6%, while
both candidates passed 388,979 actual prompt tokens and bounded deterministic
quality. See the
[SGLang/NVFP4 qualification](../findings/2026-08-15-qwen38-27b-sglang-nvfp4-qualification.md).
A matched MTP=3 follow-up then raised official-FP8 SGLang decode from 48.0 to
111.3 tok/s and Inferact NVFP4 from 57.9 to 98.1. Both retained the complete
functional and repeated deterministic-quality gate and passed a 389K
retrieval probe. The prior multimodal crash was isolated to SGLang's automatic
CUDA-IPC feature transport in this exact runtime; forcing CPU transport let
both checkpoints pass bounded image understanding and OCR with MTP enabled.
The later matched consolidation corpus added two-image ordering and supported
the human promotion of official FP8; Inferact NVFP4 remains `no-promotion`.
The then-promoted official-FP8 service subsequently qualified one video. The
32-image ceiling, concurrency above one, and host-memory pressure remain open.
See the
[MTP/multimodal qualification](../findings/2026-08-15-qwen38-27b-sglang-mtp-multimodal-qualification.md).
Official or community research used to choose a recipe; not local proof.
compatibility-only
Loaded or answered bounded probes; no qualification claim.
functional
Independent behavior gates passed for the stated contract.
capacity
Context, concurrency, throughput, latency, or residency was measured.
quality
A declared quality workload was measured with retained results.
historical-invalid
Retained run is incomplete, incomparable, or missing identity needed for reuse.
Decision labels are separate: current, rollback, challenger,
no-promotion, and rejected. In this public documentation, they describe the
latest published evidence decision for a reference configuration, not live
deployment state. A quality result can still be no-promotion; a failed load
can be rejected while remaining useful compatibility evidence.