Skip to content

Benchmark run catalog

Eight-image follow-up, 2026-09-18: GLM EXL3 r9 passed four eight-image comparison requests, one eight-image high-resolution request. Limit now eight/request; no concurrent eight-image soak.

2026-09-19 APC promotion

Date Capability Configuration Measured GPU Evidence Decision Links
2026-09-19 Huihui Qwen3.8 final direct-I/O 160K context envelope Pinned Huihui artifact, NInfer 70434721, baked image 22fa89d5, binary 82ec9634; MTP3/INT8 KV, 163,840 context/C1, 8,192 visible-output cap RTX 5090 startup; direct preflight 8/8 including 144,425-token needle and 147,941-token tool; repeated exact retrieval 9/9 at 150,058/150,144/150,124 prompt tokens qualified 160K/C1 reference; 64K remains a historical comparison; retrieval latency mixes cold/partial-cache and warm attempts Qwen dossier / context-envelope finding
2026-09-19 GLM-5.3-Flash r10 APC Same checkpoint/v84 image; EXL3 4-bpw, FP8 MLA KV, 327,680 configured tokens, C4, TP2/EP2/DCP2, no speculation 2x RTX PRO 6000 Blackwell Max-Q Four matched 32-request finalist cells passed; shared mean visible TTFT −67.84%, request throughput 3.17×; unique request throughput −2.39%; cache reuse 75.58%; final diagnostics 12/12 per arm; vision/context gates passed Explicit human approval and fresh bounded acceptance; exact r9 remains the rollback Dossier · Promotion follow-up · Historical campaign

2026-09-18 vision acceptance

Date Capability Configuration Measured GPU Evidence Decision Links
2026-09-18 Eight-image comparison and OCR GLM EXL3 r9 vision8, FP8 KV, no speculation, configured 327K/C4 2x RTX PRO 6000 Blackwell Max-Q Four ordinary comparisons, one high-resolution request, six direct preflight checks; matching Pi output has no correlated request provenance User-authorized eight-image limit; r8 one-image rollback; no concurrency soak Dossier · Finding
2026-09-18 Image/OCR and text/tools GLM EXL3 r8 vision, FP8 KV, no speculation, configured 327K/C4 2x RTX PRO 6000 Blackwell Max-Q Functional; 10/10 images; 260K retrieval passed on retry after initial refusal User-authorized vision enablement; prior text profile retained Dossier · Finding

This manually maintained catalog indexes retained, decision-relevant runs. Measured hardware names only the device that executed the workload. PRO relationship prevents a co-resident or protected PRO 6000 from being reported as measured.

Every status and decision in this catalog belongs to its dated evidence. Terms such as current, Primary, rollback, and production describe the published decision at that point in history; they do not report active routes, deployment, placement, or availability.

Research-only recipe intake

These records are indexed for decision reachability but are not benchmark runs and name no measured hardware.

Date Subject Evidence Decision Dossier / finding
2026-08-15 Qwen3.8 27B external recipe refresh external-prior; current official vLLM/SGLang recipes, exact-hardware Reddit MTP sweep, X discovery, and community runtime reports Queue matched official-FP8 MTP=3/4/5 A/B; prepare digest-pinned SGLang official-weight spike; no live or promotion change Qwen3.8 27B · recipe refresh

Cross-run benchmark interpretations

These records synthesize retained measurements without claiming a new run.

Date Subject Evidence Decision Dossier / finding
2026-09-03 GLM-5.3-Flash scheduler concurrency versus KV capacity derivative capacity comparison of the retained 2026-08-29, 2026-08-31, and 2026-09-02 local artifacts; no live request or configuration change C16 is a short-request scheduler ceiling, not sixteen full windows; keep current SGLang at qualified C1, recognize the corrected rollback's measured C2 long-context headroom, and benchmark any higher current-profile admission before changing the published contract GLM-5.3-Flash · concurrency interpretation

Failed startup attempts

Date Configuration Hardware used for startup Outcome / boundary Evidence
2026-09-17 GLM mixed 3.5-bpw EXL3, derived v84, TP2/DCP1, NVFP4 KV, configured 327K/C1 2x RTX PRO 6000 Max-Q, native Linux, 96 GB installed host RAM Host RAM/swap exhausted before readiness; no inference metrics; rejected configuration, no promotion; baseline restored GLM dossier · Finding

RTX PRO 6000 runs

Date Capability / configuration Measured hardware Evidence Decision Dossier / finding
2026-09-19 Huihui Qwen3.8 compatible-runtime baked MTP3 32K 1x RTX 5090, Windows/WSL2 direct lane baked quality gates passed; vision 12/12; capacity 12/12 at 181.46 mean decode tok/s and nominal-31K 6/6 at 174.55 qualified bounded profile; clean recreation and router preflight 6/6 passed; no actual deployment assignment disclosed Qwen3.8 27B · promotion addendum
2026-09-19 Huihui Qwen3.8 compatible-runtime historical comparison set 1x RTX 5090, Windows/WSL2 direct lane four prior follow-up capacity runs plus two baked capacity runs published to historical evidence descriptive historical evidence; failed 128-word warmups retained outside performance rows Qwen3.8 27B · promotion addendum · promotion evidence
2026-09-17 GLM mixed 3.5-bpw EXL3 R7 loader repair, TP2/DCP1, NVFP4 KV, configured 327K/C4 2x RTX PRO 6000 Max-Q, native Linux C1/C4 functional and repeated bounded quality pass; 216,307 actual-token needle; C4 capacity 4/4 then 3/4, baseline 2/4; strict failures retained Runnable experimental candidate; no overall win; baseline restored, no-promotion GLM dossier · Finding
2026-09-14 GLM-5.3-Flash EXL3 4-bpw r7, no speculation; 327,680 total / 65,536 output, C4 2x RTX PRO 6000 Blackwell Max-Q quality, functional, capacity; MMLU-Pro 90/100 one pass, agentic 30/30, SWE 4/5, strict120 120/120, post-promotion context 9/9 selected September 14 text-only lane; Pi/Hermes/OpenClaw tool checks pass; no exhaustive intelligence winner or full-window concurrency claim GLM dossier · Finding
2026-09-11 GLM v0.4.3, fixed ormandj weights, 524K/C4 2x RTX PRO 6000 Blackwell Max-Q functional, capacity, bounded quality; Pi 14/14, retrieval 12/12 user-authorized promotion; 4K output retained GLM dossier · Finding
2026-09-11 GLM rc14 Pi completion budget 4096 vs 16384; same 393K/C1 backend 2x RTX PRO 6000 Blackwell Max-Q bounded functional; 16 valid / 24 tasks, eight coding runs invalid no-promotion; retain 4096 GLM dossier · Finding
Date Capability Exact model/configuration Measured hardware Evidence Decision Dossier / finding
2026-09-19 VoiceChat feasibility stop mlx-community/NemotronLabs-VoiceChat-11B-4bit@dffd203f; MLX runtime 77a6cfca; no candidate load Apple M4 Max, 48 GiB unified memory read-only feasibility; disk arithmetic only, memory unresolved; protected audio diagnostics excluded from candidate latency/quality whole tool-capable replacement blocked by missing tool-result ingress; no-promotion Voice LLM MLX · finding
2026-09-09 ormandj v0.4.2 runtime upgrade qualification GLM-5.3-Flash c3cbb989, candidate image 877ae294, TP2/393,216/C1, P2P enabled 2× RTX PRO 6000 Blackwell Max-Q functional/thinking/multimodal/high-context pass; 4K n12 and 120K n3 matched cells; strict turnover 58/60 and repeat 57/60 user-selected retain-baseline/no-promotion; exact baseline restored GLM-5.3-Flash · finding
2026-09-09 Native NCCL P2P transport A/B GLM-5.3-Flash ormandj W4A16/NVFP4 c3cbb989, exact rc14 image 0c063795, TP2/393,216/C1; only P2P disable 1 to 0 changed 2× RTX PRO 6000 Blackwell Max-Q direct/routed gates and 380K correctness pass; 4K n12 and 120K n3 matched medians; 380K strict capacity and diagnostic marker failure retained user-authorized native P2P default; bounded transport decision GLM-5.3-Flash · finding
2026-09-08 Native Linux / retained WSL comparison GLM-5.3-Flash ormandj W4A16/NVFP4 c3cbb989, exact rc14 image 0c063795, adaptive EAGLE, FP8 KV, TP2/393,216/C1; native adds disabled NCCL P2P 2× RTX PRO 6000 Blackwell Max-Q functional, capacity, bounded quality; historical-style n3 decode +20.5–33.0%, endurance 60/60; extended context 128/150 through 376,484 actual tokens (9 empty, 13 incorrect); strict output 0/3 and unique canaries 1/10 excluded no-promotion; whole-stack migration comparison, no strict finalist qualification; routed context/reasoning limits retained GLM-5.3-Flash · finding
2026-09-04 Qwen3.8 27B comprehensive SGLang/vLLM optimization and topology campaign Inferact 6128240e, RadixArk 319f741c, kelnei 29099dc7, incoai DFlash2 dedf8df6; SGLang digest 616a3e97 / source 5f55db35 and vLLM 0.27.1 digest c2f3b1b9; 262,144 context, FP8 KV, K/chunk/compile/Mamba/target/runtime A/Bs 2x RTX PRO 6000 Max-Q under Docker Desktop/WSL2; single TP1, one TP2 service, and two independent TP1 replicas measured functional, capacity, matched performance, correctness rejection, exact restoration; sustained-output N100: TP1 764.3, TP2 587.9, DP2 1,401.8–1,423.4 aggregate tok/s; all 100/100 canaries; full mean/p50/p95/p99 TTFT, prefill, decode, TPOT/ITL, E2E; RadixArk 746.7 TTFT tradeoff; kelnei MTP2 503.4 versus 315.2 no-spec; 82K/C8 non-interactive DP2 bounded aggregate-throughput winner; TP2 rejected after 23.1% regression and repeated strict-JSON corruption; RadixArk and kelnei retained tradeoffs; no-promotion; exact GLM TP2 service and authenticated route restored Qwen3.8 27B · possibility campaign
2026-09-02 GLM-5.3-Flash SWE-bench Verified smoke ormandj/GLM-5.3-Flash-W4A16-NVFP4-K32-Experts-FP8-WO@c3cbb989, SGLang 4c2c169b, rc.14 image digest 0c063795, TP=2, FP8 KV, adaptive EAGLE [3,5], 393,216/C1, 4,096 output; thinking enabled; isolated macOS arm64 worker 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 bounded quality and benchmark-infrastructure qualification; fixed django__django-11099 attempted 1/1, officially graded 1/1, resolved 1/1; 11 routed requests; agent 34.216 s and grader 29.076 s; normalized token totals unavailable adds one-instance repository-agent evidence without changing benchmark guidance; not a full-suite score GLM-5.3-Flash · SWE smoke
2026-09-02 GLM-5.3-Flash SGLang SM120 393K promotion and client acceptance ormandj/GLM-5.3-Flash-W4A16-NVFP4-K32-Experts-FP8-WO@c3cbb989, rc.14 image digest 0c063795, TP=2, FP8 KV, adaptive EAGLE [3,5], 393,216/C1, 4,096 output, image/OCR 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 prior direct functional/capacity/quality/media/endurance qualification; model-only reserve waiver; managed direct and routed promotion gates; real Pi/OpenClaw/Hermes acceptance; 2,543 MiB/card after client workload; exact 524K rollback retained human-approved published current text/tools/image/OCR profile; Mid Mod Pi unchanged pending host-local router credential GLM-5.3-Flash · promotion
2026-09-02 GLM-5.3-Flash SGLang SM120 W4A16, adaptive-MTP A/B, context-envelope qualification, image/OCR, and endurance ormandj/GLM-5.3-Flash-W4A16-NVFP4-K32-Experts-FP8-WO@c3cbb989, rc.14 image digest 0c063795, SGLang 4c2c169b, TP=2, FP8 KV, adaptive EAGLE [3,5]; selected 245,760/max-running1, with matched 131K no-spec and policy-infeasible 393K/499,712 controls 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, matched performance, bounded quality, multimodal, endurance, reserve, and restoration; tools 20/20, coding 15/15, image/OCR 12/12, endurance 60/60; selected profile 108.57/93.35/95.00 tok/s decode at nominal 4K/120K/230K and 3,487 MiB post-workload free per card 245,760-token/C1 profile locally verified challenger, no-promotion; 393,216 and 499,712 profiles policy-infeasible; incumbent exactly restored GLM-5.3-Flash · SGLang qualification
2026-08-31 GLM-5.3-Flash xgrammar fix-forward, matched speculation A/B, 524K context, C2, image/OCR, rollback, and real clients wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3-v1@319d66a8, incoai/GLM-5.3-Flash-DFlash2@dc77ff1c, corrected runtime digest 4909e318, TP=2/DCP=2, EXL3 K3 target, FP8 DS-MLA KV, DFlash2 fixed K5, 524,288/maxseq16/batch2,048, router c16, up to 16 images 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, matched performance, bounded quality, rollback, routed/client acceptance; both arms 28/28 functional; 83.08 tok/s at 4K and pooled 69.99 at 240K; C2 nominal 250K 2/2; 2,493,817 KV tokens; real OpenClaw/Hermes/Pi pass corrected 524K DFlash2 K5 selected as current; former 1M profile first rollback; text/image/OCR, no video; DFlash2 noncommercial boundary GLM-5.3-Flash · xgrammar qualification
2026-08-30 GLM-5.3-Flash text/tools/image/OCR, 1M context, and K3/K5/chunk optimization wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3-v1@319d66a8, incoai/GLM-5.3-Flash-DFlash2@dc77ff1c, digest-pinned Purtell runtime 001a45bd, TP=2/DCP=2, EXL3 K3 target, FP8 DS-MLA KV, DFlash2 K5, 1,048,576/maxseq16/batch2,048, c16, up to 16 images 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, matched performance, bounded quality, routed/client acceptance; tools 20/20, image/OCR 12/12, quality 12/12, exact 950K-target retrieval, 82.1/67.4/67.9 tok/s at 4K/131K/240K; real Hermes/Pi/OpenClaw pass historical verified; first same-model rollback after the 2026-08-31 fix-forward; K3 alternate and batch4,096 rejection retained; text/image/OCR, no video; DFlash2 noncommercial boundary GLM-5.3-Flash · optimization and promotion
2026-08-29 GLM-5.3-Flash text/tools/image/OCR, speculation, and 524K context qualification brandonmusic/GLM-5.3-Flash-tr3-4bpw@5ab363a8, digest-pinned Purtell EXL3/B12x runtime da5cec95, TP=2/DCP=2, NVFP4 DS-MLA KV; vision/text fixed K5, adaptive K1-K5+ReplaySSM, and no-spec controls at 262,144/524,288 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, matched performance, bounded quality; vision/OCR, tools 20/20, high-reasoning coding 15/15, 250K-target / 206,296-actual retrieval, and 72.8/55.7 tok/s at 4K/128K; exact 495,045-token retrieval and 497,976-token tool use on 524K text profiles; adaptive tools 12/20 vision fixed K5 interactive and no-spec 524K maximum-context challenger profiles, no-promotion; adaptive MTP rejected; 0xSero 3.0-bpw watch-only; direct vision candidate retained for hands-on GLM-5.3-Flash · qualification
2026-08-26 Primary image/OCR/video expansion and comprehensive context sweep RadixArk Qwen3.8 Flash Next NVFP4 7b719225, digest-pinned SGLang 59f06adc, engine d91c3682, hash-gated PR #36556 QSA, TP=2, 262,144 context, c1, BF16 KV, NEXTN 3/1/4, four images or one video 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, multimodal repeatability, routed/client acceptance; direct 30/30, isolated 27/30 then 30/30, live 29/30 then 28/30; edges 8/8; 25/25 context requests through 245,000 actual prompt tokens; 155.9/114.7/112.9 median decode tok/s at 4K/128K/254K targets; 516,032 server tokens human-authorized current text/image/OCR/video Primary; 262K/c1, four-image/one-video, thinking-disabled contract Qwen3.8 Flash Next · vision promotion
2026-08-26 Text Primary QSA-fast/MTP3 fix-forward promotion RadixArk Qwen3.8 Flash Next NVFP4 7b719225, digest-pinned SGLang 59f06adc, engine d91c3682, hash-gated PR #36556 QSA fast path, TP=2, 262,144 context, c1, BF16 KV, NEXTN 3/1/4 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, matched performance, bounded quality, routed identity and fresh real-client acceptance; tools 20/20 and quality 12/12; MTP3 154.9 tok/s at 4K, 134.1 at 128K, 102.0 at 253,703 actual prompt tokens with 8,192 output request; 2.33x/1.93x matched no-spec decode human-authorized current text Primary; 262K/c1 thinking-disabled contract; multimodal unpromoted Qwen3.8 Flash Next · QSA-fast promotion
2026-08-26 Initial text Primary qualification and promotion RadixArk Qwen3.8 Flash Next NVFP4 7b719225, digest-pinned SGLang 59f06adc, engine d91c3682, TP=2, 262,144 context, c1, no speculation, portable QSA sparse decode plus NCCL logits fallback 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, bounded quality, routed and real-client acceptance; 253,325-token direct/routed retrieval with 8,192 reserve, intelligence 6/6, session 3/3, tools 20/20 plus repeated tools 3/3, Responses, OpenClaw/Hermes/Pi pass; 4K/c1 214.612 ms TTFT and 12.801 tok/s decode superseded same day by the QSA-fast/MTP3 fix-forward promotion; retained correctness baseline Qwen3.8 Flash Next · initial promotion
2026-08-21 DeepSeek 0731 Infernal Invocation r18 1M promotion Official 9e165c30, digest-pinned r18 414ec7d0, B12X W4A8/FP8 compressed MLA KV, TP=2/DCP=1, 1,048,576 context, maxseq8, batch4,096, fixed probabilistic DSpark K5; matched no-spec control 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, matched performance, bounded quality, routed and real-client acceptance; 1,040,063-token retrieval, tools 160/160 plus functional batches, agentic 12/12, c8 short and c2 at 490,861 prompt tokens/request; K5/no-spec decode 142.1/76.4 tok/s at 4K and 129.5/76.3 at 32K; Hermes/Pi/OpenClaw pass former human-approved text Primary as of 2026-08-26; retained evidence DeepSeek 0731 · promotion
2026-08-16 DeepSeek 0731 Infernal Invocation r15 393K promotion Official 9e165c30, digest-pinned r15 f1b13c86, B12X W4A8/FP8 compressed MLA KV, TP=2/DCP=1, 393,216 context, maxseq8, batch4,096, fixed probabilistic DSpark K5; matched target-only control 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 functional, capacity, matched performance, bounded quality, routed acceptance; K5/no-spec 4K decode 150.0/76.4 tok/s, direct 351,118-token and routed 340,119-token retrieval, repeated quality 12/12, routed tools and OpenClaw-compatible wire paths pass former human-approved text Primary; r33 393K fixed-port managed rollback at that date; actual Mini OpenClaw turn remained unproven DeepSeek 0731 · promotion
2026-08-16 Qwen3.8 27B SGLang video and router admission Official FP8 017b9c7a, digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 E4M3 KV, EAGLE 3/1/4, c1, CPU feature transport 1x RTX PRO 6000 Max-Q active; second equal card dormant functional, deterministic multimodal, routed acceptance; direct 30/30 with video 14/14, live admitted 28/28, overflow/malformed/SSE/tool and Primary regression pass expand current in place with vision.video, two images/one video fail-closed admission; no model restart Qwen3.8 27B · video-router finding
2026-08-15 Remote agentic and SWE-bench Verified scout Official FP8 017b9c7a, digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 E4M3 KV, EAGLE 3/1/4, c1; router-only Apple evaluation worker 1x RTX PRO 6000 Max-Q active; second equal card idle bounded quality; agentic smoke 2/2, scout 16/18, fixed SWE Verified sample 5/5 official-grader resolved retain current; debug-loop follow-up and larger stratified SWE sample recommended; no route or promotion change Qwen3.8 27B · agentic/SWE scout
2026-08-15 Qwen3.8 27B SGLang official-FP8 single-service promotion Official FP8 017b9c7a, digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 E4M3 KV, max-running1, 2K chunks, FlashInfer, EAGLE 3/1/4, five GDN states, CPU feature transport 1x RTX PRO 6000 Max-Q active; second equal card empty functional, capacity, deterministic multimodal, routed and real-client acceptance; 108K retrieval, tools 20/20, direct and routed media 18/18, Responses, Hermes and OpenClaw pass human-approved current; Primary/general-vision/OCR consolidated on one service; two-image/no-video admission Qwen3.8 27B · promotion
2026-08-15 Qwen3.8 27B SGLang single-service consolidation A/B Official BF16 1d4bf0f, official FP8 017b9c7a, and Inferact NVFP4 6128240e; digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 KV, max-running1, 2K chunks, FlashInfer, EAGLE 3/1/4, CPU feature transport 2x RTX PRO 6000 Max-Q; BF16 fixed on one card, quantized candidates sequential on the other functional, capacity, deterministic multimodal; all three pass 18/18 across six image tasks including a two-image comparison; BF16/FP8/NVFP4 media p50 0.915/0.588/0.448 s and 4K decode 62.7/111.4/97.7 tok/s Official FP8 preferred single-service challenger; NVFP4 lowest-latency third-party alternative; no-promotion; exact current vLLM split restored and routed Qwen3.8 27B · consolidation A/B
2026-08-15 Qwen3.8 27B SGLang MTP=3 and CPU-transport multimodal A/B Official FP8 017b9c7a and Inferact NVFP4 6128240e; digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 KV, max-running1, 2K chunks, FlashInfer, EAGLE 3/1/4, five GDN states; multimodal arms force CPU feature transport 2x RTX PRO 6000 Max-Q, one independent candidate per equal card followed by a swap functional, capacity, bounded deterministic quality, bounded multimodal; official FP8 MTP averages 111.3 decode tok/s and NVFP4 98.1; both pass 389K retrieval, intelligence 6/6, session 3/3, tools 3/3, image, and OCR SGLang MTP/CPU-multimodal challenger, no-promotion; official FP8 wins decode, NVFP4 wins TTFT/prefill; exact current vLLM split restored and readmitted Qwen3.8 27B · MTP/multimodal qualification
2026-08-15 Qwen3.8 27B SGLang official-FP8/NVFP4 A/B Official FP8 017b9c7a versus Inferact ModelOpt NVFP4 6128240e; digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 KV, max-running1, 2K chunks, FlashInfer, no speculation; models swapped across cards 2x RTX PRO 6000 Max-Q, one independent candidate per equal card functional, capacity, bounded deterministic quality; both placements pass full gate, tools 20/20, repeated intelligence/session/tools, and 388,979-token retrieval; NVFP4 averages 0.429 s TTFT, 8,409 prefill tok/s, and 57.9 decode tok/s versus 0.554 s, 6,512, and 48.0 NVFP4 text challenger, both no-promotion; SGLang multimodal unqualified on WSL2; exact current vLLM split restored and readmitted Qwen3.8 27B · SGLang/NVFP4 qualification
2026-08-15 Qwen3.8 27B official-FP8 MTP-depth A/B Official FP8 017b9c7a, vLLM 3a091411, TP=1, 393,216 context, FP8 KV, maxseq1, batch 4,096, MTP=4 versus MTP=5; settings swapped across cards 2x RTX PRO 6000 Max-Q, one independent candidate per equal card functional, capacity, bounded deterministic quality; both pass tools 20/20, repeated intelligence/session/tools, and 388,979-token retrieval; same-card MTP=5 decode only 0.4-1.3% above MTP=4 with no E2E win MTP=4/5 no-promotion; retain current MTP=3; exact FP8/BF16 split restored and readmitted Qwen3.8 27B · MTP-depth qualification
2026-08-14 Qwen3.8 27B split promotion Official FP8 017b9c7a text plus official BF16 1d4bf0f2 multimodal, vLLM 3a091411, TP=1 per card, 393,216 context, FP8 KV, maxseq1, batch 4,096, MTP=3; BF16 32 images/request 2x RTX PRO 6000 Max-Q, one independent serve per equal card functional, capacity, multimodal, client acceptance; FP8 93.6 tok/s at 4K, routed tools 20/20; BF16 routed media 30/30 and 32-image request 1/1; Hermes text/image and OpenClaw Primary/vision passed without fallback human-approved current split; FP8 text Primary, BF16 explicit general-vision/OCR; TP=2 long-context profiles remain experimental Qwen3.8 27B · promotion
2026-08-14 Qwen3.8 27B TP/MTP/long-context matrix Official BF16 1d4bf0f2 and official FP8 017b9c7a, vLLM 3a091411; split TP=1 at 393K and exclusive TP=2 at 393K/600K/1.01M; matched no-MTP/MTP=3, FP8 KV, maxseq1, batch 4,096 2x RTX PRO 6000 Max-Q over PCIe without P2P or NVLink functional, capacity; all 16 arms passed complete functional gates and cold retrieval at 388,979/598,729/985,107 actual prompt tokens; 4K FP8 MTP 85.9-93.6 tok/s; TP=2 control cut 393K TTFT 35-38% challenger, no-promotion; TP=2 for prefill/capacity, MTP for decode, 600K/1M batch-like; exact split baselines restored Qwen3.8 27B · matrix
2026-08-14 Qwen3.8 27B official FP8 1M-context continuation Official FP8 017b9c7a, vLLM 3a091411, 1,010,000 configured context, TP=1, FP8 KV, maxseq1, batch 4,096, text-only, no MTP/prefix cache 1x RTX PRO 6000 Max-Q; second equal card continued the BF16 serve functional, capacity; retrieval passed at 316,849, 422,449, and 633,649 actual prompt tokens, then 825,049 at 3/3 with 956.739 s mean E2E; post-stress and restored-control gates passed stable offline/batch challenger, no-promotion; original 262K FP8 lane restored, no route change Qwen3.8 27B · 1M continuation
2026-08-14 Qwen3.8 27B official qualification and setting A/Bs Official BF16 1d4bf0f2 multimodal and official FP8 017b9c7a text, vLLM 3a091411, 262,144 context, TP=1 per split lane; FP8 MTP=3, prefix-cache, and unquantized-KV one-variable arms 2x RTX PRO 6000 Max-Q, one independent card per co-resident serve functional, capacity, bounded quality, multimodal; both baselines through 241,250 actual prompt tokens, FP8 c5 30/30, BF16 media 30/30; MTP c1 94.8 tok/s, warm 30K-prefix TTFT 0.41 s challenger, no-promotion; direct baselines healthy, no route change; durable worker agentic/SWE open Qwen3.8 27B · qualification
2026-08-11 DeepSeek 0731 r33 393K Primary promotion Official revision 9e165c30, digest-pinned r33 B12X mixed NVFP4-MoE/FP8 path, FP8 DS-MLA KV, DSpark K5, TP=2, 393,216 context, maxseq16, batch 4,096, GPU-only 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 functional, capacity, bounded quality; direct 359,900-token pass, routed functional pass except legacy long-needle calibration and trivial-prompt reasoning-evidence checks; OpenClaw and Hermes 393K/32K/high client-path passes human-approved current; routed/OpenClaw/Hermes >300K and SWE score remain open DeepSeek 0731 · promotion
2026-08-10 DeepSeek 0731 r33 batch-token A/B Official revision 9e165c30, digest-pinned r33 B12X mixed NVFP4-MoE/FP8 path, FP8 DS-MLA KV, target-only/no-spec, TP=2, 131,072 context, c1; only functional change 8,192 to 4,096 batched tokens 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 functional, capacity, bounded performance; fresh bracketed starts, 6/6 preflight, matched 119,503-prompt-token pass; activation 1.73 to 1.14 GiB, minimum-rank KV 15.27 to 15.99 GiB; reported KV tokens 283,917 to 553,243 with accounting caveat 4,096 candidate healthy and direct-only at campaign close; next arm GPU-only 393K; no-promotion DeepSeek 0731 · batch-token A/B
2026-08-10 DeepSeek 0731 r33 quality-first control Official revision 9e165c30, digest-pinned r33 B12X W4A8 mixed NVFP4-MoE/FP8 runtime, FP8 DS-MLA KV, target-only/no-spec, TP=2, 131,072 context, c1; 393K FP8-KV plus 16 GiB host-offload recipe translated but not loaded 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 functional, capacity, bounded quality; preflight 6/6, repeated intelligence/session/tools pass, 119,503 actual prompt tokens at 17.445 s TTFT and 73.86 decode tok/s; context-target calibration caveat retained priority challenger, no-promotion; no route change DeepSeek 0731 · r33 control
2026-08-07 DeepSeek 0731 Vision (NVFP4) first load webbrain-one/DeepSeek-V4-Flash-0731-Vision-NVFP4 rev 3a8f168c, digest-pinned SGLang v0.5.16 image, marlin/marlin kernels, TP=2, KV fp8_e4m3, 4,096 context, --mem-fraction-static 0.97 2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive dual-gpu-exclusive compatibility-only, bounded functional (text lane), bounded negative quality (vision lane) no-promotion; failed back to the 650K Primary in the same session DeepSeek 0731 · vision first-load
2026-08-03 Remote context, agentic recovery, and SWE-bench Verified smoke Official DeepSeek 0731 revision 9e165c30, pinned r16 B12X W4A8/FP8 image, FP8 MLA KV, TP=2, DSpark K5, 650K/maxseq16; Apple evaluation worker 2x RTX PRO 6000 Max-Q over PCIe without NVLink functional, bounded quality; 8K context 1/1, tool-error retry protocol pass with final-answer fail, SWE-bench Verified 1/1 official-grader smoke benchmark substrate qualified for scout; no promotion or route decision DeepSeek 0731 · smoke finding
2026-08-02 DeepSeek 0731 Primary promotion Official revision 9e165c30, pinned r16 B12X W4A8/FP8 image, FP8 MLA KV, TP=2, DSpark K5, 650K/maxseq16/batch4096; router output cap 32768 2x RTX PRO 6000 Max-Q over PCIe without NVLink; display on AMD iGPU functional, capacity; Dark Pi, Mini Pi, Mini OpenClaw, oversized-output clamp, exact readiness and exclusive ownership current, human-approved; 1M removed from client-facing Primary after two fatal client-shaped workspace failures DeepSeek 0731 · promotion
2026-08-02 DeepSeek 0731 GPU-only 650K/1M Pi qualification Official revision 9e165c30, pinned r16 B12X W4A8/FP8 image, FP8 MLA KV, TP=2, DSpark K5; 650K/maxseq16/batch4096 and 1M/maxseq1/4/16/batch2048 2x RTX PRO 6000 Max-Q over PCIe without NVLink; display on AMD iGPU functional, capacity; 640K and 985K retrieval, matched 32K c1 speed, three-tool burst, retained maxseq1 fatal workspace failure 650K/maxseq16 preferred everyday Pi experiment; 1M/maxseq16 preferred deep-session experiment; maxseq4 passing alternative; maxseq1 rejected; all no-promotion DeepSeek 0731 · 650K/1M Pi qualification
2026-08-02 DeepSeek 0731 native KV offload and 256K Official revision 9e165c30, derived r16 B12X mmap-unpinned image 331b7925, FP8 MLA KV, TP=2, DSpark K5, 8/16 GiB CPU offload, 262,144 served tokens, c1 2x RTX PRO 6000 Max-Q over PCIe without NVLink functional, capacity; 128K store/replay, cold 125K/192K/250K ladder, 16 GiB 113,408-token CPU reload, live lifecycle regression priority challenger, no-promotion; 128K remains preferred performance lane DeepSeek 0731 · native-offload/256K qualification
2026-08-01 DeepSeek 0731 DSpark and 128K Official revision 9e165c30, pinned r16 B12X W4A8/FP8 image, FP8 KV, TP=2, DSpark K5 and same-image no-spec control, c1 2x RTX PRO 6000 Max-Q over PCIe without NVLink functional, capacity, bounded quality; low/high/max preflight, 128K, 27/27 coding-agent attempts, paired speculative A/B, per-card telemetry priority challenger, no-promotion; DSpark preferred for experiments, 3 GiB reserve failed DeepSeek 0731 · r16 qualification
2026-08-01 Exclusive TP=2 LLM campaign DeepSeek V4 Flash 0731, Inkling Small NVFP4, Qwen3.5 122B NVFP4, Nemotron 3 Super 120B NVFP4, and Laguna S 2.1 NVFP4; exact pinned revisions, c1, no speculative decode 2× RTX PRO 6000 over PCIe without NVLink external-prior, functional, capacity, quality; failed/calibration lanes retained all no-promotion; DeepSeek is the priority intelligence challenger; production aliases unchanged Dossiers · dual-PRO campaign · DeepSeek research
2026-07-29 Primary LLM promotion Agents-A1 official FP8 4d7d5938, FP8 KV, vLLM f25953cc, 262K, c1, thinking off RTX PRO 6000 functional, capacity, quality current; Qwen3.5 becomes immediate rollback Agents-A1 · promotion
2026-07-29 262K multimodal head-to-head Agents-A1 official FP8 4d7d5938, FP8 KV, vLLM f25953cc; Qwen3.5 122B NVFP4 98915d83, BF16 KV, NGC 26.06; thinking off, c1 RTX PRO 6000 functional, capacity, quality Agents-A1 wins bounded comparison; Qwen stays current; no-promotion Agents-A1 · Qwen3.5 · head-to-head
2026-07-28 Multimodal challenger Agents-A1 BF16 addff08f, official FP8 4d7d5938, and ProtoLabs NVFP4 ff24ba5c; vLLM f25953cc, FP8 KV, 131K; isolated router and FP8 MoE A/B RTX PRO 6000 functional, capacity, quality FP8 principal candidate; NVFP4 compact text Pareto; generated tune rejected; no-promotion Agents-A1 · multimodal qualification
2026-07-28 Primary LLM Qwen3.5 122B A10B NVFP4, rev 98915d837c4e7c87ac8296d02e89de19b3207e6d, BF16 KV, 262K, c1 RTX PRO 6000 functional, capacity, quality current Qwen3.5 · qualification
2026-07-27 Challenger LLM Agents-A1, rev addff08f1653ee72765c5cf458fe84556bb34f8e, thinking disabled RTX PRO 6000 functional, capacity, quality challenger, no-promotion Agents-A1 · release sweep
2026-07-27 Heavy control Laguna S 2.1 NVFP4, rev 07614121b31898586430f189d27a25a0be310843 RTX PRO 6000 functional, capacity, quality rollback Laguna · release sweep
2026-07-26 Heavy LLM Laguna S 2.1 NVFP4, pinned nightly, FP8 KV, 262K, thinking disabled RTX PRO 6000 functional, capacity, quality rollback Laguna · qualification
2026-07-18 Heavy LLM GPT-OSS Puzzle 88B, rev 9c0e0746a0d2218b28cc7b2cb3ce4e1a2f50fdb2, MXFP4/FP8 KV, 131K RTX PRO 6000 functional, capacity, quality rollback Puzzle · enablement
2026-07-17 Heavy LLM GPT-OSS Puzzle 88B compatibility and qualification sequence RTX PRO 6000 compatibility-only, functional no-promotion at that point Puzzle · qualification
2026-07-17 Heavy LLM Gemma 4 31B official QAT W4A16, vLLM 0.25.1, FP8 KV RTX PRO 6000 functional, capacity, quality rejected for latency Gemma 4 · optimization
2026-07-16 Heavy LLM Gemma 4 official/Unsloth 12B, 26B, 31B configurations RTX PRO 6000 compatibility-only, capacity, quality no-promotion / rejected Gemma 4 · vLLM 0.25.1 sweep
2026-07-16 Heavy LLM Gemma 4 Unsloth NVFP4 follow-up RTX PRO 6000 capacity, quality no-promotion Gemma 4 · follow-up
2026-07-16 Heavy LLM Gemma 4 chat-template and size bakeoff RTX PRO 6000 compatibility-only, quality no-promotion / rejected Gemma 4 · template bakeoff
2026-07-13 Heavy LLM Qwen3.6 27B container recipe and first characterization RTX PRO 6000 functional, capacity no-promotion Qwen3.6 · recipe
2026-07-12 Heavy LLM ThinkingCap Qwen3.6 27B FP8, rev e48255afd77b403446332be0f595868337b36591 RTX PRO 6000 functional, capacity, quality historical control Qwen3.6 · promotion-era record
2026-07-12 Heavy LLM Qwen3.6 27B community NVFP4+MTP, official FP8, Unsloth NVFP4, ThinkingCap RTX PRO 6000 functional, capacity, quality no-promotion Qwen3.6 · variation bakeoff
2026-07-12 Heavy LLM Qwen3.6 protocol-v2 comparison RTX PRO 6000 quality no-promotion Qwen3.6 · comparison
2026-07-12 Heavy LLM Qwen3.6 baseline RTX PRO 6000 functional, quality no-promotion Qwen3.6 · baseline
2026-07-12 Heavy LLM Qwen3.5 122B MXFP4, rev 345839ea666a70f5035672f7c88afcba6281921f RTX PRO 6000 functional, capacity, quality no-promotion Qwen3.5 · MXFP4 benchmark
2026-07-12 Heavy LLM Nemotron 3 Super 120B and Mistral Small 4 RTX PRO 6000 functional, capacity, quality no-promotion Nemotron Super, Mistral · challengers
2026-07-12 Heavy LLM Nemotron Puzzle 75B recheck RTX PRO 6000 functional, capacity, quality no-promotion Nemotron Puzzle · recheck
2026-07-12 Heavy LLM GPT-OSS 120B deterministic recheck RTX PRO 6000 capacity, quality no-promotion GPT-OSS 120B · recheck
2026-07-12 Heavy LLM Laguna XS 2.1 protocol-v2 attempts RTX PRO 6000 historical-invalid, compatibility-only rejected Laguna · eval v2
2026-07-10–11 Heavy LLM Ornith 35B FP8, MiniMax M2.7 REAP, DeepSeek V4 Flash, Nemotron Puzzle, Qwen3.6 NVFP4+MTP RTX PRO 6000 functional, capacity, quality, historical-invalid no-promotion / rejected Dossiers · Blackwell bakeoff
2026-07-10 Heavy LLM Qwen3.5 122B NVFP4, NGC 26.04, FP8 KV, 131K RTX PRO 6000 functional, capacity, quality no-promotion at that point Qwen3.5 · candidate record

RTX 5090 runs

Date Capability Exact model/configuration Measured hardware PRO relationship Evidence Decision Dossier / finding
2026-09-19 Huihui 64K text/tools/vision and native clients Pinned lyf/Qwen3.8-27B-Huihui-Abliterated-NInfer-NVFP4@18144690, NInfer 70434721, baked digest 44377d56, MTP3/INT8 KV, 65,536 context/C1 RTX 5090 unrelated functional, quality, descriptive canary-free C1 capacity; direct7/7, routed6/6, vision12/12, Hermes9/9; 24,224 MiB historical qualified 64K comparison; no endurance or comparative64K speed claim Qwen dossier / 64K finding
2026-09-12 Efficient fine-tune/pruning screen Signal27B, Swift27B, Qwopus Flash27B, Minitron20B; pinned Q6_K/Q4_0 KV, llama.cpp a298422d,64K/C1; Signal/Swift MTP3 and no-spec controls RTX 5090 unrelated functional, diagnostic quality, failed strict capacity; all tools20/20; strict thinking-on scores9/10,9/10,1/10,3/10 versus incumbent8/10; small and format-sensitive retain exact262K incumbent; all challengers no-promotion; no qualified full-contract upgrade Efficient variants · finding
2026-09-15 Managed image/video bringup and Wan graph repair FLUX.2 high 1024×1024/four-step smoke; Wan v1 bd12b2de and v2 b93170ee, c1 RTX 5090 unrelated functional image/video artifacts; one image review pass; v1 video independent failure; v2 partial recovery no-promotion; Wan v2 available=false, quality unverified FLUX.2 Klein · Wan2.2 · finding
2026-09-03 Qwen3.8 27B quant/speculation bakeoff Gittensor, cdiamond, QUASAR, CometKim, Red Hat, Telperion, and Unsloth Dynamic V3 pinned checkpoints; SGLang 0.5.18, llama.cpp b10548, NInfer 1455676b, and vLLM 0.27.1; fresh matched no-spec/spec arms at advertised 262K or bounded 64K profiles RTX 5090 unrelated functional, capacity, matched performance, bounded quality, compatibility failure, exact restoration; Gittensor target-only 50.9 ms TTFT/244,002 actual prompt; CometKim MTP3 228.0 tok/s but tools 0/3; Unsloth MTP3 137.7 tok/s/tools 20/20 at 64K; cdiamond MTP8 full-context fallback Gittensor preferred direct TTFT challenger, Unsloth preferred clean 64K speculative arm, no-promotion; dedicated zero-reserve exception only after idle baseline; exact GGUF incumbent restored Qwen3.8 27B · comparison · finding
2026-09-03 Direct text/tools performance and long context neroued/Qwen3.8-27B-nvfp4-NInfer@204e3d92, exact .ninfer SHA-256; NInfer e3aeaf8c; no-spec control versus MTP3/lm-head draft; INT8 KV, 252,928 context, c1, thinking disabled RTX 5090 unrelated functional, capacity, matched performance, bounded quality, exact restoration; MTP3 median TTFT 0.430 s, decode 165.9 tok/s, E2E 0.720 s versus 0.421 s/75.3 tok/s/1.085 s control; 201,746-token prompt plus 8,192 output cap passed; shared-prefix tools 17/20; 2,354 MiB free preferred direct text/tools performance challenger, no-promotion; normal 3 GiB reserve, admission, routed/client, broader agentic/SWE, and promotion-grade runtime gates open; exact GGUF incumbent restored Qwen3.8 27B · qualification
2026-08-28 Production Hermes text-to-image profiles and cold lifecycle FLUX.2 Klein 4B FP8 5b4408e5, Qwen3 4B encoder/FLUX.2 VAE 5f526678, ComfyUI v0.33.4, graph 991b63b8, fixed four-step draft 512×512 / standard 768×768 / high 1024×1024, c1; Anvil Serving 0.36.0 fix-forward through b46f6ce; real Hermes MCP RTX 5090 unrelated functional, routed/client acceptance, bounded capacity and quality, production cutover; 6/8 strict independent reviews pass with two retained draft fidelity/count failures; warm gateway E2E 1.352/1.242/1.650 s; exact same-job cold resume/teardown pass; final server-issued bundle regression 908.936 s E2E and 0.087 s generation; native MCP bytes matched the authenticated resource; arbitrary width and cross-boundary auth controls passed exact image workflow available=true, promoted=false; three fixed profiles only; no exact count/material guarantee; video remains quality_failed and unavailable FLUX.2 Klein · production enablement
2026-08-28 Live text-to-image gateway/client acceptance FLUX.2 Klein 4B FP8 5b4408e5, Qwen3 4B encoder/FLUX.2 VAE 5f526678, ComfyUI v0.33.4, graph 991b63b8, 512×512, four steps, c1; exact Anvil merge 19320de6; cold approval plus real Hermes MCP RTX 5090 unrelated functional, routed/client acceptance, bounded quality; two PNGs passed independent prompt-adherence review; artifact full/range/missing/signature/size/hash checks passed; cold/no-job and teardown controls passed unavailable candidate, no-promotion; broader quality and production cutover open FLUX.2 Klein · live validation
2026-08-28 Live text-to-video gateway/client acceptance Wan2.2 TI2V 5B FP16/UMT5 FP8/VAE c4f60d30, ComfyUI v0.33.4 plus VideoHelperSuite 4ee72c06, graph bd12b2de, 512×288, 17 frames, eight steps, 16 fps, c1; exact Anvil merge 19320de6; real Hermes MCP plus A2A replay RTX 5090 unrelated functional, routed/client acceptance, negative bounded quality; 117,738-byte H.264 decoded at 17 frames/1.0625 s and passed artifact controls, but independent contact-sheet review failed spatial/prompt adherence unavailable candidate, no-promotion; exact workflow version blocked on quality Wan2.2 · live validation
2026-08-28 Text-to-image generation FLUX.2 Klein 4B FP8 5b4408e5, Qwen3 4B encoder/FLUX.2 VAE 5f526678, ComfyUI v0.33.4/CUDA 13.0/PyTorch 2.13, graph 991b63b8, 512×512, four steps, c1 RTX 5090 unrelated functional, capacity; decodable PNG, 9.859 s, peak 12,919 MiB from 943 MiB baseline, queue running 1/pending 0 unavailable candidate, no-promotion; perceptual quality and routed/client acceptance open; final worker removed FLUX.2 Klein · qualification
2026-08-28 Text-to-video generation Wan2.2 TI2V 5B FP16/UMT5 FP8/VAE c4f60d30, ComfyUI v0.33.4 plus VideoHelperSuite 4ee72c06, graph bd12b2de, 512×288, 17 frames, eight steps, 16 fps, c1 RTX 5090 unrelated functional, capacity; decodable H.264 MP4, 9.092 s, peak 18,263 MiB from 943 MiB baseline, queue running 1/pending 0 unavailable candidate, no-promotion; perceptual quality, longer clips, and routed/client acceptance open; final worker removed Wan2.2 · qualification
2026-08-22 Routed real-client acceptance follow-up Unsloth Qwen3.8 27B GGUF 4ca72078, Q4_0 weights/KV/projector, Q4_0 MTP3, digest-pinned llama.cpp b10548; 262,144 backend context, c1 RTX 5090 unrelated functional, routed acceptance; OpenClaw and Hermes exact identity/no-fallback plus shell-tool/result continuation pass; wrong-selector fallback negative control detected; route declared stale 131,072-token SGLang/NVFP4 compatibility metadata short routed tools pass; truthful fingerprint and routed 250K capacity fail; no-promotion Qwen3.8 27B · 250K GGUF qualification
2026-08-21 250K Fast-tier text/tools/vision qualification Unsloth Qwen3.8 27B GGUF 4ca72078, Q4_0 weights/KV/projector, no-spec control versus Q4_0 MTP3, digest-pinned llama.cpp b10548; 262,144 context, c1 RTX 5090 unrelated functional, capacity, matched performance, bounded quality; 253,822 actual prompt tokens plus 8,192 reserve, long-tools at 110,875, tools 20/20, agentic 16/18, 101-turn endurance 3/3, images 18/18; MTP decode 104.1 versus 69.1 tok/s FAST-TIER challenger, no-promotion; Q6_K+same-MTP mathematically disqualified; exact pre-test 128K baseline restored Qwen3.8 27B · 250K GGUF qualification
2026-08-21 Recipe research and same-checkpoint MTP3/ReplaySSM A/B RadixArk Qwen3.8 27B NVFP4 554ebba9, no-speculation baseline versus native MTP 3/1/4 plus linear ReplaySSM on digest-pinned SGLang f825d729; TP=1, FP8 E4M3 KV, c1, 131,072 declared context RTX 5090 unrelated functional, capacity, matched performance, current external-source registry; candidate decode +80.5% at 4K and +67.9% at 64K, tools 20/20, but only 70,231 KV tokens and 64K E2E +1.9% slower MTP3/ReplaySSM rejected as 128K replacement; external EXL3/NInfer/vLLM leads remain research-only; exact baseline restored at 105,649 prompt tokens, no-promotion Qwen3.8 27B · recipe research
2026-08-21 DFlash2 compatibility and context-capacity gate RadixArk Qwen3.8 27B NVFP4 554ebba9 plus incoai DFlash2 dedf8df6, digest-pinned SGLang f825d729, TP=1, configured 262,144 context, FP8 E4M3 KV, c1, DFLASH/8 draft tokens; exact float32/extra_buffer_lazy arm plus BF16 single-slot tuning ladder RTX 5090 unrelated functional, capacity; exact arm allocated 24,347 KV tokens and rejected 105,649 prompt tokens; tuned BF16/no-radix/no-prefill-graph arm allocated 70,262 and passed coding, JSON, 49,549-token retrieval, tools 20/20 Both DFlash2 arms rejected as 128K replacements, no-promotion; exact stock 128K candidate restored and passed the complete gate Qwen3.8 27B · DFlash2 qualification
2026-08-20 Chat-template A/B RadixArk Qwen3.8 27B NVFP4 554ebba9, digest-pinned SGLang c4271c3, TP=1, 131,072 context, FP8 E4M3 KV, c1, no MTP; stock versus Sharp v22.1 3dc34df RTX 5090 unrelated functional, bounded diagnostic quality; Sharp preflight pass, thinking-enabled 24/30 for both with Sharp +10.8% completion tokens/+10.7% latency, thinking-disabled behavior stock 18/18 versus Sharp 15/18 with Sharp -5.1% tokens Sharp v22.1 rejected, no-promotion; exact stock 128K candidate restored Qwen3.8 27B · template A/B
2026-08-17 Long-context multimodal boundary RadixArk Qwen3.8 27B NVFP4 554ebba9, digest-pinned SGLang c4271c3, TP=1, 131,072 context, FP8 E4M3 KV, c1, no MTP, CPU feature transport RTX 5090 unrelated functional, bounded deterministic quality; retrieval at 119,675 actual prompt tokens, tools 20/20, media 30/30, boundary 4/4 at eight images / two videos retained direct 128K challenger, no-promotion; 64K recipe retained as rollback Qwen3.8 27B · 128K qualification
2026-08-17 Computer-use perception / vision / video RadixArk Qwen3.8 27B NVFP4 554ebba9, digest-pinned SGLang c4271c3, TP=1, 65,536 context, FP8 E4M3 KV, c1, no MTP, CPU feature transport RTX 5090 unrelated functional, capacity, bounded deterministic quality; 60K retrieval, tools 20/20, direct image/OCR/video pass, corpus 30/30 with video 14/14 and mixed 4/4 qualified 64K baseline; now retained as 128K rollback, no-promotion Qwen3.8 27B · 5090 qualification
2026-07-28 STT Parakeet TDT 0.6B v3 baseline RTX 5090 protected/co-resident quality, capacity current Parakeet · ASR qualification
2026-07-28 STT Qwen3-ASR 0.6B, rev 5eb144179a02acc5e5ba31e748d22b0cf3e303b0 RTX 5090 protected/co-resident quality, capacity challenger, no-promotion Qwen3-ASR · ASR qualification
2026-07-28 STT Nemotron 3.5 ASR, rev f3d333391852ba876df169dcc9ba902d25b6ab0b RTX 5090 protected/co-resident quality, capacity rejected Nemotron ASR · ASR qualification
2026-07-27 Omni/vision Nemotron Nano Omni 30B NVFP4, rev dc5f0b0bfddf8b6e0f5891475be9af05b80126fe RTX 5090 protected/co-resident functional, capacity current topology Nemotron Omni · Omni qualification
2026-07-27 Omni/voice Qwen2.5-Omni 3B, rev f75b40e3da2003cdd6e1829b1f420ca70797c34e; Parakeet; Kokoro RTX 5090 protected/co-resident functional, capacity challenger, no-promotion Qwen Omni, Parakeet, Kokoro · co-resident stack
2026-07-16 Fast LLM Gemma 4 E2B W4A16, E4B, and controls RTX 5090 topology-only quality, capacity; E2B preflights through 120K passed, strict timeout triage 0/3 E4B retained; E2B no-promotion Gemma E2B · Gemma E4B · template bakeoff
2026-07-13 Fast LLM Gemma 4 E4B Fast RTX 5090 topology-only functional historical current Gemma E4B · router promotion
2026-07-10–11 Fast/Omni Nemotron Nano/Omni 30B, Gemma 4 31B failed load, Qwen3.5 35B GGUF, Gemma E4B GGUF RTX 5090 topology-only functional, capacity, quality, historical-invalid no-promotion / rejected Qwen3.5 35B · Gemma 4 variants · RTX 5090 dossiers · Blackwell bakeoff

Apple Silicon runs

Date Capability Exact model/configuration Measured hardware Evidence Decision Dossier / finding
2026-09-19 Swift / stock Qwen3.8-27B artifact feasibility screen Pinned Swift and stock GGUF Q6_K / Q4_K_M candidates plus family-matched F16 projector; conversion runtime llama.cpp b10896 fa676981, managed binary/config unverified Apple M4 Max 40-core GPU, 48 GiB unified memory read-only disk-policy screen only; all four uncached pairs shortfall the retained reserve/allowance before temporary files; memory containment unresolved investigation stop; no download, trial, benchmark, cleanup, route, or promotion; no-promotion Efficient variants · Qwen3.8 27B · finding
2026-09-08 Local voice LLM refresh and TTS runtime update Qwen3.5-9B MLX 4-bit 8b2b98c00a6b4d291155e4890773ca8f769aee53; Qwen3.8-27B 3e6447f082e89cc7f0bc6e5441afd38dfce760ff; Qwen3.6-35B-A3B 38740b847e4cb78f352aba30aa41c76e08e6eb46; MLX-LM 0.31.3 / MLX 0.31.2; Kokoro FastAPI 0.8.2 58b08a915b3463cb76e376a2867e04f9d828f4df Apple M4 Max laptop, 48 GB unified memory 9B functional 6/6; diagnostic quality 9/12 and patchformat 0/3; spoken 33/36 strict; strict 4K/C1 capacity 0/10 with canaries 10/10; Qwen3.8 compatibility-only 5/6 groups, tools 1/3; Qwen3.6 compatibility-only 5/6 strict JSON fail but spoken 36/36 and one accepted SDK session; Kokoro 0.8.2 one-sample WER 0.0/2139.31 ms and final warm Realtime acceptance no-promotion for LLMs; no performance headline or live LLM change; candidates unloaded; Kokoro 0.8.2 and versioned TTS/proxy definitions deployed with four endpoint checks HTTP 200 Voice LLM MLX · finding