| 2026-09-19 |
VoiceChat feasibility stop |
mlx-community/NemotronLabs-VoiceChat-11B-4bit@dffd203f; MLX runtime 77a6cfca; no candidate load |
Apple M4 Max, 48 GiB unified memory |
read-only feasibility; disk arithmetic only, memory unresolved; protected audio diagnostics excluded from candidate latency/quality |
whole tool-capable replacement blocked by missing tool-result ingress; no-promotion |
Voice LLM MLX · finding |
| 2026-09-09 |
ormandj v0.4.2 runtime upgrade qualification |
GLM-5.3-Flash c3cbb989, candidate image 877ae294, TP2/393,216/C1, P2P enabled |
2× RTX PRO 6000 Blackwell Max-Q |
functional/thinking/multimodal/high-context pass; 4K n12 and 120K n3 matched cells; strict turnover 58/60 and repeat 57/60 |
user-selected retain-baseline/no-promotion; exact baseline restored |
GLM-5.3-Flash · finding |
| 2026-09-09 |
Native NCCL P2P transport A/B |
GLM-5.3-Flash ormandj W4A16/NVFP4 c3cbb989, exact rc14 image 0c063795, TP2/393,216/C1; only P2P disable 1 to 0 changed |
2× RTX PRO 6000 Blackwell Max-Q |
direct/routed gates and 380K correctness pass; 4K n12 and 120K n3 matched medians; 380K strict capacity and diagnostic marker failure retained |
user-authorized native P2P default; bounded transport decision |
GLM-5.3-Flash · finding |
| 2026-09-08 |
Native Linux / retained WSL comparison |
GLM-5.3-Flash ormandj W4A16/NVFP4 c3cbb989, exact rc14 image 0c063795, adaptive EAGLE, FP8 KV, TP2/393,216/C1; native adds disabled NCCL P2P |
2× RTX PRO 6000 Blackwell Max-Q |
functional, capacity, bounded quality; historical-style n3 decode +20.5–33.0%, endurance 60/60; extended context 128/150 through 376,484 actual tokens (9 empty, 13 incorrect); strict output 0/3 and unique canaries 1/10 excluded |
no-promotion; whole-stack migration comparison, no strict finalist qualification; routed context/reasoning limits retained |
GLM-5.3-Flash · finding |
| 2026-09-04 |
Qwen3.8 27B comprehensive SGLang/vLLM optimization and topology campaign |
Inferact 6128240e, RadixArk 319f741c, kelnei 29099dc7, incoai DFlash2 dedf8df6; SGLang digest 616a3e97 / source 5f55db35 and vLLM 0.27.1 digest c2f3b1b9; 262,144 context, FP8 KV, K/chunk/compile/Mamba/target/runtime A/Bs |
2x RTX PRO 6000 Max-Q under Docker Desktop/WSL2; single TP1, one TP2 service, and two independent TP1 replicas measured |
functional, capacity, matched performance, correctness rejection, exact restoration; sustained-output N100: TP1 764.3, TP2 587.9, DP2 1,401.8–1,423.4 aggregate tok/s; all 100/100 canaries; full mean/p50/p95/p99 TTFT, prefill, decode, TPOT/ITL, E2E; RadixArk 746.7 TTFT tradeoff; kelnei MTP2 503.4 versus 315.2 no-spec; 82K/C8 non-interactive |
DP2 bounded aggregate-throughput winner; TP2 rejected after 23.1% regression and repeated strict-JSON corruption; RadixArk and kelnei retained tradeoffs; no-promotion; exact GLM TP2 service and authenticated route restored |
Qwen3.8 27B · possibility campaign |
| 2026-09-02 |
GLM-5.3-Flash SWE-bench Verified smoke |
ormandj/GLM-5.3-Flash-W4A16-NVFP4-K32-Experts-FP8-WO@c3cbb989, SGLang 4c2c169b, rc.14 image digest 0c063795, TP=2, FP8 KV, adaptive EAGLE [3,5], 393,216/C1, 4,096 output; thinking enabled; isolated macOS arm64 worker |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
bounded quality and benchmark-infrastructure qualification; fixed django__django-11099 attempted 1/1, officially graded 1/1, resolved 1/1; 11 routed requests; agent 34.216 s and grader 29.076 s; normalized token totals unavailable |
adds one-instance repository-agent evidence without changing benchmark guidance; not a full-suite score |
GLM-5.3-Flash · SWE smoke |
| 2026-09-02 |
GLM-5.3-Flash SGLang SM120 393K promotion and client acceptance |
ormandj/GLM-5.3-Flash-W4A16-NVFP4-K32-Experts-FP8-WO@c3cbb989, rc.14 image digest 0c063795, TP=2, FP8 KV, adaptive EAGLE [3,5], 393,216/C1, 4,096 output, image/OCR |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
prior direct functional/capacity/quality/media/endurance qualification; model-only reserve waiver; managed direct and routed promotion gates; real Pi/OpenClaw/Hermes acceptance; 2,543 MiB/card after client workload; exact 524K rollback retained |
human-approved published current text/tools/image/OCR profile; Mid Mod Pi unchanged pending host-local router credential |
GLM-5.3-Flash · promotion |
| 2026-09-02 |
GLM-5.3-Flash SGLang SM120 W4A16, adaptive-MTP A/B, context-envelope qualification, image/OCR, and endurance |
ormandj/GLM-5.3-Flash-W4A16-NVFP4-K32-Experts-FP8-WO@c3cbb989, rc.14 image digest 0c063795, SGLang 4c2c169b, TP=2, FP8 KV, adaptive EAGLE [3,5]; selected 245,760/max-running1, with matched 131K no-spec and policy-infeasible 393K/499,712 controls |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, matched performance, bounded quality, multimodal, endurance, reserve, and restoration; tools 20/20, coding 15/15, image/OCR 12/12, endurance 60/60; selected profile 108.57/93.35/95.00 tok/s decode at nominal 4K/120K/230K and 3,487 MiB post-workload free per card |
245,760-token/C1 profile locally verified challenger, no-promotion; 393,216 and 499,712 profiles policy-infeasible; incumbent exactly restored |
GLM-5.3-Flash · SGLang qualification |
| 2026-08-31 |
GLM-5.3-Flash xgrammar fix-forward, matched speculation A/B, 524K context, C2, image/OCR, rollback, and real clients |
wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3-v1@319d66a8, incoai/GLM-5.3-Flash-DFlash2@dc77ff1c, corrected runtime digest 4909e318, TP=2/DCP=2, EXL3 K3 target, FP8 DS-MLA KV, DFlash2 fixed K5, 524,288/maxseq16/batch2,048, router c16, up to 16 images |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, matched performance, bounded quality, rollback, routed/client acceptance; both arms 28/28 functional; 83.08 tok/s at 4K and pooled 69.99 at 240K; C2 nominal 250K 2/2; 2,493,817 KV tokens; real OpenClaw/Hermes/Pi pass |
corrected 524K DFlash2 K5 selected as current; former 1M profile first rollback; text/image/OCR, no video; DFlash2 noncommercial boundary |
GLM-5.3-Flash · xgrammar qualification |
| 2026-08-30 |
GLM-5.3-Flash text/tools/image/OCR, 1M context, and K3/K5/chunk optimization |
wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3-v1@319d66a8, incoai/GLM-5.3-Flash-DFlash2@dc77ff1c, digest-pinned Purtell runtime 001a45bd, TP=2/DCP=2, EXL3 K3 target, FP8 DS-MLA KV, DFlash2 K5, 1,048,576/maxseq16/batch2,048, c16, up to 16 images |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, matched performance, bounded quality, routed/client acceptance; tools 20/20, image/OCR 12/12, quality 12/12, exact 950K-target retrieval, 82.1/67.4/67.9 tok/s at 4K/131K/240K; real Hermes/Pi/OpenClaw pass |
historical verified; first same-model rollback after the 2026-08-31 fix-forward; K3 alternate and batch4,096 rejection retained; text/image/OCR, no video; DFlash2 noncommercial boundary |
GLM-5.3-Flash · optimization and promotion |
| 2026-08-29 |
GLM-5.3-Flash text/tools/image/OCR, speculation, and 524K context qualification |
brandonmusic/GLM-5.3-Flash-tr3-4bpw@5ab363a8, digest-pinned Purtell EXL3/B12x runtime da5cec95, TP=2/DCP=2, NVFP4 DS-MLA KV; vision/text fixed K5, adaptive K1-K5+ReplaySSM, and no-spec controls at 262,144/524,288 |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, matched performance, bounded quality; vision/OCR, tools 20/20, high-reasoning coding 15/15, 250K-target / 206,296-actual retrieval, and 72.8/55.7 tok/s at 4K/128K; exact 495,045-token retrieval and 497,976-token tool use on 524K text profiles; adaptive tools 12/20 |
vision fixed K5 interactive and no-spec 524K maximum-context challenger profiles, no-promotion; adaptive MTP rejected; 0xSero 3.0-bpw watch-only; direct vision candidate retained for hands-on |
GLM-5.3-Flash · qualification |
| 2026-08-26 |
Primary image/OCR/video expansion and comprehensive context sweep |
RadixArk Qwen3.8 Flash Next NVFP4 7b719225, digest-pinned SGLang 59f06adc, engine d91c3682, hash-gated PR #36556 QSA, TP=2, 262,144 context, c1, BF16 KV, NEXTN 3/1/4, four images or one video |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, multimodal repeatability, routed/client acceptance; direct 30/30, isolated 27/30 then 30/30, live 29/30 then 28/30; edges 8/8; 25/25 context requests through 245,000 actual prompt tokens; 155.9/114.7/112.9 median decode tok/s at 4K/128K/254K targets; 516,032 server tokens |
human-authorized current text/image/OCR/video Primary; 262K/c1, four-image/one-video, thinking-disabled contract |
Qwen3.8 Flash Next · vision promotion |
| 2026-08-26 |
Text Primary QSA-fast/MTP3 fix-forward promotion |
RadixArk Qwen3.8 Flash Next NVFP4 7b719225, digest-pinned SGLang 59f06adc, engine d91c3682, hash-gated PR #36556 QSA fast path, TP=2, 262,144 context, c1, BF16 KV, NEXTN 3/1/4 |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, matched performance, bounded quality, routed identity and fresh real-client acceptance; tools 20/20 and quality 12/12; MTP3 154.9 tok/s at 4K, 134.1 at 128K, 102.0 at 253,703 actual prompt tokens with 8,192 output request; 2.33x/1.93x matched no-spec decode |
human-authorized current text Primary; 262K/c1 thinking-disabled contract; multimodal unpromoted |
Qwen3.8 Flash Next · QSA-fast promotion |
| 2026-08-26 |
Initial text Primary qualification and promotion |
RadixArk Qwen3.8 Flash Next NVFP4 7b719225, digest-pinned SGLang 59f06adc, engine d91c3682, TP=2, 262,144 context, c1, no speculation, portable QSA sparse decode plus NCCL logits fallback |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, bounded quality, routed and real-client acceptance; 253,325-token direct/routed retrieval with 8,192 reserve, intelligence 6/6, session 3/3, tools 20/20 plus repeated tools 3/3, Responses, OpenClaw/Hermes/Pi pass; 4K/c1 214.612 ms TTFT and 12.801 tok/s decode |
superseded same day by the QSA-fast/MTP3 fix-forward promotion; retained correctness baseline |
Qwen3.8 Flash Next · initial promotion |
| 2026-08-21 |
DeepSeek 0731 Infernal Invocation r18 1M promotion |
Official 9e165c30, digest-pinned r18 414ec7d0, B12X W4A8/FP8 compressed MLA KV, TP=2/DCP=1, 1,048,576 context, maxseq8, batch4,096, fixed probabilistic DSpark K5; matched no-spec control |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, matched performance, bounded quality, routed and real-client acceptance; 1,040,063-token retrieval, tools 160/160 plus functional batches, agentic 12/12, c8 short and c2 at 490,861 prompt tokens/request; K5/no-spec decode 142.1/76.4 tok/s at 4K and 129.5/76.3 at 32K; Hermes/Pi/OpenClaw pass |
former human-approved text Primary as of 2026-08-26; retained evidence |
DeepSeek 0731 · promotion |
| 2026-08-16 |
DeepSeek 0731 Infernal Invocation r15 393K promotion |
Official 9e165c30, digest-pinned r15 f1b13c86, B12X W4A8/FP8 compressed MLA KV, TP=2/DCP=1, 393,216 context, maxseq8, batch4,096, fixed probabilistic DSpark K5; matched target-only control |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 under WSL2 |
functional, capacity, matched performance, bounded quality, routed acceptance; K5/no-spec 4K decode 150.0/76.4 tok/s, direct 351,118-token and routed 340,119-token retrieval, repeated quality 12/12, routed tools and OpenClaw-compatible wire paths pass |
former human-approved text Primary; r33 393K fixed-port managed rollback at that date; actual Mini OpenClaw turn remained unproven |
DeepSeek 0731 · promotion |
| 2026-08-16 |
Qwen3.8 27B SGLang video and router admission |
Official FP8 017b9c7a, digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 E4M3 KV, EAGLE 3/1/4, c1, CPU feature transport |
1x RTX PRO 6000 Max-Q active; second equal card dormant |
functional, deterministic multimodal, routed acceptance; direct 30/30 with video 14/14, live admitted 28/28, overflow/malformed/SSE/tool and Primary regression pass |
expand current in place with vision.video, two images/one video fail-closed admission; no model restart |
Qwen3.8 27B · video-router finding |
| 2026-08-15 |
Remote agentic and SWE-bench Verified scout |
Official FP8 017b9c7a, digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 E4M3 KV, EAGLE 3/1/4, c1; router-only Apple evaluation worker |
1x RTX PRO 6000 Max-Q active; second equal card idle |
bounded quality; agentic smoke 2/2, scout 16/18, fixed SWE Verified sample 5/5 official-grader resolved |
retain current; debug-loop follow-up and larger stratified SWE sample recommended; no route or promotion change |
Qwen3.8 27B · agentic/SWE scout |
| 2026-08-15 |
Qwen3.8 27B SGLang official-FP8 single-service promotion |
Official FP8 017b9c7a, digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 E4M3 KV, max-running1, 2K chunks, FlashInfer, EAGLE 3/1/4, five GDN states, CPU feature transport |
1x RTX PRO 6000 Max-Q active; second equal card empty |
functional, capacity, deterministic multimodal, routed and real-client acceptance; 108K retrieval, tools 20/20, direct and routed media 18/18, Responses, Hermes and OpenClaw pass |
human-approved current; Primary/general-vision/OCR consolidated on one service; two-image/no-video admission |
Qwen3.8 27B · promotion |
| 2026-08-15 |
Qwen3.8 27B SGLang single-service consolidation A/B |
Official BF16 1d4bf0f, official FP8 017b9c7a, and Inferact NVFP4 6128240e; digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 KV, max-running1, 2K chunks, FlashInfer, EAGLE 3/1/4, CPU feature transport |
2x RTX PRO 6000 Max-Q; BF16 fixed on one card, quantized candidates sequential on the other |
functional, capacity, deterministic multimodal; all three pass 18/18 across six image tasks including a two-image comparison; BF16/FP8/NVFP4 media p50 0.915/0.588/0.448 s and 4K decode 62.7/111.4/97.7 tok/s |
Official FP8 preferred single-service challenger; NVFP4 lowest-latency third-party alternative; no-promotion; exact current vLLM split restored and routed |
Qwen3.8 27B · consolidation A/B |
| 2026-08-15 |
Qwen3.8 27B SGLang MTP=3 and CPU-transport multimodal A/B |
Official FP8 017b9c7a and Inferact NVFP4 6128240e; digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 KV, max-running1, 2K chunks, FlashInfer, EAGLE 3/1/4, five GDN states; multimodal arms force CPU feature transport |
2x RTX PRO 6000 Max-Q, one independent candidate per equal card followed by a swap |
functional, capacity, bounded deterministic quality, bounded multimodal; official FP8 MTP averages 111.3 decode tok/s and NVFP4 98.1; both pass 389K retrieval, intelligence 6/6, session 3/3, tools 3/3, image, and OCR |
SGLang MTP/CPU-multimodal challenger, no-promotion; official FP8 wins decode, NVFP4 wins TTFT/prefill; exact current vLLM split restored and readmitted |
Qwen3.8 27B · MTP/multimodal qualification |
| 2026-08-15 |
Qwen3.8 27B SGLang official-FP8/NVFP4 A/B |
Official FP8 017b9c7a versus Inferact ModelOpt NVFP4 6128240e; digest-pinned SGLang c4271c3, TP=1, 393,216 context, FP8 KV, max-running1, 2K chunks, FlashInfer, no speculation; models swapped across cards |
2x RTX PRO 6000 Max-Q, one independent candidate per equal card |
functional, capacity, bounded deterministic quality; both placements pass full gate, tools 20/20, repeated intelligence/session/tools, and 388,979-token retrieval; NVFP4 averages 0.429 s TTFT, 8,409 prefill tok/s, and 57.9 decode tok/s versus 0.554 s, 6,512, and 48.0 |
NVFP4 text challenger, both no-promotion; SGLang multimodal unqualified on WSL2; exact current vLLM split restored and readmitted |
Qwen3.8 27B · SGLang/NVFP4 qualification |
| 2026-08-15 |
Qwen3.8 27B official-FP8 MTP-depth A/B |
Official FP8 017b9c7a, vLLM 3a091411, TP=1, 393,216 context, FP8 KV, maxseq1, batch 4,096, MTP=4 versus MTP=5; settings swapped across cards |
2x RTX PRO 6000 Max-Q, one independent candidate per equal card |
functional, capacity, bounded deterministic quality; both pass tools 20/20, repeated intelligence/session/tools, and 388,979-token retrieval; same-card MTP=5 decode only 0.4-1.3% above MTP=4 with no E2E win |
MTP=4/5 no-promotion; retain current MTP=3; exact FP8/BF16 split restored and readmitted |
Qwen3.8 27B · MTP-depth qualification |
| 2026-08-14 |
Qwen3.8 27B split promotion |
Official FP8 017b9c7a text plus official BF16 1d4bf0f2 multimodal, vLLM 3a091411, TP=1 per card, 393,216 context, FP8 KV, maxseq1, batch 4,096, MTP=3; BF16 32 images/request |
2x RTX PRO 6000 Max-Q, one independent serve per equal card |
functional, capacity, multimodal, client acceptance; FP8 93.6 tok/s at 4K, routed tools 20/20; BF16 routed media 30/30 and 32-image request 1/1; Hermes text/image and OpenClaw Primary/vision passed without fallback |
human-approved current split; FP8 text Primary, BF16 explicit general-vision/OCR; TP=2 long-context profiles remain experimental |
Qwen3.8 27B · promotion |
| 2026-08-14 |
Qwen3.8 27B TP/MTP/long-context matrix |
Official BF16 1d4bf0f2 and official FP8 017b9c7a, vLLM 3a091411; split TP=1 at 393K and exclusive TP=2 at 393K/600K/1.01M; matched no-MTP/MTP=3, FP8 KV, maxseq1, batch 4,096 |
2x RTX PRO 6000 Max-Q over PCIe without P2P or NVLink |
functional, capacity; all 16 arms passed complete functional gates and cold retrieval at 388,979/598,729/985,107 actual prompt tokens; 4K FP8 MTP 85.9-93.6 tok/s; TP=2 control cut 393K TTFT 35-38% |
challenger, no-promotion; TP=2 for prefill/capacity, MTP for decode, 600K/1M batch-like; exact split baselines restored |
Qwen3.8 27B · matrix |
| 2026-08-14 |
Qwen3.8 27B official FP8 1M-context continuation |
Official FP8 017b9c7a, vLLM 3a091411, 1,010,000 configured context, TP=1, FP8 KV, maxseq1, batch 4,096, text-only, no MTP/prefix cache |
1x RTX PRO 6000 Max-Q; second equal card continued the BF16 serve |
functional, capacity; retrieval passed at 316,849, 422,449, and 633,649 actual prompt tokens, then 825,049 at 3/3 with 956.739 s mean E2E; post-stress and restored-control gates passed |
stable offline/batch challenger, no-promotion; original 262K FP8 lane restored, no route change |
Qwen3.8 27B · 1M continuation |
| 2026-08-14 |
Qwen3.8 27B official qualification and setting A/Bs |
Official BF16 1d4bf0f2 multimodal and official FP8 017b9c7a text, vLLM 3a091411, 262,144 context, TP=1 per split lane; FP8 MTP=3, prefix-cache, and unquantized-KV one-variable arms |
2x RTX PRO 6000 Max-Q, one independent card per co-resident serve |
functional, capacity, bounded quality, multimodal; both baselines through 241,250 actual prompt tokens, FP8 c5 30/30, BF16 media 30/30; MTP c1 94.8 tok/s, warm 30K-prefix TTFT 0.41 s |
challenger, no-promotion; direct baselines healthy, no route change; durable worker agentic/SWE open |
Qwen3.8 27B · qualification |
| 2026-08-11 |
DeepSeek 0731 r33 393K Primary promotion |
Official revision 9e165c30, digest-pinned r33 B12X mixed NVFP4-MoE/FP8 path, FP8 DS-MLA KV, DSpark K5, TP=2, 393,216 context, maxseq16, batch 4,096, GPU-only |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 |
functional, capacity, bounded quality; direct 359,900-token pass, routed functional pass except legacy long-needle calibration and trivial-prompt reasoning-evidence checks; OpenClaw and Hermes 393K/32K/high client-path passes |
human-approved current; routed/OpenClaw/Hermes >300K and SWE score remain open |
DeepSeek 0731 · promotion |
| 2026-08-10 |
DeepSeek 0731 r33 batch-token A/B |
Official revision 9e165c30, digest-pinned r33 B12X mixed NVFP4-MoE/FP8 path, FP8 DS-MLA KV, target-only/no-spec, TP=2, 131,072 context, c1; only functional change 8,192 to 4,096 batched tokens |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 |
functional, capacity, bounded performance; fresh bracketed starts, 6/6 preflight, matched 119,503-prompt-token pass; activation 1.73 to 1.14 GiB, minimum-rank KV 15.27 to 15.99 GiB; reported KV tokens 283,917 to 553,243 with accounting caveat |
4,096 candidate healthy and direct-only at campaign close; next arm GPU-only 393K; no-promotion |
DeepSeek 0731 · batch-token A/B |
| 2026-08-10 |
DeepSeek 0731 r33 quality-first control |
Official revision 9e165c30, digest-pinned r33 B12X W4A8 mixed NVFP4-MoE/FP8 runtime, FP8 DS-MLA KV, target-only/no-spec, TP=2, 131,072 context, c1; 393K FP8-KV plus 16 GiB host-offload recipe translated but not loaded |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive TP=2 |
functional, capacity, bounded quality; preflight 6/6, repeated intelligence/session/tools pass, 119,503 actual prompt tokens at 17.445 s TTFT and 73.86 decode tok/s; context-target calibration caveat retained |
priority challenger, no-promotion; no route change |
DeepSeek 0731 · r33 control |
| 2026-08-07 |
DeepSeek 0731 Vision (NVFP4) first load |
webbrain-one/DeepSeek-V4-Flash-0731-Vision-NVFP4 rev 3a8f168c, digest-pinned SGLang v0.5.16 image, marlin/marlin kernels, TP=2, KV fp8_e4m3, 4,096 context, --mem-fraction-static 0.97 |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; exclusive dual-gpu-exclusive |
compatibility-only, bounded functional (text lane), bounded negative quality (vision lane) |
no-promotion; failed back to the 650K Primary in the same session |
DeepSeek 0731 · vision first-load |
| 2026-08-03 |
Remote context, agentic recovery, and SWE-bench Verified smoke |
Official DeepSeek 0731 revision 9e165c30, pinned r16 B12X W4A8/FP8 image, FP8 MLA KV, TP=2, DSpark K5, 650K/maxseq16; Apple evaluation worker |
2x RTX PRO 6000 Max-Q over PCIe without NVLink |
functional, bounded quality; 8K context 1/1, tool-error retry protocol pass with final-answer fail, SWE-bench Verified 1/1 official-grader smoke |
benchmark substrate qualified for scout; no promotion or route decision |
DeepSeek 0731 · smoke finding |
| 2026-08-02 |
DeepSeek 0731 Primary promotion |
Official revision 9e165c30, pinned r16 B12X W4A8/FP8 image, FP8 MLA KV, TP=2, DSpark K5, 650K/maxseq16/batch4096; router output cap 32768 |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; display on AMD iGPU |
functional, capacity; Dark Pi, Mini Pi, Mini OpenClaw, oversized-output clamp, exact readiness and exclusive ownership |
current, human-approved; 1M removed from client-facing Primary after two fatal client-shaped workspace failures |
DeepSeek 0731 · promotion |
| 2026-08-02 |
DeepSeek 0731 GPU-only 650K/1M Pi qualification |
Official revision 9e165c30, pinned r16 B12X W4A8/FP8 image, FP8 MLA KV, TP=2, DSpark K5; 650K/maxseq16/batch4096 and 1M/maxseq1/4/16/batch2048 |
2x RTX PRO 6000 Max-Q over PCIe without NVLink; display on AMD iGPU |
functional, capacity; 640K and 985K retrieval, matched 32K c1 speed, three-tool burst, retained maxseq1 fatal workspace failure |
650K/maxseq16 preferred everyday Pi experiment; 1M/maxseq16 preferred deep-session experiment; maxseq4 passing alternative; maxseq1 rejected; all no-promotion |
DeepSeek 0731 · 650K/1M Pi qualification |
| 2026-08-02 |
DeepSeek 0731 native KV offload and 256K |
Official revision 9e165c30, derived r16 B12X mmap-unpinned image 331b7925, FP8 MLA KV, TP=2, DSpark K5, 8/16 GiB CPU offload, 262,144 served tokens, c1 |
2x RTX PRO 6000 Max-Q over PCIe without NVLink |
functional, capacity; 128K store/replay, cold 125K/192K/250K ladder, 16 GiB 113,408-token CPU reload, live lifecycle regression |
priority challenger, no-promotion; 128K remains preferred performance lane |
DeepSeek 0731 · native-offload/256K qualification |
| 2026-08-01 |
DeepSeek 0731 DSpark and 128K |
Official revision 9e165c30, pinned r16 B12X W4A8/FP8 image, FP8 KV, TP=2, DSpark K5 and same-image no-spec control, c1 |
2x RTX PRO 6000 Max-Q over PCIe without NVLink |
functional, capacity, bounded quality; low/high/max preflight, 128K, 27/27 coding-agent attempts, paired speculative A/B, per-card telemetry |
priority challenger, no-promotion; DSpark preferred for experiments, 3 GiB reserve failed |
DeepSeek 0731 · r16 qualification |
| 2026-08-01 |
Exclusive TP=2 LLM campaign |
DeepSeek V4 Flash 0731, Inkling Small NVFP4, Qwen3.5 122B NVFP4, Nemotron 3 Super 120B NVFP4, and Laguna S 2.1 NVFP4; exact pinned revisions, c1, no speculative decode |
2× RTX PRO 6000 over PCIe without NVLink |
external-prior, functional, capacity, quality; failed/calibration lanes retained |
all no-promotion; DeepSeek is the priority intelligence challenger; production aliases unchanged |
Dossiers · dual-PRO campaign · DeepSeek research |
| 2026-07-29 |
Primary LLM promotion |
Agents-A1 official FP8 4d7d5938, FP8 KV, vLLM f25953cc, 262K, c1, thinking off |
RTX PRO 6000 |
functional, capacity, quality |
current; Qwen3.5 becomes immediate rollback |
Agents-A1 · promotion |
| 2026-07-29 |
262K multimodal head-to-head |
Agents-A1 official FP8 4d7d5938, FP8 KV, vLLM f25953cc; Qwen3.5 122B NVFP4 98915d83, BF16 KV, NGC 26.06; thinking off, c1 |
RTX PRO 6000 |
functional, capacity, quality |
Agents-A1 wins bounded comparison; Qwen stays current; no-promotion |
Agents-A1 · Qwen3.5 · head-to-head |
| 2026-07-28 |
Multimodal challenger |
Agents-A1 BF16 addff08f, official FP8 4d7d5938, and ProtoLabs NVFP4 ff24ba5c; vLLM f25953cc, FP8 KV, 131K; isolated router and FP8 MoE A/B |
RTX PRO 6000 |
functional, capacity, quality |
FP8 principal candidate; NVFP4 compact text Pareto; generated tune rejected; no-promotion |
Agents-A1 · multimodal qualification |
| 2026-07-28 |
Primary LLM |
Qwen3.5 122B A10B NVFP4, rev 98915d837c4e7c87ac8296d02e89de19b3207e6d, BF16 KV, 262K, c1 |
RTX PRO 6000 |
functional, capacity, quality |
current |
Qwen3.5 · qualification |
| 2026-07-27 |
Challenger LLM |
Agents-A1, rev addff08f1653ee72765c5cf458fe84556bb34f8e, thinking disabled |
RTX PRO 6000 |
functional, capacity, quality |
challenger, no-promotion |
Agents-A1 · release sweep |
| 2026-07-27 |
Heavy control |
Laguna S 2.1 NVFP4, rev 07614121b31898586430f189d27a25a0be310843 |
RTX PRO 6000 |
functional, capacity, quality |
rollback |
Laguna · release sweep |
| 2026-07-26 |
Heavy LLM |
Laguna S 2.1 NVFP4, pinned nightly, FP8 KV, 262K, thinking disabled |
RTX PRO 6000 |
functional, capacity, quality |
rollback |
Laguna · qualification |
| 2026-07-18 |
Heavy LLM |
GPT-OSS Puzzle 88B, rev 9c0e0746a0d2218b28cc7b2cb3ce4e1a2f50fdb2, MXFP4/FP8 KV, 131K |
RTX PRO 6000 |
functional, capacity, quality |
rollback |
Puzzle · enablement |
| 2026-07-17 |
Heavy LLM |
GPT-OSS Puzzle 88B compatibility and qualification sequence |
RTX PRO 6000 |
compatibility-only, functional |
no-promotion at that point |
Puzzle · qualification |
| 2026-07-17 |
Heavy LLM |
Gemma 4 31B official QAT W4A16, vLLM 0.25.1, FP8 KV |
RTX PRO 6000 |
functional, capacity, quality |
rejected for latency |
Gemma 4 · optimization |
| 2026-07-16 |
Heavy LLM |
Gemma 4 official/Unsloth 12B, 26B, 31B configurations |
RTX PRO 6000 |
compatibility-only, capacity, quality |
no-promotion / rejected |
Gemma 4 · vLLM 0.25.1 sweep |
| 2026-07-16 |
Heavy LLM |
Gemma 4 Unsloth NVFP4 follow-up |
RTX PRO 6000 |
capacity, quality |
no-promotion |
Gemma 4 · follow-up |
| 2026-07-16 |
Heavy LLM |
Gemma 4 chat-template and size bakeoff |
RTX PRO 6000 |
compatibility-only, quality |
no-promotion / rejected |
Gemma 4 · template bakeoff |
| 2026-07-13 |
Heavy LLM |
Qwen3.6 27B container recipe and first characterization |
RTX PRO 6000 |
functional, capacity |
no-promotion |
Qwen3.6 · recipe |
| 2026-07-12 |
Heavy LLM |
ThinkingCap Qwen3.6 27B FP8, rev e48255afd77b403446332be0f595868337b36591 |
RTX PRO 6000 |
functional, capacity, quality |
historical control |
Qwen3.6 · promotion-era record |
| 2026-07-12 |
Heavy LLM |
Qwen3.6 27B community NVFP4+MTP, official FP8, Unsloth NVFP4, ThinkingCap |
RTX PRO 6000 |
functional, capacity, quality |
no-promotion |
Qwen3.6 · variation bakeoff |
| 2026-07-12 |
Heavy LLM |
Qwen3.6 protocol-v2 comparison |
RTX PRO 6000 |
quality |
no-promotion |
Qwen3.6 · comparison |
| 2026-07-12 |
Heavy LLM |
Qwen3.6 baseline |
RTX PRO 6000 |
functional, quality |
no-promotion |
Qwen3.6 · baseline |
| 2026-07-12 |
Heavy LLM |
Qwen3.5 122B MXFP4, rev 345839ea666a70f5035672f7c88afcba6281921f |
RTX PRO 6000 |
functional, capacity, quality |
no-promotion |
Qwen3.5 · MXFP4 benchmark |
| 2026-07-12 |
Heavy LLM |
Nemotron 3 Super 120B and Mistral Small 4 |
RTX PRO 6000 |
functional, capacity, quality |
no-promotion |
Nemotron Super, Mistral · challengers |
| 2026-07-12 |
Heavy LLM |
Nemotron Puzzle 75B recheck |
RTX PRO 6000 |
functional, capacity, quality |
no-promotion |
Nemotron Puzzle · recheck |
| 2026-07-12 |
Heavy LLM |
GPT-OSS 120B deterministic recheck |
RTX PRO 6000 |
capacity, quality |
no-promotion |
GPT-OSS 120B · recheck |
| 2026-07-12 |
Heavy LLM |
Laguna XS 2.1 protocol-v2 attempts |
RTX PRO 6000 |
historical-invalid, compatibility-only |
rejected |
Laguna · eval v2 |
| 2026-07-10–11 |
Heavy LLM |
Ornith 35B FP8, MiniMax M2.7 REAP, DeepSeek V4 Flash, Nemotron Puzzle, Qwen3.6 NVFP4+MTP |
RTX PRO 6000 |
functional, capacity, quality, historical-invalid |
no-promotion / rejected |
Dossiers · Blackwell bakeoff |
| 2026-07-10 |
Heavy LLM |
Qwen3.5 122B NVFP4, NGC 26.04, FP8 KV, 131K |
RTX PRO 6000 |
functional, capacity, quality |
no-promotion at that point |
Qwen3.5 · candidate record |