GLM-5.3-Flash¶
Eight-image follow-up, 2026-09-18: GLM EXL3 r9 passed four eight-image comparison requests, one eight-image high-resolution request. Limit now eight/request; no concurrent eight-image soak.
Current status and review date¶
2026-09-18 update: Bounded vision acceptance passed 10/10 image attempts and routed checks and a matching Pi output transcript with unchanged 4-bpw weights, FP8 KV and disabled speculation. Eight images/request; configured 327K/C4 retained. A 259,922-token synthetic retrieval passed on retry after an initial refusal; this is not a full-window soak. The following September 14 snapshot remains the prior text-only evidence.
2026-09-19 APC campaign: matched shared-prefix mean visible TTFT improved 67.84% and request throughput rose 3.17×; unique-prefix throughput fell 2.39%, within the 5% gate. Final diagnostic quality passed 12/12 per arm, with image and long-context gates passed. Exact r9 restoration and direct/routed acceptance are verified. At the historical campaign close, independent review recommended a bounded human-approved APC promotion and no selected configuration changed. See the dated finding.
2026-09-19 APC promotion follow-up: explicit human approval and fresh managed acceptance promoted the bounded r10 APC configuration. This changes the current documented decision, while the campaign finding remains its historical restored-r9 close and r9 remains the exact rollback. See the follow-up.
Decision snapshot
- Product role: current bounded text, tools, image, and OCR reference.
- Selected or best-qualified configuration: GLM Flash EXL3 r10 APC, 4-bpw EXL3, FP8 MLA KV, 327,680 configured tokens, C4, with no speculation.
- Measured hardware: two RTX PRO 6000 Blackwell Max-Q cards, native Linux.
- Evidence: APC promotion follow-up and historical campaign: shared-prefix visible TTFT −67.84%, unique-prefix request throughput −2.39% within the 5% gate, diagnostic quality 12/12 per arm, plus bounded image and long-context checks.
- Decision: explicitly human-approved bounded r10 APC promotion with fresh acceptance. Exact r9 is the documented rollback; active route and startup assignments are private.
- Important limitation: performance uses synthetic 32-request finalist cells at about 25K actual prompt tokens, not a general speed or intelligence claim. No full-window concurrent capacity, video, broad SWE, soak, or fresh reboot proof is retained.
- Review dates: locally measured and reviewed 2026-09-19 UTC; the earlier 4-bpw qualification and mixed-quant startup failure remain retained history.
Review narrative¶
2026-09-17 — mixed 3.5-bpw startup failure¶
Status: failed compatibility-only attempt; rejected configuration, no-promotion. Measured: one managed load of pinned mixed K3/K4 EXL3 on dual RTX PRO 6000 Max-Q, TP2/DCP1, NVFP4 KV, configured 327,680 tokens/C1. Host RAM and 8 GiB swap exhausted before readiness; the desktop session shut down. Limits: zero candidate inference requests; no quality, context, latency, or GPU-fit conclusion. Retry requires host-memory containment and loader diagnosis. The exact selected 4-bpw baseline was restored and passed direct and routed short checks. Evidence: dated finding and raw excerpts.
2026-09-17 — bounded R7 loader fix-forward¶
Status: C1/C4 functional; experimental and no-promotion. Measured: the pinned mixed K3/K4 EXL3 candidate loaded under managed containment, passed C4 preflight and three-repetition chat/tool/intelligence scouts, and retrieved exact ZEBRA at 216,307 actual input tokens. One strict C4 capacity cell passed 4/4 and reported 98.17 tok/s only for that cell; a second passed 3/4 after a 65-versus-64 code-word failure. The matched baseline also passed only 2/4 after two 63-versus-64 failures. Limits: no overall win, full 327,680 context, C4 long-context capacity, representative repository/session/client acceptance, or repeated strict-capacity improvement. Evidence: dated finding and sanitized artifacts.
2026-09-14 — selected EXL3 r7 no-spec text lane¶
Status: quality, functional, and bounded capacity; selected on September 14 for text only. Measured: 4-bpw EXL3 at 327,680 total tokens with a 65,536-token output reserve and C4. The fixed one-pass MMLU-Pro scout scored 90/100; native agentic completed 30/30; the frozen official SWE scout resolved 4/5; strict120 completed 120/120; and Pi, Hermes, and OpenClaw tool checks passed. Post-promotion context completed 9/9 at 255,647–255,672 actual prompt tokens plus the 65,536-token reserve. The one matched no-spec C4 population had 38.908 tok/s median decode. Limits: no exhaustive intelligence winner, broad speed ranking, full-window concurrency soak, or fresh boot/reboot proof. Evidence: dated finding and sanitized evidence index.
2026-09-11 — completion-budget diagnosis¶
Status: bounded functional, no-promotion. Measured: real Pi on native Linux, same rc14/393K/C1 backend; 4096 versus 16384 completed 6/8 versus 7/8 valid tasks. Limits: eight coding runs invalid; larger budget did not rescue the reasoning fixture in both repetitions. Keep the existing output default. Evidence: dated finding. Dossier review for this bounded result: 2026-09-11; earlier qualification dates remain separate.
2026-09-11 — operator-reported daily experience, with confirmed deltas versus the previous serving¶
Status: operational narrative plus controlled-comparison deltas. Previous
serving: SGLang rc14 on Windows/WSL2, TP2, 393,216-token context, C1 —
deployed 2026-09-02 and serving until the native migration. Confirmed deltas
(same rc14 image, same weights, same hardware, both sides): per-request decode
medians rose from 112.07 → 149.02 tok/s at the 4K target (+33.0%), 96.17 →
125.03 at 120K (+30.0%), 102.42 → 124.46 at 262K (+21.5%), and 99.79 → 120.29
at 380K (+20.5%); the 60-sample endurance lane rose 102.19 → 142.64 tok/s
(+39.5%). Time-to-first-token at the 4K target fell from 0.178 s to 0.123 s
median, with the p95 pair at 0.193 s → 0.129 s; long-context TTFT stayed
prefill-bound and effectively unchanged (effective prefill ±2.3%). The
v0.4.3 runtime then added the capacity gains on top: the 524,288-token shared
pool with C4 four-request scheduling in place of C1 serialization, stock
template support for enable_thinking=false, and host-IPC deployment parity.
After promotion, the transition's final axis was restored: GPU peer-to-peer.
The two RTX PRO 6000 Max-Q cards share one PCIe host bridge (no NVLink on
Max-Q) and the driver reports bidirectional PCIe P2P OK; WSL2 could not use it
(the translated-IOMMU diagnosis). The 2026-09-08 comparison therefore
deliberately ran with NCCL_P2P_DISABLE=1 on both sides — its deltas are the
OS-migration effect alone. The P2P axis was then restored and measured separately: the
2026-09-09 authorized transport A/B re-enabled NCCL P2P on the native recipe
with matched evidence (4K n=12: median TTFT −9.0%, E2E −7.4%; 120K n=3:
TTFT −8.9%) and retained it as the default. The live v0.4.3 serving adds the
FlashInfer PCIe IPC all-reduce on top of that transport (startup proof
2026-09-11 09:45 UTC, max_numel=786432, custom all-reduce disabled in its
favor); that layer currently runs untuned seed configurations per the engine
log, so the FlashInfer workspace tune and a dedicated IPC all-reduce
measurement are the recorded follow-up.
Reported experience: the served v0.4.3 configuration is the operator's daily driver for agentic coding sessions, long-document review, and routed tool work through the Capability Gateway, with dependable day-to-day response quality, tool calling, and long-context behavior since the promotion gate. The operator further reports it as probably the strongest model they have used for this work — the first local configuration where the routine hard parts of agentic work complete locally, with escalation to the remote lead models becoming the exception for genuinely novel problems rather than the normal unblocking step. This is subjective operator experience with no escalation-rate metric behind it; it is recorded here because it is the qualification program's purpose outcome, and it should be re-examined as v0.4.3 operational traffic accumulates in Grafana.
Limits: the deltas are local whole-stack measurements from the 2026-09-08 controlled comparison — not a universal Linux speedup claim. The Grafana/Prometheus operational scrape (started 2026-09-09 09:52 UTC) covers the final WSL2 serving window through the 2026-09-11 07:50–08:10 UTC cutover: on real agent traffic the WSL2 baseline recorded TTFT median 0.49 s / p95 0.53 s and an engine inter-token gap of 6.0 ms median over about 185 scrape windows, while the native-era live window holds only the qualification campaign's traffic (TTFT p95 0.16 s, inter-token 5.9 ms median, about 30 minutes of samples), so like-for-like operational era comparison needs more accumulated v0.4.3 traffic; the confirmed cross-era deltas remain the controlled A/B. Subjective quality judgments remain operator experience; the measured record above and in the linked dated findings remains the decision evidence for any configuration change.
2026-08-29 — initial Cardillo/Purtell qualification¶
The brandonmusic/GLM-5.3-Flash-tr3-4bpw 262K/524K campaign remains retained
historical evidence and the 262K image remains the one-week rollback. Adaptive
K1-K5 plus ReplaySSM stays rejected for tool-call corruption. The unserved
0xSero 3.0-bpw release remains watch-only.
2026-08-30 — 1M optimization¶
The DFlash2 fixed-K5 profile with 2,048-token batching became the selected same-model configuration in the dated campaign. The matched K3 arm remained a verified alternate, while the 4,096-token scheduler-chunk arm was rejected.
2026-08-31 — xgrammar fix-forward at 524K¶
The fix-forward qualification selected the exact
wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3-v1 target with the DFlash2 fixed-K5
draft and a corrected xgrammar runtime as the replacement default. The
published contract covers text, tools, image, and OCR at 524,288 configured
tokens, router concurrency 16, and up to 16 images. Video is unsupported. The
former 1M K5 profile remains the first rollback.
2026-09-02 — pinned SGLang SM120 qualification and promotion¶
The exact ormandj W4A16/NVFP4 checkpoint and rc14 SGLang image were translated from an external native-Linux recipe to WSL2 and qualified under managed Anvil Serving lifecycle. A matched 131K adaptive-MTP/no-spec A/B showed adaptive MTP improving decode by 39.7% at 4K and 29.2% at the 120K target, with the 2.8% long-prefill difference inside the 3% equivalence band. The 393,216-token/C1 lane passed direct function, tools, long tool, image/OCR, reasoning, 15/15 coding, 12/12 media, 60/60 endurance, and the full 4K/120K/262K/380K performance sweep through 304,491 measured prompt tokens. It initially failed only the standing 3 GiB reserve policy at 2,101 MiB/card. After the operator declared the pair model-only and waived that reserve, a managed promotion passed direct, routed, and real Pi/OpenClaw/Hermes gates. The exact 524K incumbent remains the immediate rollback. The 245,760/C1 lane is a conservative verified fallback; the 499K lanes remain rejected or unverified.
2026-09-02 — isolated-worker SWE-bench smoke¶
An isolated macOS arm64 worker ran the fixed django__django-11099 smoke
through the declared routed endpoint with thinking enabled. Mini-SWE-agent made 11 routed model
requests, submitted a patch, and the pinned official SWE-bench grader resolved
the instance. The run attempted, graded, and resolved 1/1 in 63.495 seconds of
worker time. This adds bounded repository-agent evidence but is not a
representative full-suite score. The campaign also fixed the managed harness's
isolated Python environment, detached Docker discovery, ambient global config,
venv validation, and missing-trajectory classification; see the
SWE smoke finding.
2026-09-03 — concurrency and KV-capacity interpretation¶
A cross-run review separated scheduler ceilings from KV-resident capacity and measured long-context concurrency. The current SGLang profile remains C1 with a configured 393,216-token shared pool. The corrected 524K rollback reports 2,493,817 KV tokens and proved C2 at 206,630 prompt tokens/request. The historical BrandonMusic fixed-K5 524K lane completed C16 only with 4K prompts; its reported pool was 565,898 tokens, or 1.08 complete windows. Thus maxseq16 is not evidence for sixteen full 524K conversations. See the cross-run interpretation.
2026-09-08 — native Linux comparison¶
The existing pinned SGLang profile on the same physical GPU pair showed 20.5–33.0% higher historical-style decode than retained WSL cells. The comparison preserves variable short outputs and shared-prefix cache history; it does not establish strict controlled performance or an OS-only speedup. Coding, images, agentic and the identical official SWE smoke passed bounded checks, but the SWE agent trajectory became slower. Failed strict-output, canary, routed context/reasoning and literal image checks remain visible. No new promotion occurred. See the native Linux comparison.
Immutable identity¶
The selected 2026-09-14 text lane retains brandonmusic/GLM-5.3-Flash-tr3-4bpw
at a5fee929cf4888b1824323e33e8a19b60129e025 and
verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-language-only@sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692.
See the native promoted-context artifact
and configuration identity.
The selected native v0.4.3 profile uses ormandj revision
c3cbb9891b67c741bcbf6b176dd7af9265b069db, image
sha256:ec4243f940a179a27fea21895077efd47cd050501f99a1d2a5fecf7df2e7be71,
served identity glm-5.3-flash and derived template SHA256
3406cc91800b56c06439e65aaea1929e6bc803cdb62f7f995f25ab082efd3408.
Full configuration.
The remaining identities in this section are retained historical profiles.
- Target:
wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3-v1 - Target revision:
319d66a8b53092b491f698440ecea781e4ddd4e4 - Draft:
incoai/GLM-5.3-Flash-DFlash2 - Draft revision:
dc77ff1c99eeb2df044ee3d4f0094eb033fee410 - Runtime image:
sha256:4909e318ba1348a179824e210f90c268d6fc68e8b4e514af4782e26e6a1e5939 - Runtime source base:
vllm-project/vllm@487ecf187, with the hash-gated xgrammar reasoning-end/speculative-validation patch underconfigs/runtime-patches/vllm/487ecf187-xgrammar-spec-reasoning-end/ - Runtime-reported vLLM:
0.1.dev20051+g487ecf187 - Served identity:
glm53-flash-exl3-k3-dflash2-k5-fp8-tp2-524k-vision-xgfix - Quantization: EXL3/MCG K3 routed experts with native attention/shared/vision/MTP tensors; FP8 DS-MLA target KV; BF16 DFlash2 draft KV
The 2026-09-02 current SGLang profile has a separate immutable identity:
- Target:
ormandj/GLM-5.3-Flash-W4A16-NVFP4-K32-Experts-FP8-WO - Target revision:
c3cbb9891b67c741bcbf6b176dd7af9265b069db - Runtime image:
ghcr.io/ormandj/sglang-glm53-flash-sm120@sha256:0c0637959c3931829f05154087bbefd2c50003fb9b2010200ce0ec82f4d71a53 - Runtime source label:
ormandj/sglang-glm53-flash-sm120@a547c90 - SGLang source label:
sgl-project/sglang@4c2c169 - Served identity:
glm53-flash-ormandj-sglang-sm120-tp2-393k-c1-adaptive-mtp - Quantization: ModelOpt W4A16/NVFP4 K32 experts, FP8 weight-only dense tensors, FP8 KV, and adaptive EAGLE speculative decoding
The target, draft, and runtime are third-party artifacts. The DFlash2 draft is CC-BY-NC-ND-4.0, so the combined recipe is evaluation/noncommercial unless separate permission is obtained.
Tested hardware and topology¶
The 2026-09-08 native Linux comparison measures the same physical dual-PRO
pair at TP2/393,216/C1. It preserves the exact SGLang image and checkpoint,
adds NCCL_P2P_DISABLE=1, and retains the existing compatibility patches.
OS, driver/container stack and cache history differ from the historical
baseline. The reconstruction boundary is recorded in the
comparison configuration.
The following topology describes the retained EXL3/WSL qualification lane.
Two RTX PRO 6000 Blackwell Max-Q cards, 96 GB each, under Windows 11, Docker Desktop, and WSL2. They are assigned exclusively to one TP=2/DCP=2 owner over PCIe without NVLink. Aggregate VRAM is 192 GB, not unified memory. The selected profile uses the V2 runner, B12x sparse MLA attention, 2,048-token batching, maxseq16, prefix caching, the visual tower, and 524,288 configured context.
CUDA-IPC peer access is unavailable in this WSL2 topology. The qualified translation uses PyNCCL and disables custom PCIe all-reduce, B12X DCP A2A, and top-k owner exchange while preserving EXL3 K3 compute, sparse B12X MLA, FP8 target KV, and the DFlash2 draft.
Engine, quantization, KV, context, and concurrency recipe¶
| Recipe | Role | Context | Decision |
|---|---|---|---|
| EXL3 r7 no-spec (public-safe configuration in dated finding) | selected text-only reference | 327,680 total / 65,536 output, C4 | current; bounded evidence |
| SGLang W4A16 adaptive MTP | historical 2026-09-02 text/image/OCR selection | 393,216 | verified, historical |
| DFlash2 K5, corrected xgrammar | historical rollback | 524,288 | verified, historical |
| SGLang W4A16 adaptive MTP, conservative | conservative fallback | 245,760 | verified, fallback |
| Matched no-speculation control | reliability/performance control | 524,288 | verified, control |
| Former DFlash2 K5, batch 2,048 | first rollback | 1,048,576 | historical verified, rollback |
| DFlash2 K3, batch 2,048 | former high-concurrency alternate | 1,048,576 | historical verified |
| K5, batch 4,096 | matched scheduler-chunk trial | 1,048,576 | rejected; no durable recipe |
| TR3 vision fixed K5 | prior interactive profile and image rollback | 262,144 | historical verified |
| TR3 no spec | prior maximum-context/headroom lane | 524,288 | historical verified |
The current EXL3 r7 no-spec lane admits C4 at 327,680 total tokens with a 65,536-token output reserve. It completed 9/9 post-promotion context cases at 255,647–255,672 actual prompt tokens plus reserve; this does not establish concurrent full-window capacity. See the dated finding.
Historical capacity profiles¶
The corrected 524K engine reported 19.18 GiB of available KV memory per rank and 2,493,817 KV tokens, or 4.76 complete configured windows. Router c16 is a scheduling ceiling for short requests, not proof of sixteen simultaneous 524K prompts. The measured long-context concurrency gate is C2 at a nominal 250K target. The locally qualified media contract is up to 16 images and zero videos.
The historical SGLang profile uses C1, 4,096 maximum output tokens, adaptive EAGLE, FP8 KV, and the multimodal visual tower. At 393,216 configured tokens it retained 2,101 MiB free on each card after qualification and 2,543 MiB after the promotion's routed/client work. The operator explicitly waived the standing 3,072 MiB reserve for this model-only pair without waiving runtime safety evidence. The 245,760 profile retained 3,487 MiB/card and remains the conservative fallback. The 499,712/C4 profile is rejected and the C1 variant is unverified.
For planning only, the historical 393,216-token shared pool divides to 196,608 total tokens/request at C2 and 98,304 at C4 before output, media, protocol, and runtime headroom. A first C2 text gate at 180,000 prompt + 4,096 output per request would leave 25,024 tokens in the shared pool. This is not qualified: that historical recipe also fixes one running request and batch-size-one decode graphs, so increasing concurrency requires a managed candidate load and the full functional, capacity, quality, endurance, and post-workload VRAM gates.
Evidence by measurement class¶
Selected EXL3 r7 no-spec, 2026-09-14¶
Status: quality, functional, bounded capacity; September 14 text-only reference.
Measured: MMLU-Pro 90/100 one pass, agentic 30/30, frozen SWE 4/5,
strict120 120/120 at C4, and context 9/9 at 255,647–255,672 prompt tokens plus
65,536 reserve. Limits: one 38.908 tok/s C4 median, no exhaustive
intelligence ranking, full-window soak, or fresh boot/reboot proof.
Evidence: finding and sanitized evidence.
Native Linux comparison, 2026-09-08¶
Status: functional, capacity, bounded quality; no-promotion.
Measured: warm direct C1/n3 shared-prefix, variable-short-output decode
149.02/125.03/124.46/120.29 tok/s at 4K/120K/262K/380K targets, versus WSL
112.07/96.17/102.42/99.79; long prefill approximately ±2.3%; endurance
60/60 at 142.64 versus 102.19 tok/s. Coding 15/15, image corpus 12/12,
agentic 30/30, official SWE smoke 1/1.
Limits: strict output 0/3 and unique canaries 1/10 excluded; no finalist
performance qualification. SWE agent time 115.69 versus 34.22 seconds,
22 versus 11 requests. Routed 380K admission and reasoning evidence failed;
direct image phrase sensitivity also failed. No OS-only causal claim.
The extended default-thinking context suite passed 128/150 through 376,484
actual prompt tokens. Nominal 8K/32K/131K/262K/380K buckets passed
27/18/24/30/29 of 30; failures were 9 empty length-terminated responses and
13 incorrect visible answers. The non-monotonic curve means the native
threshold-derived effective_context=8192 is not a physical cap. No matched
Windows suite exists, so this is not a Linux-caused quality regression.
Evidence: context summary;
dated finding and native artifacts.
Retained EXL3 and SGLang qualification history¶
functional: both matched arms passed 28/28 direct observations, including strict JSON, 206,296-actual-token retrieval, tools 20/20, a structured tool after the long prompt, streaming tools, tool-result continuation, and Responses. The selected arm also passed image understanding and OCR.capacity: C2 at a nominal 250K target / 206,630 actual prompt tokens per request completed 2/2 with an 8,192-token API completion allowance; the engine reported 2,493,817 KV tokens.- bounded
quality: intelligence 6/6, session 3/3, tools 3/3, and the 4K, 131K, and 240K context cases passed with no retained failure. performance: the matched DFlash2 arm measured 83.08 tok/s at 4K and a pooled 69.99 tok/s across two five-request 240K runs, versus 42.61 and 43.63 for no speculation. The two 240K run-level medians were 64.51 and 70.50.concurrency: the selected profile completed the measured 250K-class C2 gate 2/2. Router c16 remains a short-request scheduling ceiling.- multimodal: direct image understanding and verbatim OCR passed; up to 16 images are admitted. Video remains disabled.
- client acceptance: real OpenClaw and Hermes shell-tool continuations passed
with exact identity and no fallback. Pi 0.84.2's initial normal
extension-loaded PTY path made exactly one
readcall, recovered an unseen marker, and emitted zero error events; the goal-closure recheck on Pi 0.84.4 retained 524,288/8,192 and passed another real PTY tool nonce.
For the 2026-09-02 SGLang profile, thinking-disabled preflight passed all ten observations, tools passed 20/20, and long needle/tool cases passed, and image understanding/OCR passed. Thinking-enabled smoke, JSON, tools 5/5, tool-result continuation, and Responses passed with dedicated reasoning. Bounded coding passed 15/15 at a 2,048-token visible budget, the media corpus passed 12/12, and endurance passed 60/60. The three-run capacity sweep measured 112.07/96.17/102.42/99.79 tok/s decode and 16,729/5,749/5,579/5,457 tok/s effective prefill at nominal 4K/120K/262K/380K targets respectively. Managed direct and routed gates plus real Pi/OpenClaw/Hermes acceptance passed after promotion.
Matched local comparison¶
| Requested context | Matched no spec | Corrected DFlash2 K5 | Former 1M K5 | Current vs no spec | Current vs former 1M |
|---|---|---|---|---|---|
| 4K | 42.61 tok/s | 83.08 tok/s | 82.1 tok/s | +95.0% | +1.2% |
| 240K | 43.63 tok/s | 69.99 tok/s pooled | 67.9 tok/s | +60.4% | +3.1% |
The 240K selected value pools ten requests across two retained runs; their run-level p50 values were 64.51 and 70.50 tok/s. This comparison is local to the exact derived runtime and dual-Max-Q/WSL2 topology; it is not a general model-intelligence ranking.
Decision and promotion state¶
The 2026-09-14 EXL3 r7 no-spec lane was the selected September 14 text-only reference.
It does not establish an exhaustive intelligence winner, a general performance
ranking, or full-window concurrent capacity. The 2026-09-08 migrated-stack
measurement is no-promotion; all earlier selections below are retained
historical qualification and promotion decisions.
The 393K/C1 SGLang adaptive-MTP profile was the selected published text/image/OCR default in its dated record. The corrected 524K K5 profile was its immediate exact rollback; it retains measured C2 headroom at a 250K-class prompt and its corrected structured-output behavior. Raising its batching to 4,096 remains rejected.
The historical SGLang contract advertises 393,216 context, 4,096 maximum output,
C1, and image/OCR without video. During its recorded promotion,
llm.primary, llm.secondary, llm.auxiliary, llm.voice,
vision.general, and vision.ocr select the same exact service during this
evaluation. Qwen3.8 Flash Next remains the retained video-capable rollback.
The 524K EXL3/DFlash2 profile is its historical same-model rollback, and the
245,760/C1 SGLang lane is its conservative same-engine fallback.
Failures and gotchas¶
-
Current EXL3 r7 limits (2026-09-14): no exhaustive intelligence comparison, full-window concurrency soak, or fresh boot/reboot test is retained. The 38.908 tok/s C4 no-spec median is one matched population, not a general speed claim.
-
Native comparison limits (2026-09-08): strict controlled output passed 0/3; unique natural canaries 1/10. Both populations are excluded. Routed 380K text fails secondary media admission with HTTP 413; routed reasoning checks lack the required dedicated channel despite correct answers. Corresponding direct gates pass. One direct general-image literal phrase check fails on table formatting despite the separate 12/12 image corpus. These are retained failures, not relaxed validators or new qualifications.
-
DFlash2's published license is noncommercial/no-derivatives. Obtain separate permission before commercial use.
- The target/draft/runtime are community artifacts, not stock-vLLM support.
- The original DFlash2 runtime failed structured generation after reasoning termination. Only the digest-pinned corrected xgrammar image is qualified.
- Video is unsupported; the DFlash2 drafter receives text-only inputs on image calls while the target processes the image.
- WSL2 peer IPC failed, requiring the qualified PyNCCL transport translation.
- The runtime suggested a larger fixed KV pool, but it was not A/B tested; the selected 0.95 utilization retains measured operating reserve.
- No missing MoE/GEMM tune warning appeared. First-request JIT warnings are warm-up observations, not evidence that a kernel tune would help.
- The high-control quality request was accepted but independent token-level
reasoning telemetry was unavailable; the artifact says
requested_unverified. - The rollback profile needs its already-qualified visible/reasoning budget; an artificially small visible cap can end in reasoning without an answer.
- No exact Docker-image removal product surface exists. The previous GLM image remains the intended one-week rollback; no broad prune was used.
- The rc14 SGLang image required three narrow WSL2 fix-forwards:
NCCL_CUMEM_ENABLE=0, disabling expandable-segment allocation, and a hash-gated fallback from the symmetric-memory logits gatherer to ordinary NCCL. The derived chat template is also hash-gated so thinking-disabled requests do not leak internal reasoning. - The 393K/C1 profile's original reserve-policy failure was reclassified under the recorded model-only waiver; the physical 2,101 MiB/card measurement is unchanged. The 499K/C4 lane is rejected and the 499K/C1 lane is unverified.
- Sparse-MLA CPB calibration rejected implausible fitted results and retained the engine heuristic. Optional NCCL plugin warnings were benign. No kernel tune was stored because no exact default-versus-tuned end-to-end A/B showed an improvement.
Dated run history¶
| Date | Event | Result |
|---|---|---|
| 2026-09-14 | Intelligence and context qualification, EXL3 r7 no-spec | Selected September 14 text-only lane: MMLU-Pro 90/100 one pass, agentic 30/30, frozen SWE 4/5, strict120 120/120, context 9/9 at 255,647–255,672 prompt tokens plus 65,536 reserve, and Pi, Hermes, and OpenClaw tool checks pass; one C4 no-spec population 38.908 tok/s median; no exhaustive intelligence, full-window soak, or fresh boot/reboot claim; finding and evidence. |
| 2026-09-09 | ormandj v0.4.2 runtime qualification | Bounded direct gates and matched capacity improve, but strict turnover is 58/60 then 57/60; user selected retain-baseline/no-promotion and the exact baseline was restored; finding |
| 2026-09-09 | Native NCCL P2P transport A/B | User-authorized native default retains P2P enabled after direct/routed gates and bounded 4K n12/120K n3 latency evidence; both 380K strict-output cells and the raw PowerShell diagnostic marker failure remain explicit; finding |
| 2026-09-08 | Native Linux versus retained WSL on the same dual-card GLM image/model | C1/n3 historical-style decode +20.5–33.0%; bounded quality and SWE smoke pass, extended context 128/150 (9 empty, 13 incorrect), strict-output/canary and routed failures retained; whole-stack comparison, no-promotion; finding |
| 2026-09-03 | Cross-run scheduler, KV-capacity, and measured-concurrency reconciliation | Current SGLang remains qualified at 393K/C1; corrected 524K rollback retains measured C2 at 206,630 prompt tokens/request; historical BrandonMusic C16 is short-request evidence, not full-window concurrency; interpretation and artifact links |
| 2026-09-02 | Isolated-worker SWE-bench Verified smoke and harness fix-forward | Fixed django__django-11099 attempted 1/1, officially graded 1/1, and resolved 1/1 through 11 routed requests; one-instance smoke only; finding and sanitized evidence |
| 2026-09-02 | Human-approved 393K/C1 reserve reclassification, managed promotion fix-forward, routed gates, and real-client acceptance | Published current text/tools/image/OCR profile; 304,491 actual prompt tokens, 112.07/96.17/102.42/99.79 tok/s at 4K/120K/262K/380K targets, direct and routed gates, Pi/OpenClaw/Hermes pass, 2,543 MiB/card after client work, exact 524K rollback retained; promotion finding |
| 2026-09-02 | Pinned ormandj SGLang SM120 translation, WSL2 fix-forward, matched adaptive/no-spec A/B, 240K full qualification, larger-envelope feasibility, endurance, and exact incumbent restoration | 245,760/c1 adaptive MTP selected as a verified challenger with no promotion; 108.57/93.35/95.00 tok/s decode at 4K/120K/230K, quality 15/15, media 12/12, endurance 60/60, and 3,487 MiB/card post-workload reserve; finding and raw artifacts |
| 2026-08-31 | xgrammar fix-forward image, matched 524K no-spec/DFlash2 A/B, C2 250K-class gate, rollback drill, router, and real-client forward restore | corrected DFlash2 K5 selected at 524K; 83.08 tok/s at 4K and pooled 69.99 at 240K; C2 2/2; full finding and raw artifacts |
| 2026-08-30 | Current-source refresh, K5/K3/batch4,096 A/B, 4K-500K performance, 950K retrieval, image/OCR, quality, router, and real-client promotion | K5/batch2,048 selected as current one-week default; K3 verified alternate; batch4,096 rejected; full finding and raw artifacts |
| 2026-08-29 | Cardillo/Purtell translation, adaptive/fixed/no-spec A/B, vision/OCR, 250K and near-500K capacity | TR3 vision fixed K5 and no-spec 524K qualified as challengers; adaptive MTP rejected; 0xSero 3.0-bpw watch-only; historical finding |