GLM-5.3-Flash r10 APC campaign¶
Date: 2026-09-19. Decision: recommend a bounded, human-approved APC promotion. The campaign is complete; r9 is restored and APC has not been promoted. The later promotion follow-up records explicit approval and fresh acceptance; this finding remains the historical campaign close.
Result card¶
APC reduced mean visible response latency from 23.7 to 7.6 seconds on the shared-prefix workload. Fresh-prefix request throughput was 2.39% lower, within the frozen 5% limit. Both recipes passed the final repeated functional checks.
| Setup | Value |
|---|---|
| Model | brandonmusic/GLM-5.3-Flash-tr3-4bpw@a5fee929cf4888b1824323e33e8a19b60129e025 |
| Hardware | Two RTX PRO 6000 Blackwell Max-Q GPUs; TP2/EP2/DCP2 |
| Runtime | v84 image sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692; EXL3 4-bpw, FP8 MLA KV, no speculation |
| Contract | 327,680 configured context; C4; batch 2,048; eight images; native Max reasoning |
| Change | r9 APC disabled versus r10 APC enabled with Mamba alignment |
| Evidence | Local functional, capacity and bounded diagnostic quality |
| Decision | Challenger recommended for human approval; current baseline restored |
| Matched workload, 32 requests per arm | r9 baseline | r10 APC | Change |
|---|---|---|---|
| Shared prefix: mean visible TTFT | 23.70 s | 7.62 s | −67.84% |
| Shared prefix: mean first output | 9.22 s | 2.48 s | −73.15% |
| Shared prefix: completed requests/s | 0.146 | 0.462 | 3.17× |
| Unique prefix: mean visible TTFT | 24.10 s | 25.07 s | +4.02% |
| Unique prefix: completed requests/s | 0.144 | 0.141 | −2.39% |
Why it matters: repeated agent turns can reuse long prompt prefixes. APC improved that path with the checkpoint unchanged; fresh requests showed a small cost. This campaign does not establish an intelligence change.
Caveat: these are synthetic format/canary requests, not repository-editing tasks. One 32-request cell per arm and cache mode is bounded evidence, not a universal latency guarantee.
Evidence index · Artifact manifest · Publication summary · Independent review
Recipe research and exact configuration¶
The September 19 source check found no newer matching SM120 APC image from the runtime publisher. The newer DGX Spark work targets a different architecture and was a recipe lead, not a substitute for this two-GPU test. Dated source registry.
Both arms retain checkpoint, image, precision, topology, context, concurrency, batching and multimodal settings. The candidate also uses a tested 52 GiB RAM / 1 MiB swap assertion before model execution. The original baseline's unlimited-swap observation did not reproduce after managed recreation: the restored baseline enforces zero swap. Its cause remains unproven. No global host policy changed.
APC cache capacity is 1,394,028 tokens versus 1,495,412 without APC, a 6.78% reduction. The imported runtime commit is 6dc2f516688fe6f84c6994dcd20fddf296853a6c; its package reports 0.1.dev20111+g7f1e92bec.d20260827. Both identities are retained. Configuration.
Method and results¶
Finalists use nominal 32K context (about 25K actual prompt tokens), C4, strict 16-word output, request canaries, 4,096 completion tokens, temperature 0, top-p 0.95 and Max reasoning. Shared and unique populations remain separate. The shared APC pool was warm; isolated counters measured 605,696 hits across 801,376 queried tokens, or 75.58%. Controls ran on the restored baseline. All four finalist cells passed 32/32.
Reasoning volume varied: APC completion-token means increased 12.56% for shared and 5.42% for unique requests. Completed requests/s and visible latency therefore lead the comparison. First-output timing ends at the first reasoning or content delta. Visible TTFT ends at the first visible content delta and includes the delay spent on preceding reasoning. With 32 observations, p99 is descriptive rather than a stable tail estimate. Derived comparison and source hashes.
Final quality runs passed 12/12 per arm: six intelligence-marker, three tool and three session checks across three repetitions. Native template rendering verifies the Max-effort prompt control, not compliance with an internal reasoning budget. Both arms passed four eight-image cases. Candidate image/OCR, a 309,422-token needle and a 110,760-token tool call passed. No OOM was observed. These diagnostics do not establish broad intelligence or SWE parity.
Failures and limits¶
The original 64-word baseline failed two of eight exact word counts and is excluded from performance claims. The revised protocol remains strict. The earlier eight-request unique-prefix scout showed an 11.8% throughput regression; it remains visible, while the frozen 32-request finalist supplies the non-regression decision. Original two-repetition quality runs were below the runner's minimum and were superseded by matched three-repetition runs.
The candidate unique finalist omitted its embedded configuration fingerprint; an external artifact-hash and retained identity binding closes traceability without rewriting the native run. Largest measured context is 309,422 at C1; four simultaneous maximum-context requests and video were not tested. Friction and coverage.
Decision and restoration¶
Both frozen performance gates pass: at least 15% shared-prefix TTFT improvement and no more than 5% unique-prefix request-throughput regression. Independent review recommends a bounded APC promotion after human approval. Rollback conditions include correctness, containment or routed-readiness failure, or a reproduced unique-throughput regression above 5%.
The exact r9 recipe is restored, its primary route readmitted, direct and authenticated routed preflight passed, and the APC container and listener are absent. Router configuration and client catalogs were not changed. Production cutover and any required client convergence remain a separately approved operation. Restoration receipt.