GLM-5.3-Flash DCP1 r11 qualification evidence¶
Campaign: 2026-09-23-glm-dcp1-qualification
State: publication_verified
Decision: exact DCP1 batch-2,048/C4/0.97 r11 selected and deployed after
accepted direct, production-routed, and native-client checks. The merged Pages
publication returned HTTP 200 and Chrome readback verified the published page.
This is a public-safe publication-verified evidence selection. Native result schemas remain native. Local endpoints, paths, and runtime identifiers were removed; reviewed synthetic benchmark output remains where it is needed to audit a claim. The artifact manifest has ten retained roles. The final solver review, client/control-plane review, dashboard publication receipt, and public HTTP/browser readback are retained.
Retained interim evidence¶
- DCP1 201K unique-prefix 8K cohort, repeated-prefix cohort, and 310K/C2 overlap: retained native stability schemas.
- Original 201K/4K unique-prefix cohort: retained failed third anchor with no visible answer. The later 8K cohort does not replace this result.
- Short quality screen and agentic gate: 12/12 and 18/18 respectively. These are separate gates, not a general model ranking.
- Natural-output token count: 18,016 visible tokens with a complete marker. The exact reviewed output and independent coherence review bind this separate, stop-finish natural-output gate; no integrated application-correctness claim follows.
- 64K native diagnostic, full synthetic visible
output, summary,
execution record, before/after
identities / after,
and independent review: the
65,536-token completion cap ended with
lengthafter 5,921 visible tokens (24,123 characters), three complete chapters, and a partial fourth; the completion marker is absent. The native failure remainscompletion_budget_exhausted_after_visible_output. Full reasoning is not retained; only 252,061 reasoning characters and null reasoning-token metadata remain. The review found coherent, non-degenerated visible text through its truncation, allowing the predeclared routed diagnostic to proceed without a waiver. It also records a material one-attempt visible-yield regression versus the matched incumbent (24,123 versus 70,772 characters; three versus eleven complete chapters), with no supported causal explanation or reasoning-loop claim. - Isolated routed preflight and its same-identity execution receipt retain 28 passed observations across eight functional families, including a 20-request tool batch. This is a routed functional result, not a capacity ranking.
- Routed admission native result, execution
receipt, active snapshots,
terminal decisions, and router errors
bind an exclusive offered-eight, nominal-32K window to eight terminal rows
with one configuration. Six served requests passed identity, own-canary, and
strict-output checks. Two HTTP 503s are explicitly
admission_timeout, so the native exit-1 performance result is ineligible and the optional 8/8 completion objective failed. The retained initial false gate documents the absent native active config field; gate v2 corrects the join by binding the eight gateway IDs to the eight terminal rows with the same configuration, without a rerun. It closes the configured C4 dispatch/admission boundary only: sampled dispatched plus streaming attained and did not exceed four across 75 samples. It does not qualify C8, sustained queueing, cancellation, or performance. - Final GPU window summary reports 4,273
samples per GPU, a 1,721 MiB sampled minimum free memory on each card, and
maximum temperatures of 87/85 C. The source CSV is hash-bound by the summary;
one-second sampling is not a continuous minimum. Candidate status
records 52 GiB peak host memory at its limit, 13,361
maxevents, and zero OOM events; pressure remains a caveat, not an OOM or general reliability claim. - The final independent qualification review recommends controlled promotion of the exact DCP1 batch-2,048/C4/0.97 recipe under the retained user authorization; r11 was subsequently deployed. DCP1 has 825,268 KV tokens versus DCP2's 1,416,244 (a 41.73% reduction), so four full 327K windows are unsupported. It is a mitigation, not a root-cause fix or crash-rate result, and it is not a speed winner.
- r11 startup, scenario binding,
direct preflight, and production-routed
preflight retain the reviewed r11
identity and 28/28 passed direct plus production-routed functional checks.
The subsequent native-client acceptance is retained below. The historical
pre-deployment review's
promoted=falsefield remains dated evidence rather than the current deployment state. - Native client acceptance records 13 passed
semantic paths, not 13 strict-format passes. Five Pi paths and Pi Web returned
fenced JSON, as did two secondary Hermes vision responses; Mac OpenClaw also
prepended prose before correct JSON. These remain visible strict-format
failures while their semantic/tool continuation paths pass. OpenClaw's actual
reserve is 20K; retired upstream reserve knobs are not represented as a 65,536
reserve. Pi Web's
messageCount=0is a metadata quirk because its retained native transcript passed. Controller mount verification records read-only mounts and unchanged siblings; all fleet repeat targets converged withchanged=0. The public WebUI browser path passed its user-completed login, calculate-tool, and exact1591check. The independent control-plane review accepts semantic client and controller convergence with those format, correlation, Pi Web metadata, and OpenClaw reserve caveats. The later public HTTP/browser readback closes publication verification. - The publication receipt verifies nine imported dashboard runs against 99 source-matched metrics, nine authenticated Workbench evidence cards, byte-identical repeat imports, and model readiness of one. The earlier receipt is retained as the pre-readback snapshot. The publication closure receipt records the merged public page's HTTP 200 and Chrome readback, plus the successful Pages run. The observability repair did not restart inference or the router and made no new benchmark request.
- Native SWE result: all five fixed tasks were submitted and officially graded, with four resolved. The absolute-gate disposition, environment comparison, and independent review retain the CPython/package mismatch: the absolute at-least-3/5 gate passes, but the 4/5 versus 3/5 baseline comparison is ineligible. The earlier false composite gate remains diagnostic evidence.
- Offered-C8 controls, seed 1801, and failed seed 2801: the last is performance-ineligible and closes the dependent C8 arm. The independent review confirms that the failed direct diagnostic stops C8 without by itself disqualifying the unchanged production-envelope C4 recipe.
- The original selected batch-2,048/C4 performance cells are preserved for the
dashboard importer at
performance-dcp1[-repeat1|-repeat2]-{c1-32k-unique,c4-128k-unique,c4-128k-shared}.json. They remain separate cache/context groups with descriptive 16-request cell metrics; they do not rank batch 4,096, include primes, or incorporate C8. - The native-derived chart manifest, SVG, and graph data were rendered twice with identical bytes. They plot only those nine eligible cells and preserve their unique/shared cache separation.
- First C4 reload window: sampled minimum free memory was 1,755 MiB on each GPU, maximum temperatures were 88/86 C, and the largest sample gap was 0.256 s. It covers warm-up, offered-C8 controls, and agentic work only; it is not a continuous minimum or the final SWE/64K window.
- Search closure and sanitization receipts: scope and public-byte provenance for this stage.
Gaps retained for campaign close¶
| Gate | State | Limit |
|---|---|---|
| Frozen SWE | complete | Absolute 4/5 passes at-least-3/5; paired 4/5 versus 3/5 comparison is ineligible. |
| 64K diagnostic | complete, bounded | The native completion-budget failure remains; coherent truncation satisfies the separate diagnostic rule. One-attempt visible yield regressed versus the matched incumbent. |
| Routed dispatch-cap C4 and real client | C4 boundary complete | 28 routed preflight observations passed; the offered-eight diagnostic preserves two bounded admission timeouts and no C8 claim. |
| Restoration | complete | r11 deployment and exact r10 rollback are retained. |
| Native-client acceptance | complete with format caveats | All 13 semantic paths passed; strict-format failures remain explicit. |
| Deployment review | complete | The historical recommendation and pre-deployment promoted=false record are retained; r11 was subsequently selected and deployed. |
| Public post/readback | complete | Pages run succeeded; the merged public page returned HTTP 200 and Chrome confirmed its title and core published facts. |
The shareable sanitized managed recipe is candidate-dcp1.public.toml. It preserves the managed schema, containment, environment, and flags while using the generic cache path and port. Its materialization prerequisite is documented in the prior recipe-reconstruction section; review a managed preview before any load.