Skip to content

GLM-5.3-Flash DCP1 r11 qualification evidence

Campaign: 2026-09-23-glm-dcp1-qualification State: publication_verified Decision: exact DCP1 batch-2,048/C4/0.97 r11 selected and deployed after accepted direct, production-routed, and native-client checks. The merged Pages publication returned HTTP 200 and Chrome readback verified the published page.

This is a public-safe publication-verified evidence selection. Native result schemas remain native. Local endpoints, paths, and runtime identifiers were removed; reviewed synthetic benchmark output remains where it is needed to audit a claim. The artifact manifest has ten retained roles. The final solver review, client/control-plane review, dashboard publication receipt, and public HTTP/browser readback are retained.

Retained interim evidence

  • DCP1 201K unique-prefix 8K cohort, repeated-prefix cohort, and 310K/C2 overlap: retained native stability schemas.
  • Original 201K/4K unique-prefix cohort: retained failed third anchor with no visible answer. The later 8K cohort does not replace this result.
  • Short quality screen and agentic gate: 12/12 and 18/18 respectively. These are separate gates, not a general model ranking.
  • Natural-output token count: 18,016 visible tokens with a complete marker. The exact reviewed output and independent coherence review bind this separate, stop-finish natural-output gate; no integrated application-correctness claim follows.
  • 64K native diagnostic, full synthetic visible output, summary, execution record, before/after identities / after, and independent review: the 65,536-token completion cap ended with length after 5,921 visible tokens (24,123 characters), three complete chapters, and a partial fourth; the completion marker is absent. The native failure remains completion_budget_exhausted_after_visible_output. Full reasoning is not retained; only 252,061 reasoning characters and null reasoning-token metadata remain. The review found coherent, non-degenerated visible text through its truncation, allowing the predeclared routed diagnostic to proceed without a waiver. It also records a material one-attempt visible-yield regression versus the matched incumbent (24,123 versus 70,772 characters; three versus eleven complete chapters), with no supported causal explanation or reasoning-loop claim.
  • Isolated routed preflight and its same-identity execution receipt retain 28 passed observations across eight functional families, including a 20-request tool batch. This is a routed functional result, not a capacity ranking.
  • Routed admission native result, execution receipt, active snapshots, terminal decisions, and router errors bind an exclusive offered-eight, nominal-32K window to eight terminal rows with one configuration. Six served requests passed identity, own-canary, and strict-output checks. Two HTTP 503s are explicitly admission_timeout, so the native exit-1 performance result is ineligible and the optional 8/8 completion objective failed. The retained initial false gate documents the absent native active config field; gate v2 corrects the join by binding the eight gateway IDs to the eight terminal rows with the same configuration, without a rerun. It closes the configured C4 dispatch/admission boundary only: sampled dispatched plus streaming attained and did not exceed four across 75 samples. It does not qualify C8, sustained queueing, cancellation, or performance.
  • Final GPU window summary reports 4,273 samples per GPU, a 1,721 MiB sampled minimum free memory on each card, and maximum temperatures of 87/85 C. The source CSV is hash-bound by the summary; one-second sampling is not a continuous minimum. Candidate status records 52 GiB peak host memory at its limit, 13,361 max events, and zero OOM events; pressure remains a caveat, not an OOM or general reliability claim.
  • The final independent qualification review recommends controlled promotion of the exact DCP1 batch-2,048/C4/0.97 recipe under the retained user authorization; r11 was subsequently deployed. DCP1 has 825,268 KV tokens versus DCP2's 1,416,244 (a 41.73% reduction), so four full 327K windows are unsupported. It is a mitigation, not a root-cause fix or crash-rate result, and it is not a speed winner.
  • r11 startup, scenario binding, direct preflight, and production-routed preflight retain the reviewed r11 identity and 28/28 passed direct plus production-routed functional checks. The subsequent native-client acceptance is retained below. The historical pre-deployment review's promoted=false field remains dated evidence rather than the current deployment state.
  • Native client acceptance records 13 passed semantic paths, not 13 strict-format passes. Five Pi paths and Pi Web returned fenced JSON, as did two secondary Hermes vision responses; Mac OpenClaw also prepended prose before correct JSON. These remain visible strict-format failures while their semantic/tool continuation paths pass. OpenClaw's actual reserve is 20K; retired upstream reserve knobs are not represented as a 65,536 reserve. Pi Web's messageCount=0 is a metadata quirk because its retained native transcript passed. Controller mount verification records read-only mounts and unchanged siblings; all fleet repeat targets converged with changed=0. The public WebUI browser path passed its user-completed login, calculate-tool, and exact 1591 check. The independent control-plane review accepts semantic client and controller convergence with those format, correlation, Pi Web metadata, and OpenClaw reserve caveats. The later public HTTP/browser readback closes publication verification.
  • The publication receipt verifies nine imported dashboard runs against 99 source-matched metrics, nine authenticated Workbench evidence cards, byte-identical repeat imports, and model readiness of one. The earlier receipt is retained as the pre-readback snapshot. The publication closure receipt records the merged public page's HTTP 200 and Chrome readback, plus the successful Pages run. The observability repair did not restart inference or the router and made no new benchmark request.
  • Native SWE result: all five fixed tasks were submitted and officially graded, with four resolved. The absolute-gate disposition, environment comparison, and independent review retain the CPython/package mismatch: the absolute at-least-3/5 gate passes, but the 4/5 versus 3/5 baseline comparison is ineligible. The earlier false composite gate remains diagnostic evidence.
  • Offered-C8 controls, seed 1801, and failed seed 2801: the last is performance-ineligible and closes the dependent C8 arm. The independent review confirms that the failed direct diagnostic stops C8 without by itself disqualifying the unchanged production-envelope C4 recipe.
  • The original selected batch-2,048/C4 performance cells are preserved for the dashboard importer at performance-dcp1[-repeat1|-repeat2]-{c1-32k-unique,c4-128k-unique,c4-128k-shared}.json. They remain separate cache/context groups with descriptive 16-request cell metrics; they do not rank batch 4,096, include primes, or incorporate C8.
  • The native-derived chart manifest, SVG, and graph data were rendered twice with identical bytes. They plot only those nine eligible cells and preserve their unique/shared cache separation.
  • First C4 reload window: sampled minimum free memory was 1,755 MiB on each GPU, maximum temperatures were 88/86 C, and the largest sample gap was 0.256 s. It covers warm-up, offered-C8 controls, and agentic work only; it is not a continuous minimum or the final SWE/64K window.
  • Search closure and sanitization receipts: scope and public-byte provenance for this stage.

Gaps retained for campaign close

Gate State Limit
Frozen SWE complete Absolute 4/5 passes at-least-3/5; paired 4/5 versus 3/5 comparison is ineligible.
64K diagnostic complete, bounded The native completion-budget failure remains; coherent truncation satisfies the separate diagnostic rule. One-attempt visible yield regressed versus the matched incumbent.
Routed dispatch-cap C4 and real client C4 boundary complete 28 routed preflight observations passed; the offered-eight diagnostic preserves two bounded admission timeouts and no C8 claim.
Restoration complete r11 deployment and exact r10 rollback are retained.
Native-client acceptance complete with format caveats All 13 semantic paths passed; strict-format failures remain explicit.
Deployment review complete The historical recommendation and pre-deployment promoted=false record are retained; r11 was subsequently selected and deployed.
Public post/readback complete Pages run succeeded; the merged public page returned HTTP 200 and Chrome confirmed its title and core published facts.

The shareable sanitized managed recipe is candidate-dcp1.public.toml. It preserves the managed schema, containment, environment, and flags while using the generic cache path and port. Its materialization prerequisite is documented in the prior recipe-reconstruction section; review a managed preview before any load.