{
  "campaign_id": "2026-09-13-intelligence-context-scout",
  "complete": true,
  "editorial_note": "2026-09-17: corrected stale derivative publication prose. Retained promotion, bounded client acceptance, restoration, and post-promotion context receipts were already present; original native scores and historical outcomes are unchanged.",
  "decision": {
    "decision_labels": [
      "current"
    ],
    "evidence_labels": [
      "quality",
      "external-prior"
    ],
    "next_gate": "None for this closed qualification snapshot; preserve rollback and evaluate new candidates separately.",
    "promotion_authorization_scope": "The user authorized disruptive GPU turnover, full evaluation, and promotion work; retained receipts record the resulting selected-current promotion, bounded client acceptance, and post-promotion context evidence.",
    "promotion_authorized": true,
    "promotion_state": "promoted, admitted, and client-accepted"
  },
  "headline_results": [
    {
      "candidate": "GLM Flash EXL3 r6",
      "caveat": "5 exact-choice failures; 4 reasoning-budget exhaustions without visible answer; four-way dispatch is not speed evidence",
      "completion": "complete fixed population, one pass",
      "official_raw_count": "91/100"
    },
    {
      "candidate": "GLM Flash EXL3 r7 route2",
      "caveat": "9 deterministic exact-choice failures; 100 nonempty FINAL replies, all stop; 0 blank, length, or transport failures. Strict120 route2 is matched to a no-spec C4 control: 120/120 performance-eligible in each lane, one-pair median decode 59.662 versus 38.908 tok/s; not a general speed claim.",
      "completion": "complete fixed population, separate one pass",
      "official_raw_count": "91/100"
    },
    {
      "candidate": "GLM Flash EXL3 r7 no-speculation",
      "caveat": "10 deterministic exact-choice failures; 100 nonempty FINAL replies, all stop; 0 blank, length, or transport failures. Native deep agentic completed 30/30; frozen five-case official SWE completed 4/5 resolved, 5/5 graded, 0 errors. The small coding scout does not establish a broad coding ranking.",
      "completion": "complete fixed population, separate one pass",
      "official_raw_count": "90/100"
    },
    {
      "candidate": "Qwen Flash-Next EXL3",
      "caveat": "9 exact-choice failures, each finish stop; 0 blank, reasoning-budget, or transport failures; configuration differs from GLM",
      "completion": "complete fixed population, one pass",
      "official_raw_count": "91/100"
    },
    {
      "candidate": "Qwen3.8 27B FP8 baseline",
      "caveat": "not a 100-item result; stopped before remaining seven shards",
      "completion": "partial: three retained shards",
      "official_raw_count": "26/30"
    }
  ],
  "identity": {
    "repository_dirty_state": "selected-current publication snapshot with retained promotion and post-promotion evidence",
    "repository_revision": "3c1bd508f638447fabc0ba2cfef9bc20b84ef9c1"
  },
  "limitations": [
    "The selected-current record remains bounded: a broad matched candidate comparison and fresh boot/reboot proof are not retained. GLM r7 MTP3 and no-spec strict120 each completed 120/120 performance-eligible C4 requests; median decode was 59.662 versus 38.908 tok/s in one matched pair, not a general speed claim. No-spec native context passed 9/9 after managed cold reload at >321K total reserved tokens, but runtime cache allocation does not prove concurrent full-window correctness. Its full100 scout is 90/100; native deep agentic completed 30/30. The first SWE wrapper failed before model requests because its cached environment inventory mismatched. The frozen retry completed 4/5 resolved; the five-case scout remains bounded. The declared Qwen FP8 rollback was not exercised after final promotion. The earlier MTP3 258K context stage passed 7/9 with two unresolved retrieval failures. Next has configuration-specific 9/9 context, 21/30 agentic, 4/5 SWE, and 12/12 image/OCR observations.",
    "Pinned item 8310 has no valid listed answer; official raw counts are retained without adjustment. A separate sensitivity would subtract one r7 and Next success and remove one r6 operational blank failure.",
    "69 sanitized native artifacts plus one derived environment diagnosis are retained with source hashes and a field-redaction audit; private originals remain outside the public bundle.",
    "Next strict120 retained 120 stop responses but zero performance-eligible responses because a direct-stream diagnostic found two leading line feeds before every canary. This is a client/harness formatting-contract failure, not a model attribution."
  ],
  "schema": "anvil-serving.benchmark-decision-summary/v1",
  "workload": {
    "completion_cap_tokens": 65536,
    "priority_order": [
      "intelligence",
      "stability",
      "usable context",
      "speed"
    ],
    "repetitions": 1,
    "suite": "Pinned 100-item stratified MMLU-Pro test selection"
  }
}
