Coverage and remaining gates¶
| Requested outcome | Retained result | Limit / next gate |
|---|---|---|
| Skill before live work | Independent paired decision review, fresh transfer case and follow-up scenario-diff review accepted | Behavioral instruction tests, not model-quality measurements |
| Enhanced testing | Real CLI and synthetic HTTP regressions; exact response IDs, errors, deadlines and failed-result retention | Local vLLM tokenizer adapter only; client overlap is not scheduler proof |
| Investigate current crashes | Large serial control passed; overlapping 175K/32K reproduced engine death and two Xid31 faults | First faulty kernel unresolved; no CUDA core dump/sanitizer |
| Configuration improvement | DCP1 passed matched small and three matched large overlaps | Experimental mitigation, lower KV capacity, no promotion or reliability-rate claim |
| Hardware envelope | Measured startup KV pool and repeated 175K/32K point | 201K time guard refused; 310K, full-window C4 and cache churn unrun |
| Shareable recipe | Pinned parent/candidate, native failures/results, notebook and complete publication matrix | Operator must materialize exact weights and bind local identities; no full replacement qualification |
| Restoration | Exact original identity and exclusive mode restored; direct/routed smoke and JSON passed | Recovery does not fix the reproduced overlap fault |
The bounded diagnostic campaign is closed. Broader finalist performance, vision/OCR, real-client, agentic/SWE and natural long-output quality gates are not run for this candidate. Retain every failure and do not substitute historical baseline quality for candidate qualification. No chart pack applies: no comparable performance cells were measured.