Skip to content

Coverage and remaining gates

Requested outcome Retained result Limit / next gate
Skill before live work Independent paired decision review, fresh transfer case and follow-up scenario-diff review accepted Behavioral instruction tests, not model-quality measurements
Enhanced testing Real CLI and synthetic HTTP regressions; exact response IDs, errors, deadlines and failed-result retention Local vLLM tokenizer adapter only; client overlap is not scheduler proof
Investigate current crashes Large serial control passed; overlapping 175K/32K reproduced engine death and two Xid31 faults First faulty kernel unresolved; no CUDA core dump/sanitizer
Configuration improvement DCP1 passed matched small and three matched large overlaps Experimental mitigation, lower KV capacity, no promotion or reliability-rate claim
Hardware envelope Measured startup KV pool and repeated 175K/32K point 201K time guard refused; 310K, full-window C4 and cache churn unrun
Shareable recipe Pinned parent/candidate, native failures/results, notebook and complete publication matrix Operator must materialize exact weights and bind local identities; no full replacement qualification
Restoration Exact original identity and exclusive mode restored; direct/routed smoke and JSON passed Recovery does not fix the reproduced overlap fault

The bounded diagnostic campaign is closed. Broader finalist performance, vision/OCR, real-client, agentic/SWE and natural long-output quality gates are not run for this candidate. Retain every failure and do not substitute historical baseline quality for candidate qualification. No chart pack applies: no comparable performance cells were measured.