Coverage and gaps¶
| Request or gate | Evidence and outcome |
|---|---|
| Exact migrated configuration | Identity; matching model/image, explicit NCCL and host-stack differences |
| Historical speed comparison | Four matched prompt cells, C1 n3 each; descriptive variable-output baseline |
| Controlled output and unique cache | Strict scout 0/3; unique natural 1/10 canary compliant; dependent cells stopped |
| Direct protocol gates | Disabled 9/10; enabled passes |
| Coding quality | Five cases x3, 15/15 |
| Image/OCR corpus | Six cases x2, 12/12 |
| Extended context | Native sweep 128/150; 9 empty length-terminated and 13 incorrect visible answers; non-monotonic curve |
| Agentic recovery and long sessions | Deep profile, 30/30 |
| Matched repository-agent smoke | Official SWE grader, resolved 1/1; 22 requests versus historical 11 |
| Sustained reliability | Endurance, 60/60 |
| Routed protocol gates | Disabled 9/10; nominal 380K HTTP413; 260K passes; enabled reasoning evidence fails projection |
| Post-run state | Verified unchanged serving state; routed smoke/JSON passed |
The Windows arm is retained evidence rather than a newly rerun installation. Strict controlled-output performance qualification did not pass. Cache state, driver and transport prevent an OS-only causal claim. C1 is the deployed engine ceiling; higher concurrency, full SWE-bench, video and speculation retuning are outside this deployed-profile comparison. No profile promotion is authorized.