Skip to content

Coverage and gaps

The bounded native-host transport comparison keeps the GLM model/image, TP2, context393216, C1, adaptive MTP, dtypes, parsers, router contract and cuMem0 fixed. Both arms ran smoke, JSON, needle, tools, long tools, streaming tools, tool result, Responses, image and OCR. Capacity uses fixed32-word outputs, strict validation and request canaries: 12 requests at a4096 target and3 at120000.

The380000 target capacity cell fails exact word count in both arms (33vs32); neither timing is eligible for a performance headline. Separate380000-target needle and long-tool assertions test high-context correctness. This does not qualify the full393216-token advertised boundary, C>1, sustained endurance, another runtime/model, or future GPU/topology/boot configurations.

The two-repetition marker suite is diagnostic only. The candidate case has0.5 raw marker pass rate; independent Sol review establishes a keyword false negative. Its failed artifact remains unchanged. The suite does not execute model-generated code and cannot establish broad coding quality.

The unique-prefix cache policy does not prove cold cache. Metadata is incomplete; short-cell scout warming and partial cache reporting are disclosed. No p99 or service-tail claim is made from these small populations.