Skip to content

Mixed 3.5-bpw GLM: repaired loader and C4 evaluation

The repaired model ran at C1 and C4. Functional checks and repeated bounded quality passed; long retrieval passed at 216,307 actual input tokens. Strict capacity was inconsistent. The baseline was restored; no promotion occurred.

Primary Node ran the managed trials; Mini was not involved. Public artifacts redact operator paths, endpoint identities, container IDs, and GPU UUIDs. Native benchmark artifacts retain their full schemas, visible outputs, per-request validation, timings, and failures. Only endpoint and operator identities are redacted. The recipe remaps its endpoint port to a synthetic example value while preserving the load-bearing model, runtime, memory, and scheduling settings.