Skip to content

Huihui Qwen3.8 NInfer 64K C1 qualification reference

Date: 2026-09-19. Status: qualified reference evidence; operator assignment and promotion records are private.

Result card

Setup Result
Profile Pinned NInfer 70434721, RTX 5090, NVFP4/mixed FP8 weights, INT8 group-64 KV, MTP3, 65,536-token context, C1, no thinking, vision
Functional Direct preflight 7/7; routed preflight 6/6; repeated quality core/session passed; vision 12/12
Client compatibility Native Hermes 9/9 with configured provider; tool-result continuation retained
Capacity Strict short input 12/12: 182.1 mean decode tokens/s; long input 6/6: 167.1, with 60,769-60,776 actual prompt tokens
Memory 24,224 MiB after workloads; 4,287 MiB below the managed 28,511 MiB allocation budget
Limits No endurance or transient peak measurement; no matched 64K no-spec control; interactive browser acceptance unverified

The 32K profile was below Hermes's 64,000-token minimum. The separate 64K profile resolved 65,536 KV tokens and passed a quality request with 62,455 actual prompt tokens. These are bounded long-context observations, not proof of maximum stable prompt. The historical GGUF control used less GPU memory; this qualification is a capability/speed tradeoff, not a footprint win.

Initial context-limit and queue-limit failures caused by omitted profile overrides remain in the bundle. The corrected 60K nominal needle and C1 checks passed. Capacity used 32 exact output words and unique prompts without request canaries. Small sample counts do not establish tail latency or a causal MTP speedup.

Public evidence describes qualification only. Operator authorization, route assignment, recovery selection, media state, client distribution, dashboard state, and fleet results are retained privately.

Evidence: bundle, manifest, sanitized recipe, and publication summary.