Qwen3.5 122B¶
Current status and review date¶
Decision snapshot
- Product role: Retained qualified, former Primary rollback and promotion-era evidence; not the immediate selected deployment.
- Selected or best-qualified configuration: NVIDIA ModelOpt NVFP4 on
NGC vLLM 26.06 with BF16 KV, 262,144 served tokens, concurrency one, one
image, no video, and MTP disabled. The single-card profile carries the
dated
rollbacklabel; the matched TP=2 lane isno-promotion. - Measured hardware: One- and two-card RTX PRO 6000 lanes on Fakoli Dark; TP=2 used two equal cards over PCIe without NVLink.
- Evidence: Image/OCR, tools, repeated protocol-v3, 240K retrieval, and matched c1 evidence; TP=2 added 125,444-token and 32K measurements.
- Decision: Retain the qualified recipe and its historical rollback label; do not present it as the immediate restore target or auto-promote the TP=2 result.
- Important limitation: Near-ceiling prefill is slow, concurrency two was not qualified, and video failed before inference because the tested image lacked an H.264 decoder.
- Review dates: Retained evidence cutoff: 2026-08-01. Dossier-format review: 2026-08-31.
Review narrative¶
2026-07-10 — Initial NVFP4 candidate¶
The initial NVIDIA NVFP4 candidate established Qwen3.5 122B as a large-model lead on one RTX PRO 6000. External NVIDIA and dated community recipes informed the candidate, but only local measurements count as qualification evidence.
2026-07-12 — MXFP4 control remained no-promotion¶
The historical MXFP4 lane used vLLM nightly, compressed-tensors MXFP4 with a Marlin fallback, FP8 KV, and 131,072 served tokens. Its result did not justify promotion and remains a historical control.
2026-07-28–29 — Single-card qualification and promotion-era comparison¶
The single-card NVIDIA lane qualified image/OCR, tools, repeated protocol-v3, long-context retrieval, and near-ceiling capacity with BF16 KV and c1 admission. In the matched 2026-07-29 lane it passed 240K retrieval, 20/20 tools, and 12/12 image attempts; the 231,426-token prompt measured 68.91 seconds TTFT and 60.3 tok/s decode. Agents-A1 later won that bounded comparison and passed the complete repeated gate. In that dated campaign, Qwen remained a managed rollback alongside other recovery profiles; the label is retained as history, not as a claim about the current live route.
2026-08-01 — Dual-PRO topology qualification¶
The exclusive TP=2 campaign preserved the ModelOpt NVFP4 and BF16-KV controls while changing the distributed topology where possible. Smoke, JSON, 30K retrieval, tools, and the repeated quality gate passed. A 125,444-token prompt measured 14.59 seconds TTFT, 8,570 effective prefill tok/s, and 65.0 tok/s decode; the 32K lane measured 2.32 seconds TTFT and 67.5 tok/s decode. This is additional topology evidence, not a promotion.
Immutable identity¶
NVIDIA NVFP4¶
- Checkpoint:
nvidia/Qwen3.5-122B-A10B-NVFP4. - Revision:
98915d837c4e7c87ac8296d02e89de19b3207e6d. - Executed NGC vLLM 26.06 image digest: Not recorded in this dossier.
Historical MXFP4¶
- Checkpoint:
olka-fi/Qwen3.5-122B-A10B-MXFP4. - Revision:
345839ea666a70f5035672f7c88afcba6281921f. - Executed vLLM-nightly image digest: Not recorded in this dossier.
Tested hardware and topology¶
Single-card lane¶
- Host label: Primary Node.
- Hardware: one RTX PRO 6000.
- Admission: one request.
Dual-card lane¶
- Hardware: two equal RTX PRO 6000 cards.
- Topology: exclusive TP=2 over PCIe without NVLink.
- Co-residency: every other inference workload was offline.
Other accelerator products were not tested by these retained campaigns.
Engine, quantization, KV, context, and concurrency recipe¶
Retained single-card NVFP4 profile¶
- Runtime: NVIDIA NGC vLLM 26.06.
- Quantization: ModelOpt NVFP4.
- KV cache: BF16.
- Served context: 262,144 tokens.
- Admission: one sequence, one image, and no video.
- Speculation: MTP disabled.
The public serve-recipe registry retains the reconstructable single-card profile.
Qualified TP=2 lane¶
- Runtime and model controls: same NGC 26.06 ModelOpt NVFP4/BF16-KV lane.
- Topology: TP=2 across both RTX PRO 6000 cards.
- Admission: one running request.
- Decision:
no-promotion.
The public TP=2 campaign registry records this experiment lane.
Historical MXFP4 control¶
- Runtime: vLLM nightly.
- Quantization: compressed-tensors MXFP4 with Marlin fallback.
- KV cache: FP8.
- Served context: 131,072 tokens.
- Maximum sequences: two.
Evidence by measurement class¶
External prior¶
NVIDIA's model card and dated community single-card recipes informed the
candidate. They are external-prior, not local measurement.
Single-card functional, capacity, and quality¶
- Image/OCR: pass; 12/12 image attempts in the matched lane.
- Tools: 10/10 in qualification and 20/20 in the matched lane.
- Repeated protocol-v3: pass.
- Retrieval: 128K and 240K passes.
- At 231,426 prompt tokens: 68.91 seconds TTFT and 60.3 tok/s decode.
Dual-card functional, capacity, quality, and performance¶
- Smoke, JSON, 30K retrieval, and tools: pass.
- Repeated gate: intelligence 6/6, session 3/3, tools 3/3.
- At 125,444 prompt tokens: 14.59 seconds TTFT, 8,570 effective prefill tok/s, and 65.0 tok/s decode.
- At 32K: 2.32 seconds TTFT and 67.5 tok/s decode.
Decision and promotion state¶
Retained qualified recipe¶
- The single-card profile retains the
rollbacklabel from its dated campaign. - The TP=2 result is additional topology evidence and remains
no-promotion. - The MXFP4 control remains
no-promotion.
Current-state boundary¶
The latest public benchmark index classifies Qwen3.5 as retained qualified or promotion-era evidence, not the immediate selected deployment. Documentation does not authorize a route, deployment, or restoration change.
Failures and gotchas¶
Runtime and media¶
- Preserve BF16 KV, c1 admission, the one-image limit, and the no-video boundary for the qualified NVFP4 profile.
- When video was enabled in isolation, the exact NGC 26.06 image lacked an H.264 decoder. Every video-containing corpus request failed before model inference. This is a runtime packaging failure, not proof that the model architecture cannot understand video.
Performance and evidence boundaries¶
- Near-ceiling prefill is slow.
- Concurrency two was not tested as a qualified lane.
- Controlled long-generation decode was not retained.
- The executed runtime image digests are not recorded in this dossier.