Skip to content

Qwen3.5 122B

Current status and review date

Decision snapshot

  • Product role: Retained qualified, former Primary rollback and promotion-era evidence; not the immediate selected deployment.
  • Selected or best-qualified configuration: NVIDIA ModelOpt NVFP4 on NGC vLLM 26.06 with BF16 KV, 262,144 served tokens, concurrency one, one image, no video, and MTP disabled. The single-card profile carries the dated rollback label; the matched TP=2 lane is no-promotion.
  • Measured hardware: One- and two-card RTX PRO 6000 lanes on Fakoli Dark; TP=2 used two equal cards over PCIe without NVLink.
  • Evidence: Image/OCR, tools, repeated protocol-v3, 240K retrieval, and matched c1 evidence; TP=2 added 125,444-token and 32K measurements.
  • Decision: Retain the qualified recipe and its historical rollback label; do not present it as the immediate restore target or auto-promote the TP=2 result.
  • Important limitation: Near-ceiling prefill is slow, concurrency two was not qualified, and video failed before inference because the tested image lacked an H.264 decoder.
  • Review dates: Retained evidence cutoff: 2026-08-01. Dossier-format review: 2026-08-31.

Review narrative

2026-07-10 — Initial NVFP4 candidate

The initial NVIDIA NVFP4 candidate established Qwen3.5 122B as a large-model lead on one RTX PRO 6000. External NVIDIA and dated community recipes informed the candidate, but only local measurements count as qualification evidence.

2026-07-12 — MXFP4 control remained no-promotion

The historical MXFP4 lane used vLLM nightly, compressed-tensors MXFP4 with a Marlin fallback, FP8 KV, and 131,072 served tokens. Its result did not justify promotion and remains a historical control.

2026-07-28–29 — Single-card qualification and promotion-era comparison

The single-card NVIDIA lane qualified image/OCR, tools, repeated protocol-v3, long-context retrieval, and near-ceiling capacity with BF16 KV and c1 admission. In the matched 2026-07-29 lane it passed 240K retrieval, 20/20 tools, and 12/12 image attempts; the 231,426-token prompt measured 68.91 seconds TTFT and 60.3 tok/s decode. Agents-A1 later won that bounded comparison and passed the complete repeated gate. In that dated campaign, Qwen remained a managed rollback alongside other recovery profiles; the label is retained as history, not as a claim about the current live route.

2026-08-01 — Dual-PRO topology qualification

The exclusive TP=2 campaign preserved the ModelOpt NVFP4 and BF16-KV controls while changing the distributed topology where possible. Smoke, JSON, 30K retrieval, tools, and the repeated quality gate passed. A 125,444-token prompt measured 14.59 seconds TTFT, 8,570 effective prefill tok/s, and 65.0 tok/s decode; the 32K lane measured 2.32 seconds TTFT and 67.5 tok/s decode. This is additional topology evidence, not a promotion.

Immutable identity

NVIDIA NVFP4

  • Checkpoint: nvidia/Qwen3.5-122B-A10B-NVFP4.
  • Revision: 98915d837c4e7c87ac8296d02e89de19b3207e6d.
  • Executed NGC vLLM 26.06 image digest: Not recorded in this dossier.

Historical MXFP4

  • Checkpoint: olka-fi/Qwen3.5-122B-A10B-MXFP4.
  • Revision: 345839ea666a70f5035672f7c88afcba6281921f.
  • Executed vLLM-nightly image digest: Not recorded in this dossier.

Tested hardware and topology

Single-card lane

  • Host label: Primary Node.
  • Hardware: one RTX PRO 6000.
  • Admission: one request.

Dual-card lane

  • Hardware: two equal RTX PRO 6000 cards.
  • Topology: exclusive TP=2 over PCIe without NVLink.
  • Co-residency: every other inference workload was offline.

Other accelerator products were not tested by these retained campaigns.

Engine, quantization, KV, context, and concurrency recipe

Retained single-card NVFP4 profile

  • Runtime: NVIDIA NGC vLLM 26.06.
  • Quantization: ModelOpt NVFP4.
  • KV cache: BF16.
  • Served context: 262,144 tokens.
  • Admission: one sequence, one image, and no video.
  • Speculation: MTP disabled.

The public serve-recipe registry retains the reconstructable single-card profile.

Qualified TP=2 lane

  • Runtime and model controls: same NGC 26.06 ModelOpt NVFP4/BF16-KV lane.
  • Topology: TP=2 across both RTX PRO 6000 cards.
  • Admission: one running request.
  • Decision: no-promotion.

The public TP=2 campaign registry records this experiment lane.

Historical MXFP4 control

  • Runtime: vLLM nightly.
  • Quantization: compressed-tensors MXFP4 with Marlin fallback.
  • KV cache: FP8.
  • Served context: 131,072 tokens.
  • Maximum sequences: two.

Evidence by measurement class

External prior

NVIDIA's model card and dated community single-card recipes informed the candidate. They are external-prior, not local measurement.

Single-card functional, capacity, and quality

  • Image/OCR: pass; 12/12 image attempts in the matched lane.
  • Tools: 10/10 in qualification and 20/20 in the matched lane.
  • Repeated protocol-v3: pass.
  • Retrieval: 128K and 240K passes.
  • At 231,426 prompt tokens: 68.91 seconds TTFT and 60.3 tok/s decode.

Dual-card functional, capacity, quality, and performance

  • Smoke, JSON, 30K retrieval, and tools: pass.
  • Repeated gate: intelligence 6/6, session 3/3, tools 3/3.
  • At 125,444 prompt tokens: 14.59 seconds TTFT, 8,570 effective prefill tok/s, and 65.0 tok/s decode.
  • At 32K: 2.32 seconds TTFT and 67.5 tok/s decode.

Decision and promotion state

Retained qualified recipe

  • The single-card profile retains the rollback label from its dated campaign.
  • The TP=2 result is additional topology evidence and remains no-promotion.
  • The MXFP4 control remains no-promotion.

Current-state boundary

The latest public benchmark index classifies Qwen3.5 as retained qualified or promotion-era evidence, not the immediate selected deployment. Documentation does not authorize a route, deployment, or restoration change.

Failures and gotchas

Runtime and media

  • Preserve BF16 KV, c1 admission, the one-image limit, and the no-video boundary for the qualified NVFP4 profile.
  • When video was enabled in isolation, the exact NGC 26.06 image lacked an H.264 decoder. Every video-containing corpus request failed before model inference. This is a runtime packaging failure, not proof that the model architecture cannot understand video.

Performance and evidence boundaries

  • Near-ceiling prefill is slow.
  • Concurrency two was not tested as a qualified lane.
  • Controlled long-generation decode was not retained.
  • The executed runtime image digests are not recorded in this dossier.

Dated run history