Skip to content

Qwen3.8 27B PRO 6000 comprehensive campaign evidence

This sanitized bundle supports the dated finding.

Campaign boundary

  • Campaign: 2026-09-04-qwen38-27b-pro6000-possibility
  • Measured hardware: two equal RTX PRO 6000 Blackwell Max-Q GPUs under Docker Desktop/WSL2; single TP1, dual independent TP1, and TP2 measured
  • Measurement: direct online streaming; natural-completion screening plus a matched 100-request unique-canary sustained-output workload
  • Decision: DP2 bounded aggregate-throughput winner; TP2 rejected; no promotion
  • Not measured: broad quality, agentic/SWE, multimodal, routed/client, load-balancer/failover, and complete power/energy telemetry

Common campaign artifacts

Headline matched artifacts

Each headline artifact contains request-level timings plus mean, p50, p95, p99, confidence interval, and standard deviation for TTFT, effective prefill, decode, TPOT/mean ITL, and E2E. Aggregate throughput is output tokens divided by the concurrent run wall clock.

Functional and correctness artifacts

TP2's performance artifact completed 100/100 canaries, but its strict JSON gate emitted a duplicate object around a literal closing think delimiter and failed again in isolation. It is rejected.

SGLang optimization artifacts

The original target-only/DFlash2 K8 matrix remains retained as the baseline; see files prefixed nospec-, nospec-mamba96-, and dflash-k8-.

Topology, checkpoint, and alternate-runtime artifacts

  • TP2: files prefixed opt-tp2-k12-chunk1k-
  • DP2: files prefixed dp2-
  • RadixArk target-only/K8/K12: files prefixed radixark-
  • kelnei/vLLM MTP2 and no-spec: files prefixed kelnei-vllm0271-

The 82K/C8 artifacts are negative interactive-latency evidence. They are not warm-prefix results and must not be used to claim responsive long-context C8.

Graphs and restoration

The graph is derivative; raw JSON is authoritative. Publication does not authorize a serve, route, client, or deployment change.