Skip to content

Qwen3.8 PRO campaign recipes

Choose by hardware and workload, then inspect the recipe and its limits. These are published evidence decisions; each result applies to its recorded configuration.

Reviewed 2026-09-05. Comparison history · Container reconstruction · Measurement definitions.

How to compare these measurements

Request percentiles describe this sample, including at n=100; they do not establish stable production tails. TPOT and mean ITL in capacity-v3 are the same per-request average, not token-arrival intervals. Effective prefill includes queueing and first-output work. Compare only matched workload and configuration populations.

Primary Node / RTX PRO 6000 Blackwell Max-Q / TP=1 · Catalog topology annotation: TP=1 on one GPU

Inferact · SGLang TP1 K12/1K

challenger · qwen38-27b-inferact-nvfp4-pro6000-dflash2-k12-chunk1k

Prompt target 4096 tokens; C8; 100 requests; unique cache; request canaries=True.

TTFT P50
2,331.3 ms
TTFT P99
5,513.3 ms
Decode P50
182.5 tok/s
TPOT (mean ITL proxy) P50
5.43 ms/token
E2E P50
4,856.4 ms
Aggregate output
764.3 tok/s

n=100/100; failed=0; C8; prompt target=4096; configured context=262144; output cap=512; response word target=256; output policy=not recorded (legacy); shared prefix target=not recorded; cache=unique; thinking=disabled; sampling=not recorded (legacy).

Actual prompt tokens min/P50/max: 3,613.0 / 3,655.0 / 3,689.0. Actual completion tokens min/mean/max: 512.0 / 512.0 / 512.0.

Strengths

  • Selected one-card sustained-output and end-to-end latency arm in the matched campaign.
  • All 100 request canaries completed and passed.

Limits and untested work

  • No promotion; broad quality, routed clients, multimodal behavior, and complete power telemetry remain open.
  • The 82K/C8 stress result was non-interactive and is not a responsive long-context claim.
Evidence provenance

Source SHA-256: df868e28df0b8ae38ad985368e3d7423977f555b9994c9309d659bb5d55ed9b8. Metric values are generated from the linked native artifact. Recipe selection and strengths/limits are reviewed catalog annotations.

Recipe selector: Inferact/Qwen3.8-27B-NVFP4-PRO6000-DFlash2-K12-Chunk1K. Recipe SHA-256: f0cbb351810e67d0dab326ba3fa54bb0fb1f0f51709e514d0ae9ba3b8639080b.

Native identity: model=qwen38-27b-inferact-nvfp4-pro6000-dflash2-k12-chunk1k; engine=sglang; gpu=Primary Node / RTX PRO 6000 Blackwell Max-Q / TP=1; measurement_protocol=capacity-v3

NVIDIA RTX PRO 6000 Max-Q · Catalog topology annotation: TP=1 on one GPU

RadixArk · SGLang TP1 K8

challenger · qwen38-27b-radixark-nvfp4-sglang-pro6000-c8-dflash2

Prompt target 4096 tokens; C8; 100 requests; unique cache; request canaries=True.

TTFT P50
1,239.0 ms
TTFT P99
4,030.2 ms
Decode P50
122.5 tok/s
TPOT (mean ITL proxy) P50
8.13 ms/token
E2E P50
5,186.3 ms
Aggregate output
746.7 tok/s

n=100/100; failed=0; C8; prompt target=4096; configured context=262144; output cap=512; response word target=256; output policy=not recorded (legacy); shared prefix target=not recorded; cache=unique; thinking=disabled; sampling=not recorded (legacy).

Actual prompt tokens min/P50/max: 3,613.0 / 3,655.0 / 3,689.0. Actual completion tokens min/mean/max: 267.0 / 494.4 / 512.0.

Strengths

  • Lower time to first token than the selected Inferact sustained-output arm.
  • All 100 request canaries completed and passed.

Limits and untested work

  • Lower decode rate and higher end-to-end latency than the Inferact finalist on the matched workload.
  • No promotion; broad quality and client-path qualification remain open.
Evidence provenance

Source SHA-256: a7f65d4ce9b50c87befe503ae8a95d97d3c4d1b77b47a15513a17b5924b407c6. Metric values are generated from the linked native artifact. Recipe selection and strengths/limits are reviewed catalog annotations.

Recipe selector: RadixArk/Qwen3.8-27B-NVFP4-PRO6000-C8-DFlash2. Recipe SHA-256: a65340e4166c54c49d3299b5b6cf7f7f9695768dd7d67955ec3e0dab17784838.

Native identity: model=qwen38-27b-radixark-nvfp4-sglang-pro6000-c8-dflash2; engine=sglang; gpu=NVIDIA RTX PRO 6000 Max-Q; measurement_protocol=capacity-v3

Primary Node / RTX PRO 6000 Blackwell Max-Q / TP=1 · Catalog topology annotation: TP=1 on one GPU

kelnei · vLLM MTP2

challenger · qwen38-27b-kelnei-nvfp4-vllm0271-mtp2

Prompt target 4096 tokens; C8; 100 requests; unique cache; request canaries=True.

TTFT P50
483.3 ms
TTFT P99
3,424.8 ms
Decode P50
72.9 tok/s
TPOT (mean ITL proxy) P50
13.71 ms/token
E2E P50
7,858.4 ms
Aggregate output
503.4 tok/s

n=100/100; failed=0; C8; prompt target=4096; configured context=262144; output cap=512; response word target=256; output policy=not recorded (legacy); shared prefix target=not recorded; cache=unique; thinking=disabled; sampling=not recorded (legacy).

Actual prompt tokens min/P50/max: 3,613.0 / 3,655.0 / 3,689.0. Actual completion tokens min/mean/max: 310.0 / 509.3 / 512.0.

Strengths

  • Verified alternate-runtime speculative gain over the exact vLLM no-speculation control.
  • Smoke, JSON, Responses, shared-prefix tools, and all 100 canaries passed.

Limits and untested work

  • Did not beat the SGLang finalist on aggregate output or end-to-end latency.
  • No promotion; broad quality and routed-client gates remain open.
Evidence provenance

Source SHA-256: 4e81ab5f553dda7eea2aa5f2fe1d9553f0815b3fa64edadf21dd5f085dd6a4b1. Metric values are generated from the linked native artifact. Recipe selection and strengths/limits are reviewed catalog annotations.

Recipe selector: kelnei/Qwen3.8-27B-NVFP4-vLLM0271-MTP2-PRO6000. Recipe SHA-256: 5d6efbce2f5d7570d147961195b94ecb551b01a3bf632c4f221149048c5cfe0a.

Native identity: model=qwen38-27b-kelnei-nvfp4-vllm0271-mtp2; engine=vllm-0.27.1; gpu=Primary Node / RTX PRO 6000 Blackwell Max-Q / TP=1; measurement_protocol=capacity-v3

Primary Node / RTX PRO 6000 Blackwell Max-Q / TP=1 · Catalog topology annotation: TP=1 on one GPU

kelnei · vLLM no-spec control

no-promotion · qwen38-27b-kelnei-nvfp4-vllm0271-nospec

Prompt target 4096 tokens; C8; 100 requests; unique cache; request canaries=True.

TTFT P50
1,555.0 ms
TTFT P99
3,102.2 ms
Decode P50
46.2 tok/s
TPOT (mean ITL proxy) P50
21.64 ms/token
E2E P50
12,617.4 ms
Aggregate output
315.2 tok/s

n=100/100; failed=0; C8; prompt target=4096; configured context=262144; output cap=512; response word target=256; output policy=not recorded (legacy); shared prefix target=not recorded; cache=unique; thinking=disabled; sampling=not recorded (legacy).

Actual prompt tokens min/P50/max: 3,613.0 / 3,655.0 / 3,689.0. Actual completion tokens min/mean/max: 387.0 / 510.4 / 512.0.

Strengths

  • Exact same-checkpoint and same-runtime control for the vLLM MTP2 arm.
  • Smoke, JSON, Responses, shared-prefix tools, and all 100 canaries passed.

Limits and untested work

  • Substantially slower than the matched MTP2 arm on this workload.
  • Control evidence only; not a deployment recommendation.
Evidence provenance

Source SHA-256: 24e03aa1f10f729d27d5758e25d0d39a174730175ee6cfccfdd37f96078ced2c. Metric values are generated from the linked native artifact. Recipe selection and strengths/limits are reviewed catalog annotations.

Recipe selector: kelnei/Qwen3.8-27B-NVFP4-vLLM0271-NoSpec-PRO6000. Recipe SHA-256: a490aee8a48457651c6b2f32a54b5b73524c3a9b44631359710df2535d8fb6be.

Native identity: model=qwen38-27b-kelnei-nvfp4-vllm0271-nospec; engine=vllm-0.27.1; gpu=Primary Node / RTX PRO 6000 Blackwell Max-Q / TP=1; measurement_protocol=capacity-v3

2x NVIDIA RTX PRO 6000 Max-Q · Catalog topology annotation: TP=2 over PCIe without NVLink

Inferact · SGLang TP2 control (rejected)

rejected · qwen38-27b-inferact-nvfp4-pro6000-tp2-dflash2-k12-chunk1k

Prompt target 4096 tokens; C8; 100 requests; unique cache; request canaries=True.

TTFT P50
3,184.4 ms
TTFT P99
6,890.1 ms
Decode P50
150.2 tok/s
TPOT (mean ITL proxy) P50
6.56 ms/token
E2E P50
6,047.0 ms
Aggregate output
587.9 tok/s

n=100/100; failed=0; C8; prompt target=4096; configured context=262144; output cap=512; response word target=256; output policy=not recorded (legacy); shared prefix target=not recorded; cache=unique; thinking=disabled; sampling=not recorded (legacy).

Actual prompt tokens min/P50/max: 3,613.0 / 3,655.0 / 3,689.0. Actual completion tokens min/mean/max: 512.0 / 512.0 / 512.0.

Strengths

  • Retained as a matched topology control with complete latency distributions.
  • All 100 request canaries completed and passed.

Limits and untested work

  • Slower than one TP1 service and two independent TP1 replicas on this host.
  • Strict structured JSON failed twice with duplicated output and a leaked delimiter.
Evidence provenance

Source SHA-256: 91e65903e38ad9b9339d2973526f7e9437989e79bf1810dbb03d8fe5f08ba74b. Metric values are generated from the linked native artifact. Recipe selection and strengths/limits are reviewed catalog annotations.

Recipe selector: Inferact/Qwen3.8-27B-NVFP4-PRO6000-TP2-DFlash2-K12-Chunk1K. Recipe SHA-256: f0cbb351810e67d0dab326ba3fa54bb0fb1f0f51709e514d0ae9ba3b8639080b.

Native identity: model=qwen38-27b-inferact-nvfp4-pro6000-tp2-dflash2-k12-chunk1k; engine=sglang; gpu=2x NVIDIA RTX PRO 6000 Max-Q; measurement_protocol=capacity-v3

Catalog SHA-256: fa82f917ff8c16f22bc95d4bda32fd35f96dc39b1b49c9e278a3a9b04c5b035b.