RTX 5090 benchmark view¶
Hardware: NVIDIA GeForce RTX 5090, 32,607 MiB, sm_120. Measured hosts: current isolated Windows 11 / Docker Desktop / WSL2 qualification lane plus historical Primary Node captures. Evidence reviewed through: 2026-09-19.
The historical GGUF measurement host exposed 32 GiB physical RAM and about 16.2 GiB to Docker/WSL, rather than the campaign's original 96 GiB planning assumption. The successful 262K result therefore did not rely on hidden host RAM or weight offload.
Side-by-side speed and recipe links for every configuration measured on this card: model comparison table. The newest Qwen-only matrix is the Qwen3.8 27B quant/speculation comparison.
Current qualification and historical services¶
The context measurements used one RTX 5090 without a co-resident media-generation workload. They do not establish concurrent model/media capacity. The qualified 160K profile separately passed vision 18/18.
The September 12 efficient-variant screen tested Signal 27B, Swift 27B, Qwopus Flash 27B, and Minitron 20B with pinned llama.cpp Q6_K/64K/C1 recipes, plus Signal/Swift no-spec controls. All four passed functional checks and tools 20/20. Strict ten-question thinking-on scores were 9/10, 9/10, 1/10, and 3/10, versus 8/10 for the existing 262K incumbent; format and budget failures count. Signal/Swift used about 16% fewer completion tokens in that sample, but output failures and missing full-contract gates prevented replacement. The exact incumbent was restored. See the dedicated dossier; prior runtime and quantization recommendations below are historical evidence, not rerun here.
The Huihui NInfer MTP3 160K/C1 profile is a qualified reference for this measured GPU. Direct correctness, vision, depth retrieval and unique-canary capacity passed; deployment state is private.
The August 17 RadixArk Qwen3.8 27B NVFP4 qualification covered text, tools, image, OCR, and native
video without involving an RTX PRO 6000. Older rows preserve measurements made
while a 5090 was installed in Primary Node; they are not current topology claims.
The later Sharp v22.1 chat-template A/B was also isolated to this direct lane;
it did not change the preferred stock-template challenger configuration.
The August 21 DFlash2 diagnosis was likewise isolated and direct. The exact
float32 selector and the best BF16 single-request arm both missed the retained
128K capacity contract, so neither changed the selected stock recipe.
The same-day MTP3/ReplaySSM A/B then measured a large decode gain but only
70,231 KV tokens and a 1.9% 64K end-to-end regression. The external recipe
refresh found no full-context profile with a clear combined speed, fidelity,
tool-use, and runtime-safety win.
The later managed llama.cpp GGUF campaign changed that capacity conclusion:
Unsloth Q4_0 plus its Q4_0 MTP head passed exact retrieval at 253,822 actual
prompt tokens while preserving an 8,192-token output reserve. It is now the
preferred 5090 FAST-TIER challenger. Promotion remains deferred pending the
independent SWE worker, a correct container health probe, and resolution of
runtime position warnings.
The August 22 real-client follow-up passed OpenClaw and Hermes routed identity,
fallback detection, and shell-tool/result continuation. It did not establish a
250K routed contract because the bounded route still declared the earlier 128K
SGLang/NVFP4 compatibility metadata and video admission. A truthful 262K,
image-only router fingerprint remains a promotion gate.
On September 19, an isolated direct 8K/C1 scout tested the pinned Huihui NInfer
NVFP4 candidate. It passed protocol checks and a corrected 12/12 image-only
subset, but strict tools were 0/3 at both default and low-effort diagnostics;
the incumbent control was 3/3. The original image attempt was 0/12 because
vision was disabled and is retained as configuration-failure evidence. No
matched speed, memory, or context comparison was run, so the candidate is
rejected/no-promotion; existing recommendations do not change. Managed
restoration is verified; immutable runtime recreation remains open. See the
dated finding.
A compatible version-2 runtime follow-up recovered the original argument
contract failure: smoke 2/2, core quality 9/9 in each off/low lane, C1 protocol
5/5, and vision 12/12 passed; the 20-request burst remained 17/20 because three
requests received HTTP 429. CRLF-to-LF boundary behavior remains. A 12/12
matched no-spec/MTP3 strict 32-word cells measure 71.4/176.2 mean decode tok/s.
A separate GGUF profile measured 107.7 tok/s but has a different checkpoint,
runtime, and quantization. Huihui MTP3 used 21,852 MiB post-workload versus
GGUF 19,022 MiB, so it misses the footprint requirement. MTP3 32K is unpaired.
Managed restoration is verified and the candidate remains no-promotion.
See the follow-up finding.
The compatible-runtime follow-up also passed managed recreation and bounded router/client compatibility. Six capacity artifacts were eligible for descriptive reporting; failed 128-word warmups remain excluded from performance rows and retained as failures. See the compatibility addendum.
On September 3, the exact NInfer NVFP4 artifact and runtime revision completed
a matched no-speculation/MTP3 direct qualification. MTP3 delivered 0.430-second
median TTFT, 165.9 tok/s median decode, and 0.720-second median E2E at 4K/C1,
versus 0.421 seconds, 75.3 tok/s, and 1.085 seconds for the control. It also
returned the exact marker from a 201,746-token API-reported prompt with an
8,192-token completion cap. This made it the preferred measured direct
text/tools performance challenger at that point, while GGUF remained the broader-capability
incumbent with image/OCR, agentic, endurance, routed-client, and deeper actual-
prompt evidence. NInfer remains no-promotion: shared-prefix tools were 17/20,
only 2,354 MiB remained free, and routed/client plus broader gates are open.
A later same-day twelve-arm quant/speculation bakeoff measured the idle card before changing the reserve: no GPU processes, 0 MiB reported used, and 32,187 MiB free. Gittensor's target-only RTX5090 NVFP4/SGLang profile then won warm TTFT at 50.9 ms, passed a 244,002-token actual prompt, and completed C2 at 128.7 aggregate output tok/s. Its advertised DSpark arm failed on incompatible target/draft shapes and its FP8 KV path used default 1.0 scales. It is now the preferred direct TTFT challenger, while GGUF remains the broad-capability incumbent. CometKim MTP3 reached 228.0 tok/s decode but failed strict tools 0/3. The final Unsloth Dynamic V3.0 NVFP4 MTP3 arm passed tools 20/20 and measured 137.7 tok/s short decode plus 127.5 tok/s at a 53,706-token prompt, with 3,198 MiB free. It is the strongest clean 64K speculative arm. The zero added reserve was a dedicated-campaign exception only.
The card previously had two operator-selectable shapes:
- an exclusive Nemotron Nano Omni 30B stack for auxiliary text, general vision, and OCR; or
- Qwen2.5-Omni 3B co-resident with dedicated Parakeet STT and Kokoro TTS.
Embeddings/reranking and ComfyUI remain separately admitted stacks. The PRO
6000 Primary may remain running during these tests, but that makes it
protected/co-resident, not measured.
On 2026-08-28, the isolated lane also qualified the first bounded ComfyUI generation candidates. FLUX.2 Klein 4B FP8 produced a decodable 512×512 PNG in 9.859 seconds at 12,919 MiB peak. Wan2.2 TI2V 5B produced a decodable 17-frame, 512×288 H.264 MP4 in 9.092 seconds at 18,263 MiB peak from a clean 943 MiB worker baseline. These are functional/capacity results only: both workflows remain unavailable pending independent perceptual review and no router or client path was exercised. A later exact-build cross-host pass exercised real Hermes MCP image/video jobs, A2A replay, artifact delivery, cold approval/unavailable behavior, and managed teardown. Two FLUX.2 images passed bounded independent review. The new Wan2.2 sample decoded correctly but failed spatial and prompt-adherence review. Both workflows remained unavailable at that point.
On 2026-09-15, a managed-media bringup left the worker healthy and running on the 5090 with its declared 4,096 MiB reserve. One fixed FLUX.2 high-profile image and Wan v1/v2 functional video jobs completed. Wan v1 independently failed all sampled frames; Wan v2 restored recognizable content but remained partial and unavailable. These defect-diagnostic runs are not comparable performance cells. See the media bringup finding.
The earlier August 28 production gate enabled only the exact FLUX.2 image workflow at
three fixed c1 profiles: 512×512, 768×768, and 1024×1024, all four steps. Real
Hermes produced one warm image at each size and five cold 512×512 regressions.
Six of eight passed strict independent review; two technically valid draft
samples retained origami/material and exact-count failures. The cold path
passed explicit approval, managed build/start, exact same-job resume, artifact
return, and teardown. Its 2,045.626-second incident E2E includes multiple
defects repaired during the run; the final server-issued exact-resume regression
measured 908.936 seconds E2E and 0.087 seconds generation. Its native MCP image
matched the authenticated artifact resource byte for byte. Wan2.2 v1 remained
quality_failed and unavailable; the September 15 v2 repair remains
quality_unverified and unavailable.
Last measured and challenger state¶
| Capability | Model | Decision | Boundary |
|---|---|---|---|
| Qualified text/tools/image reference | Huihui Qwen3.8 NInfer MTP3 160K/C1 | qualified 160K reference; retained 64K comparison | Direct preflight 8/8, vision 18/18, depth retrieval 9/9, unique short capacity 12/12 and long 3/3; 150,144 maximum measured input with 8,192 output allowance; no matched speed or endurance claim |
| Lowest direct text TTFT and long context | Qwen3.8 27B Gittensor RTX5090 target-only | preferred measured TTFT challenger, no-promotion |
Direct 262K/C2; 50.9 ms warm median TTFT, 79.5 tok/s decode, 244,002 actual prompt; DSpark shape failure and FP8 KV scale caveat open |
| Clean 64K speculative text lane | Qwen3.8 27B Unsloth Dynamic V3 NVFP4 MTP3 | preferred clean 64K speculative challenger, no-promotion |
388.7 ms TTFT; 137.7 tok/s short and 127.5 tok/s at 53,706 prompt tokens; tools 20/20; 3,198 MiB free |
| Text-to-image generation | FLUX.2 Klein 4B | historical qualification; simultaneous model/media capacity unmeasured | Direct and real-Hermes c1; fixed 512/768/1024-pixel, four-step profiles; 6/8 strict bounded reviews pass with two retained draft failures; arbitrary dimensions rejected |
| Text-to-video generation | Wan2.2 TI2V 5B | unavailable candidate, no-promotion |
v1 prompt/spatial review failed; September 15 v2 small/standard clips restored recognizable content, but temporal and broader quality remain unverified |
| High direct text decode and long context | Qwen3.8 27B NInfer NVFP4 + MTP3 | retained measured decode challenger, no-promotion |
Direct c1/252,928; 0.430-second median TTFT and 165.9 tok/s decode; 201,746 actual prompt + 8,192 output cap; shared-prefix tools 17/20; 2,354 MiB free; routed/client and broader gates open |
| Broader text/tools, long context, image/OCR | Qwen3.8 27B Unsloth Q4_0 + MTP3 | retained 5090 GGUF incumbent and broad-capability challenger, no-promotion |
Direct c1/262K and real-client short/tool route pass; 253,822 actual prompt + 8,192 output reserve; routed 250K blocked by stale 128K metadata; native video unsupported |
| Computer-use perception, vision, OCR, video | Qwen3.8 27B RadixArk NVFP4 | preferred 5090 challenger, no-promotion |
Direct only; c1/128K; eight images / two videos; GUI action loop and routed admission open |
| Auxiliary text, vision, OCR | Nemotron Nano Omni 30B | historical current topology |
Audio input not qualified |
| Smaller co-resident Omni | Qwen2.5-Omni 3B | challenger, no-promotion |
Text/image/OCR passed; audio was noisy |
| STT | Parakeet TDT 0.6B v3 | historical current |
Routed default at capture time |
| STT | Qwen3-ASR 0.6B | challenger, no-promotion |
Non-inferior WER margin; not routed |
| STT | Nemotron 3.5 ASR | rejected |
WER regression exceeded margin |
| TTS | Kokoro | historical current |
Co-resident voice component at capture time |
Comparable quality and capacity¶
| Model/configuration | Evidence | Measured result | Decision |
|---|---|---|---|
| Huihui NVFP4/mixed FP8, NInfer 70434721, INT8 KV, MTP3, 160K/C1 | functional, bounded quality, unique-canary capacity |
Direct 8/8, vision 18/18, depth retrieval 9/9; descriptive short 179.7767 / long 158.789 mean decode tokens/s (n=12 / n=3) | qualified reference; earlier 64K comparison; no matched speed or endurance claim |
| Huihui NVFP4/mixed FP8, NInfer 70434721, INT8 KV, MTP3, 64K/C1 | functional, bounded quality; non-comparative canary-free C1 capacity |
Direct7/7, routed6/6, vision12/12, Hermes9/9; descriptive short182.1/long167.1 mean decode tokens/s (n12/n6); 24,224 MiB post-workload | historical comparison; no 64K matched no-spec or endurance claim |
| Qwen3.8 27B Gittensor RTX5090 NVFP4 target-only, SGLang | functional, capacity, matched performance, bounded deterministic quality |
50.9/277.0 ms warm TTFT p50/p95; 79.5 tok/s decode; 244,002-token actual prompt; C2 128.7 aggregate tok/s; coding/triage/tools 3/3 | preferred direct TTFT challenger, no-promotion; advertised DSpark incompatible; FP8 KV scales and broader/routed gates open |
| Qwen3.8 27B CometKim NVFP4Full MTP3, NInfer | functional, capacity, matched performance, negative bounded quality |
326.3 ms TTFT, 228.0 tok/s warm decode, 244,002-token actual prompt; coding/triage 3/3; strict tools 0/3 | decode research lead only; rejected as general-purpose winner |
| Qwen3.8 27B Unsloth Dynamic V3 NVFP4 MTP3, vLLM | functional, capacity, matched performance |
388.7 ms TTFT, 137.7 tok/s warm decode, 127.5 tok/s at 53,706 prompt tokens, C2 70.5 aggregate tok/s, tools 20/20, 83.2% token acceptance | strongest clean 64K speculative arm; pinned-runtime and broader-quality/routed gates open |
| FLUX.2 Klein 4B FP8, ComfyUI v0.33.4 | functional, capacity, routed/client acceptance, bounded quality, production cutover |
Direct decodable PNG at 9.859 s/12,919 MiB peak; real-Hermes fixed 512/768/1024-pixel profiles; 6/8 strict bounded reviews pass and two draft fidelity/count failures are retained; warm gateway E2E 1.242–1.650 s; exact cold resume/teardown pass; c1 | historically qualified at three profiles; simultaneous model/media capacity unmeasured |
| Wan2.2 TI2V 5B FP16, ComfyUI v0.33.4 | functional, capacity, routed/client acceptance, negative bounded quality |
Direct decodable MP4 at 9.092 s/18,263 MiB peak; exact-build real-Hermes 117,738-byte H.264 decode pass; prompt/spatial review fail; c1 | unavailable candidate, no-promotion; revise as a new workflow version |
| Qwen3.8 27B NInfer NVFP4 + MTP3 | functional, capacity, matched performance, bounded deterministic quality |
MTP3 versus no spec: 0.430/0.421-second median TTFT, 165.9/75.3 tok/s decode, and 0.720/1.085-second E2E; 201,746-token prompt plus 8,192 output cap; coding/triage/tools 3/3 each; shared-prefix burst 17/20; 29,834 MiB used / 2,354 free | retained direct decode challenger, no-promotion; ordinary reserve, admission, runtime-image, routed/client, multimodal, broad agentic/SWE, and endurance gates open |
| Qwen3.8 27B Unsloth GGUF Q4_0 + MTP3, llama.cpp | functional, capacity, matched performance, bounded deterministic quality |
253,822 actual prompt tokens plus 8,192 reserve; tools 20/20 and long-tools at 110,875; agentic 16/18; neutral 101-turn endurance 3/3; images 18/18; 104.1 tok/s short decode; 25,408 MiB startup VRAM | retained GGUF incumbent and broad-capability challenger, no-promotion; Q6_K+same-MTP infeasible under reserve policy |
| Qwen3.8 27B RadixArk NVFP4, SGLang TP=1 | functional, bounded deterministic quality |
119,675-token retrieval and tools 20/20; direct image/OCR/video; media 30/30; count boundaries 4/4 at eight images / two videos; 20.14 GB weights and 3,928 MiB free after startup | challenger, no-promotion |
| Qwen3.8 27B RadixArk NVFP4 plus MTP3/ReplaySSM | functional, capacity, matched performance |
Decode 138.19 tok/s at 4K and 117.10 at 64K versus 76.54/69.75 baseline; tools 20/20; only 70,231 KV tokens; 64K E2E 1.9% slower | rejected as 128K replacement; exact baseline restored, no-promotion |
| Qwen3.8 27B RadixArk NVFP4 plus DFlash2 | functional, capacity |
Exact float32 selector: 24,347 KV tokens and 24,341-token input ceiling. Best BF16/single-slot/no-radix/no-prefill-graph arm: 70,262 KV tokens, 49,549-token retrieval pass, tools 20/20; fixed DFlash intermediate BF16 SSM state 1.12 GB | Both arms rejected as 128K replacements; short-context lead only, no-promotion |
| Qwen3.8 27B RadixArk NVFP4, Sharp v22.1 template A/B | functional, bounded diagnostic quality |
Sharp preflight passed; thinking-enabled MMLU-Pro stayed 24/30 with +10.8% tokens and +10.7% latency; thinking-disabled behavior used 5.1% fewer tokens but passed 15/18 versus stock 18/18 | Sharp rejected; stock retained |
| Nemotron Nano Omni 30B NVFP4 | functional, capacity |
Text/JSON/4K retrieval/tools plus image/OCR passed; 65,536 context, 2 sequences | current topology |
| Qwen2.5-Omni 3B | functional, capacity |
Text/image/OCR and basic audio input passed; 24,576 MiB reservation | no-promotion |
| Parakeet, sequential | quality, capacity |
3.343% micro-WER; p95 177.87 ms | current |
| Qwen3-ASR, sequential | quality, capacity |
3.621% micro-WER; p95 113.58 ms | challenger, no-promotion |
| Nemotron 3.5 ASR, sequential | quality, capacity |
6.685% micro-WER; p95 225.45 ms | rejected |
Historical controls¶
| Model / configuration | Retained result | Limitation / decision |
|---|---|---|
| Qwen3.5 35B-A3B Q4_K_M | July 11 tools 20/20; 59,053-token actual prompt; 56.279 aggregate tok/s in a ten-request short probe | Missing immutable model/runtime pins; historical no-promotion |
| Gemma 4 E2B W4A16 | July 16 preflights through 120K passed; strict timeout triage 0/3 | no-promotion; short-generation capacity is not controlled decode |
| Gemma 4 E4B Fast | July 13 low-latency/router control | Historical selection, not current deployment |
| Gemma 4 31B | July Fast-lane load failed | Separate from the larger-card measurements in its dossier |
Later 5090 work established the Omni shapes above. Older voice research remains in the chronological findings index.
Run history¶
See RTX 5090 runs. The August 28 ComfyUI media qualification, exact-build live validation, and Hermes image production enablement; the September 3 NInfer qualification; August 21 GGUF qualification, recipe research, and DFlash2 diagnosis; August 20 template A/B; and August 17 Qwen rows classify the PRO relationship as unrelated. The July 28 ASR rows explicitly classify the PRO 6000 as protected.
2026-09-19: Huihui 64K C1 promotion¶
The earlier Huihui NInfer MTP3 qualification used 65,536 context tokens at C1. Direct preflight passed 7/7, routed preflight 6/6, vision 12/12, and native Hermes 9/9 with the exact local alias. Descriptive, canary-free C1 short-input capacity passed 12/12 at 182.1 mean decode tokens/s; the long-input cell passed 6/6 at 167.1, with 60,769-60,776 actual prompt tokens. Post-workload GPU use was 24,224 MiB. The 32K results remain historical evidence. Endurance, interactive browser acceptance, and a matched 64K no-spec control remain unmeasured. See the 64K finding.
2026-09-19: Huihui direct-I/O 160K context envelope (accepted)¶
The earlier 64K/C1 results remain a separate comparison. A separate final baked 163,840-context/C1 direct-I/O profile passed startup, direct preflight 8/8, and repeated exact retrieval 9/9 at 150,058, 150,144, and 150,124 prompt tokens; the maximum retained input is 150,144 with an 8,192-token output allowance and approximately 4.5 GiB GPU free at startup. It used MTP3 with INT8 KV and an 8,192-token visible-output cap. Startup minima were 4,606 MiB GPU free and 14,165.406 MiB Windows available; long-capacity minima were 4,598 MiB and 13,669.93 MiB respectively. The first retrieval of each case was cold or partially cached and later attempts warm; this does not support a throughput or cold-latency comparison. The subsequent direct-I/O verifier avoided the observed host-cache pressure; these sequential runs do not establish exclusive causality. The tested reference uses 163,840 context tokens and an 8,192-token output allowance; bounded native client compatibility is recorded separately. See the context-envelope finding.