Skip to content

RTX 5090 benchmark view

Hardware: NVIDIA GeForce RTX 5090, 32,607 MiB, sm_120. Measured hosts: current isolated Windows 11 / Docker Desktop / WSL2 qualification lane plus historical Primary Node captures. Evidence reviewed through: 2026-09-19.

The historical GGUF measurement host exposed 32 GiB physical RAM and about 16.2 GiB to Docker/WSL, rather than the campaign's original 96 GiB planning assumption. The successful 262K result therefore did not rely on hidden host RAM or weight offload.

Side-by-side speed and recipe links for every configuration measured on this card: model comparison table. The newest Qwen-only matrix is the Qwen3.8 27B quant/speculation comparison.

Current qualification and historical services

The context measurements used one RTX 5090 without a co-resident media-generation workload. They do not establish concurrent model/media capacity. The qualified 160K profile separately passed vision 18/18.

The September 12 efficient-variant screen tested Signal 27B, Swift 27B, Qwopus Flash 27B, and Minitron 20B with pinned llama.cpp Q6_K/64K/C1 recipes, plus Signal/Swift no-spec controls. All four passed functional checks and tools 20/20. Strict ten-question thinking-on scores were 9/10, 9/10, 1/10, and 3/10, versus 8/10 for the existing 262K incumbent; format and budget failures count. Signal/Swift used about 16% fewer completion tokens in that sample, but output failures and missing full-contract gates prevented replacement. The exact incumbent was restored. See the dedicated dossier; prior runtime and quantization recommendations below are historical evidence, not rerun here.

The Huihui NInfer MTP3 160K/C1 profile is a qualified reference for this measured GPU. Direct correctness, vision, depth retrieval and unique-canary capacity passed; deployment state is private. The August 17 RadixArk Qwen3.8 27B NVFP4 qualification covered text, tools, image, OCR, and native video without involving an RTX PRO 6000. Older rows preserve measurements made while a 5090 was installed in Primary Node; they are not current topology claims. The later Sharp v22.1 chat-template A/B was also isolated to this direct lane; it did not change the preferred stock-template challenger configuration. The August 21 DFlash2 diagnosis was likewise isolated and direct. The exact float32 selector and the best BF16 single-request arm both missed the retained 128K capacity contract, so neither changed the selected stock recipe. The same-day MTP3/ReplaySSM A/B then measured a large decode gain but only 70,231 KV tokens and a 1.9% 64K end-to-end regression. The external recipe refresh found no full-context profile with a clear combined speed, fidelity, tool-use, and runtime-safety win. The later managed llama.cpp GGUF campaign changed that capacity conclusion: Unsloth Q4_0 plus its Q4_0 MTP head passed exact retrieval at 253,822 actual prompt tokens while preserving an 8,192-token output reserve. It is now the preferred 5090 FAST-TIER challenger. Promotion remains deferred pending the independent SWE worker, a correct container health probe, and resolution of runtime position warnings. The August 22 real-client follow-up passed OpenClaw and Hermes routed identity, fallback detection, and shell-tool/result continuation. It did not establish a 250K routed contract because the bounded route still declared the earlier 128K SGLang/NVFP4 compatibility metadata and video admission. A truthful 262K, image-only router fingerprint remains a promotion gate.

On September 19, an isolated direct 8K/C1 scout tested the pinned Huihui NInfer NVFP4 candidate. It passed protocol checks and a corrected 12/12 image-only subset, but strict tools were 0/3 at both default and low-effort diagnostics; the incumbent control was 3/3. The original image attempt was 0/12 because vision was disabled and is retained as configuration-failure evidence. No matched speed, memory, or context comparison was run, so the candidate is rejected/no-promotion; existing recommendations do not change. Managed restoration is verified; immutable runtime recreation remains open. See the dated finding.

A compatible version-2 runtime follow-up recovered the original argument contract failure: smoke 2/2, core quality 9/9 in each off/low lane, C1 protocol 5/5, and vision 12/12 passed; the 20-request burst remained 17/20 because three requests received HTTP 429. CRLF-to-LF boundary behavior remains. A 12/12 matched no-spec/MTP3 strict 32-word cells measure 71.4/176.2 mean decode tok/s. A separate GGUF profile measured 107.7 tok/s but has a different checkpoint, runtime, and quantization. Huihui MTP3 used 21,852 MiB post-workload versus GGUF 19,022 MiB, so it misses the footprint requirement. MTP3 32K is unpaired. Managed restoration is verified and the candidate remains no-promotion. See the follow-up finding.

The compatible-runtime follow-up also passed managed recreation and bounded router/client compatibility. Six capacity artifacts were eligible for descriptive reporting; failed 128-word warmups remain excluded from performance rows and retained as failures. See the compatibility addendum.

On September 3, the exact NInfer NVFP4 artifact and runtime revision completed a matched no-speculation/MTP3 direct qualification. MTP3 delivered 0.430-second median TTFT, 165.9 tok/s median decode, and 0.720-second median E2E at 4K/C1, versus 0.421 seconds, 75.3 tok/s, and 1.085 seconds for the control. It also returned the exact marker from a 201,746-token API-reported prompt with an 8,192-token completion cap. This made it the preferred measured direct text/tools performance challenger at that point, while GGUF remained the broader-capability incumbent with image/OCR, agentic, endurance, routed-client, and deeper actual- prompt evidence. NInfer remains no-promotion: shared-prefix tools were 17/20, only 2,354 MiB remained free, and routed/client plus broader gates are open.

A later same-day twelve-arm quant/speculation bakeoff measured the idle card before changing the reserve: no GPU processes, 0 MiB reported used, and 32,187 MiB free. Gittensor's target-only RTX5090 NVFP4/SGLang profile then won warm TTFT at 50.9 ms, passed a 244,002-token actual prompt, and completed C2 at 128.7 aggregate output tok/s. Its advertised DSpark arm failed on incompatible target/draft shapes and its FP8 KV path used default 1.0 scales. It is now the preferred direct TTFT challenger, while GGUF remains the broad-capability incumbent. CometKim MTP3 reached 228.0 tok/s decode but failed strict tools 0/3. The final Unsloth Dynamic V3.0 NVFP4 MTP3 arm passed tools 20/20 and measured 137.7 tok/s short decode plus 127.5 tok/s at a 53,706-token prompt, with 3,198 MiB free. It is the strongest clean 64K speculative arm. The zero added reserve was a dedicated-campaign exception only.

The card previously had two operator-selectable shapes:

  1. an exclusive Nemotron Nano Omni 30B stack for auxiliary text, general vision, and OCR; or
  2. Qwen2.5-Omni 3B co-resident with dedicated Parakeet STT and Kokoro TTS.

Embeddings/reranking and ComfyUI remain separately admitted stacks. The PRO 6000 Primary may remain running during these tests, but that makes it protected/co-resident, not measured.

On 2026-08-28, the isolated lane also qualified the first bounded ComfyUI generation candidates. FLUX.2 Klein 4B FP8 produced a decodable 512×512 PNG in 9.859 seconds at 12,919 MiB peak. Wan2.2 TI2V 5B produced a decodable 17-frame, 512×288 H.264 MP4 in 9.092 seconds at 18,263 MiB peak from a clean 943 MiB worker baseline. These are functional/capacity results only: both workflows remain unavailable pending independent perceptual review and no router or client path was exercised. A later exact-build cross-host pass exercised real Hermes MCP image/video jobs, A2A replay, artifact delivery, cold approval/unavailable behavior, and managed teardown. Two FLUX.2 images passed bounded independent review. The new Wan2.2 sample decoded correctly but failed spatial and prompt-adherence review. Both workflows remained unavailable at that point.

On 2026-09-15, a managed-media bringup left the worker healthy and running on the 5090 with its declared 4,096 MiB reserve. One fixed FLUX.2 high-profile image and Wan v1/v2 functional video jobs completed. Wan v1 independently failed all sampled frames; Wan v2 restored recognizable content but remained partial and unavailable. These defect-diagnostic runs are not comparable performance cells. See the media bringup finding.

The earlier August 28 production gate enabled only the exact FLUX.2 image workflow at three fixed c1 profiles: 512×512, 768×768, and 1024×1024, all four steps. Real Hermes produced one warm image at each size and five cold 512×512 regressions. Six of eight passed strict independent review; two technically valid draft samples retained origami/material and exact-count failures. The cold path passed explicit approval, managed build/start, exact same-job resume, artifact return, and teardown. Its 2,045.626-second incident E2E includes multiple defects repaired during the run; the final server-issued exact-resume regression measured 908.936 seconds E2E and 0.087 seconds generation. Its native MCP image matched the authenticated artifact resource byte for byte. Wan2.2 v1 remained quality_failed and unavailable; the September 15 v2 repair remains quality_unverified and unavailable.

Last measured and challenger state

Capability Model Decision Boundary
Qualified text/tools/image reference Huihui Qwen3.8 NInfer MTP3 160K/C1 qualified 160K reference; retained 64K comparison Direct preflight 8/8, vision 18/18, depth retrieval 9/9, unique short capacity 12/12 and long 3/3; 150,144 maximum measured input with 8,192 output allowance; no matched speed or endurance claim
Lowest direct text TTFT and long context Qwen3.8 27B Gittensor RTX5090 target-only preferred measured TTFT challenger, no-promotion Direct 262K/C2; 50.9 ms warm median TTFT, 79.5 tok/s decode, 244,002 actual prompt; DSpark shape failure and FP8 KV scale caveat open
Clean 64K speculative text lane Qwen3.8 27B Unsloth Dynamic V3 NVFP4 MTP3 preferred clean 64K speculative challenger, no-promotion 388.7 ms TTFT; 137.7 tok/s short and 127.5 tok/s at 53,706 prompt tokens; tools 20/20; 3,198 MiB free
Text-to-image generation FLUX.2 Klein 4B historical qualification; simultaneous model/media capacity unmeasured Direct and real-Hermes c1; fixed 512/768/1024-pixel, four-step profiles; 6/8 strict bounded reviews pass with two retained draft failures; arbitrary dimensions rejected
Text-to-video generation Wan2.2 TI2V 5B unavailable candidate, no-promotion v1 prompt/spatial review failed; September 15 v2 small/standard clips restored recognizable content, but temporal and broader quality remain unverified
High direct text decode and long context Qwen3.8 27B NInfer NVFP4 + MTP3 retained measured decode challenger, no-promotion Direct c1/252,928; 0.430-second median TTFT and 165.9 tok/s decode; 201,746 actual prompt + 8,192 output cap; shared-prefix tools 17/20; 2,354 MiB free; routed/client and broader gates open
Broader text/tools, long context, image/OCR Qwen3.8 27B Unsloth Q4_0 + MTP3 retained 5090 GGUF incumbent and broad-capability challenger, no-promotion Direct c1/262K and real-client short/tool route pass; 253,822 actual prompt + 8,192 output reserve; routed 250K blocked by stale 128K metadata; native video unsupported
Computer-use perception, vision, OCR, video Qwen3.8 27B RadixArk NVFP4 preferred 5090 challenger, no-promotion Direct only; c1/128K; eight images / two videos; GUI action loop and routed admission open
Auxiliary text, vision, OCR Nemotron Nano Omni 30B historical current topology Audio input not qualified
Smaller co-resident Omni Qwen2.5-Omni 3B challenger, no-promotion Text/image/OCR passed; audio was noisy
STT Parakeet TDT 0.6B v3 historical current Routed default at capture time
STT Qwen3-ASR 0.6B challenger, no-promotion Non-inferior WER margin; not routed
STT Nemotron 3.5 ASR rejected WER regression exceeded margin
TTS Kokoro historical current Co-resident voice component at capture time

Comparable quality and capacity

Model/configuration Evidence Measured result Decision
Huihui NVFP4/mixed FP8, NInfer 70434721, INT8 KV, MTP3, 160K/C1 functional, bounded quality, unique-canary capacity Direct 8/8, vision 18/18, depth retrieval 9/9; descriptive short 179.7767 / long 158.789 mean decode tokens/s (n=12 / n=3) qualified reference; earlier 64K comparison; no matched speed or endurance claim
Huihui NVFP4/mixed FP8, NInfer 70434721, INT8 KV, MTP3, 64K/C1 functional, bounded quality; non-comparative canary-free C1 capacity Direct7/7, routed6/6, vision12/12, Hermes9/9; descriptive short182.1/long167.1 mean decode tokens/s (n12/n6); 24,224 MiB post-workload historical comparison; no 64K matched no-spec or endurance claim
Qwen3.8 27B Gittensor RTX5090 NVFP4 target-only, SGLang functional, capacity, matched performance, bounded deterministic quality 50.9/277.0 ms warm TTFT p50/p95; 79.5 tok/s decode; 244,002-token actual prompt; C2 128.7 aggregate tok/s; coding/triage/tools 3/3 preferred direct TTFT challenger, no-promotion; advertised DSpark incompatible; FP8 KV scales and broader/routed gates open
Qwen3.8 27B CometKim NVFP4Full MTP3, NInfer functional, capacity, matched performance, negative bounded quality 326.3 ms TTFT, 228.0 tok/s warm decode, 244,002-token actual prompt; coding/triage 3/3; strict tools 0/3 decode research lead only; rejected as general-purpose winner
Qwen3.8 27B Unsloth Dynamic V3 NVFP4 MTP3, vLLM functional, capacity, matched performance 388.7 ms TTFT, 137.7 tok/s warm decode, 127.5 tok/s at 53,706 prompt tokens, C2 70.5 aggregate tok/s, tools 20/20, 83.2% token acceptance strongest clean 64K speculative arm; pinned-runtime and broader-quality/routed gates open
FLUX.2 Klein 4B FP8, ComfyUI v0.33.4 functional, capacity, routed/client acceptance, bounded quality, production cutover Direct decodable PNG at 9.859 s/12,919 MiB peak; real-Hermes fixed 512/768/1024-pixel profiles; 6/8 strict bounded reviews pass and two draft fidelity/count failures are retained; warm gateway E2E 1.242–1.650 s; exact cold resume/teardown pass; c1 historically qualified at three profiles; simultaneous model/media capacity unmeasured
Wan2.2 TI2V 5B FP16, ComfyUI v0.33.4 functional, capacity, routed/client acceptance, negative bounded quality Direct decodable MP4 at 9.092 s/18,263 MiB peak; exact-build real-Hermes 117,738-byte H.264 decode pass; prompt/spatial review fail; c1 unavailable candidate, no-promotion; revise as a new workflow version
Qwen3.8 27B NInfer NVFP4 + MTP3 functional, capacity, matched performance, bounded deterministic quality MTP3 versus no spec: 0.430/0.421-second median TTFT, 165.9/75.3 tok/s decode, and 0.720/1.085-second E2E; 201,746-token prompt plus 8,192 output cap; coding/triage/tools 3/3 each; shared-prefix burst 17/20; 29,834 MiB used / 2,354 free retained direct decode challenger, no-promotion; ordinary reserve, admission, runtime-image, routed/client, multimodal, broad agentic/SWE, and endurance gates open
Qwen3.8 27B Unsloth GGUF Q4_0 + MTP3, llama.cpp functional, capacity, matched performance, bounded deterministic quality 253,822 actual prompt tokens plus 8,192 reserve; tools 20/20 and long-tools at 110,875; agentic 16/18; neutral 101-turn endurance 3/3; images 18/18; 104.1 tok/s short decode; 25,408 MiB startup VRAM retained GGUF incumbent and broad-capability challenger, no-promotion; Q6_K+same-MTP infeasible under reserve policy
Qwen3.8 27B RadixArk NVFP4, SGLang TP=1 functional, bounded deterministic quality 119,675-token retrieval and tools 20/20; direct image/OCR/video; media 30/30; count boundaries 4/4 at eight images / two videos; 20.14 GB weights and 3,928 MiB free after startup challenger, no-promotion
Qwen3.8 27B RadixArk NVFP4 plus MTP3/ReplaySSM functional, capacity, matched performance Decode 138.19 tok/s at 4K and 117.10 at 64K versus 76.54/69.75 baseline; tools 20/20; only 70,231 KV tokens; 64K E2E 1.9% slower rejected as 128K replacement; exact baseline restored, no-promotion
Qwen3.8 27B RadixArk NVFP4 plus DFlash2 functional, capacity Exact float32 selector: 24,347 KV tokens and 24,341-token input ceiling. Best BF16/single-slot/no-radix/no-prefill-graph arm: 70,262 KV tokens, 49,549-token retrieval pass, tools 20/20; fixed DFlash intermediate BF16 SSM state 1.12 GB Both arms rejected as 128K replacements; short-context lead only, no-promotion
Qwen3.8 27B RadixArk NVFP4, Sharp v22.1 template A/B functional, bounded diagnostic quality Sharp preflight passed; thinking-enabled MMLU-Pro stayed 24/30 with +10.8% tokens and +10.7% latency; thinking-disabled behavior used 5.1% fewer tokens but passed 15/18 versus stock 18/18 Sharp rejected; stock retained
Nemotron Nano Omni 30B NVFP4 functional, capacity Text/JSON/4K retrieval/tools plus image/OCR passed; 65,536 context, 2 sequences current topology
Qwen2.5-Omni 3B functional, capacity Text/image/OCR and basic audio input passed; 24,576 MiB reservation no-promotion
Parakeet, sequential quality, capacity 3.343% micro-WER; p95 177.87 ms current
Qwen3-ASR, sequential quality, capacity 3.621% micro-WER; p95 113.58 ms challenger, no-promotion
Nemotron 3.5 ASR, sequential quality, capacity 6.685% micro-WER; p95 225.45 ms rejected

Historical controls

Model / configuration Retained result Limitation / decision
Qwen3.5 35B-A3B Q4_K_M July 11 tools 20/20; 59,053-token actual prompt; 56.279 aggregate tok/s in a ten-request short probe Missing immutable model/runtime pins; historical no-promotion
Gemma 4 E2B W4A16 July 16 preflights through 120K passed; strict timeout triage 0/3 no-promotion; short-generation capacity is not controlled decode
Gemma 4 E4B Fast July 13 low-latency/router control Historical selection, not current deployment
Gemma 4 31B July Fast-lane load failed Separate from the larger-card measurements in its dossier

Later 5090 work established the Omni shapes above. Older voice research remains in the chronological findings index.

Run history

See RTX 5090 runs. The August 28 ComfyUI media qualification, exact-build live validation, and Hermes image production enablement; the September 3 NInfer qualification; August 21 GGUF qualification, recipe research, and DFlash2 diagnosis; August 20 template A/B; and August 17 Qwen rows classify the PRO relationship as unrelated. The July 28 ASR rows explicitly classify the PRO 6000 as protected.

2026-09-19: Huihui 64K C1 promotion

The earlier Huihui NInfer MTP3 qualification used 65,536 context tokens at C1. Direct preflight passed 7/7, routed preflight 6/6, vision 12/12, and native Hermes 9/9 with the exact local alias. Descriptive, canary-free C1 short-input capacity passed 12/12 at 182.1 mean decode tokens/s; the long-input cell passed 6/6 at 167.1, with 60,769-60,776 actual prompt tokens. Post-workload GPU use was 24,224 MiB. The 32K results remain historical evidence. Endurance, interactive browser acceptance, and a matched 64K no-spec control remain unmeasured. See the 64K finding.

2026-09-19: Huihui direct-I/O 160K context envelope (accepted)

The earlier 64K/C1 results remain a separate comparison. A separate final baked 163,840-context/C1 direct-I/O profile passed startup, direct preflight 8/8, and repeated exact retrieval 9/9 at 150,058, 150,144, and 150,124 prompt tokens; the maximum retained input is 150,144 with an 8,192-token output allowance and approximately 4.5 GiB GPU free at startup. It used MTP3 with INT8 KV and an 8,192-token visible-output cap. Startup minima were 4,606 MiB GPU free and 14,165.406 MiB Windows available; long-capacity minima were 4,598 MiB and 13,669.93 MiB respectively. The first retrieval of each case was cold or partially cached and later attempts warm; this does not support a throughput or cold-latency comparison. The subsequent direct-I/O verifier avoided the observed host-cache pressure; these sequential runs do not establish exclusive causality. The tested reference uses 163,840 context tokens and an 8,192-token output allowance; bounded native client compatibility is recorded separately. See the context-envelope finding.