Skip to content

Fast-Tier LLM Bakeoff Plan and Candidate Registry

Date: 2026-07-08

This finding anchors the Anvil PRD fast-tier-llm-bakeoff-2026-07. It is the source-backed candidate registry and scoring rubric for the long-running RTX 5090 Fast-tier bakeoff. The entries below are priors, not promotion evidence, until a local Primary Node run records live health, preflight, voice, context, tool, and benchmark artifacts.

Operating Constraints

  • Companion Node stays model-free for reference OpenClaw Talk validation. It may run OpenClaw Gateway, Anvil Voice Realtime/proxy, Claude Code, and Codex, but not STT, TTS, or LLM model serves.
  • Primary Node owns the RTX 5090 Fast candidate serves, router, STT/TTS endpoints, and benchmark execution.
  • Heavy is not part of this bakeoff and must remain available for larger, more capable models.
  • Reddit and community posts are recipe priors only. Official model cards and engine docs are the source of truth for model facts; local benchmark artifacts are the source of truth for promotion.
  • A candidate may test at most three engine/quantization variants. Prefer a small number of well-sourced runs over broad, low-signal churn.

Candidate Matrix

Candidate Required Primary role Official source Community or recipe prior Initial variants
nvidia/Qwen3.6-27B-NVFP4 yes, control Current Fast baseline HF + vLLM recipe RTX 5090 vLLM/MTP Reddit runs current vLLM 32K, vLLM long-context, llama.cpp/GGUF only if a matching quant is sourced
nvidia/Qwen3.6-35B-A3B-NVFP4 yes Qwen MoE quality/context candidate HF + vLLM recipe HF/Reddit release thread vLLM NVFP4 32K, vLLM 64K/128K if memory allows, SGLang only if recipe is verified
nvidia/Gemma-4-31B-IT-NVFP4 yes dense/near-dense Gemma candidate HF + NVIDIA NIM card RTX PRO/Gemma NVFP4 Reddit reports vLLM NVFP4 32K with FP8 KV, vLLM longer context, llama.cpp GGUF Q5/Q6 if sourced
zai-org/GLM-4.7-Flash yes tool/agent MoE candidate HF model card Reddit GLM support/caveat threads plus vLLM/SGLang/llama.cpp article SGLang/vLLM 32K, llama.cpp GGUF only with caveat tracking, one longer-context attempt if stable
mistralai/Devstral-Small-2-24B-Instruct-2512 yes coding-agent/tool-use candidate HF + Mistral Devstral page Reddit release thread, GGUF card, Unsloth llama-server docs vLLM FP8 32K, llama.cpp GGUF Q5/Q6, longer-context run if stable
Qwen/Qwen3-30B-A3B-Instruct-2507 optional fallback speed/quality reference HF model card/discussion Qwen/GLM comparison posts Use only if a required model cannot run or leaves a clear coverage gap

Scoring Rubric

Score each locally tested candidate/config on 100 points after hard gates pass:

Category Points What to measure
Voice latency 30 LLM-stage latency, TTFT/TTFA, total STT-to-LLM-to-TTS turn latency, repeated-turn stability
Intelligence and tool quality 30 coding/editing prompts, tool-call schema adherence, practical instruction following, Haiku-like usefulness
Usable context 15 accepted context target, 32K/64K/128K sweep result, long-context retrieval/needle behavior
Agent and multi-turn reliability 15 session memory, transcript behavior, no empty/repeated/thinking-only responses, stable tool history
Operational fit 10 load time, VRAM headroom, health behavior, recipe reproducibility, rollback safety

Hard gates before scoring:

  • Loads on Primary Node's RTX 5090.
  • Does not disrupt Heavy.
  • Completes at least one OpenClaw/Anvil Voice cycle.
  • Supports usable tool calls.
  • Preserves session/chat behavior.
  • Restores the production Fast baseline after disruptive experiments.

Promotion rule: promote only if the candidate beats the current nvidia/Qwen3.6-27B-NVFP4 baseline on total score, passes every hard gate, and does not introduce unacceptable voice, tool, context, or operational regressions. Otherwise keep the current baseline and mark the best alternate as verified/non-promoted or needs-more-data.

Live Results

Evidence directory: docs/findings/fast-tier-bakeoff-evidence/ Matrix typed summary: t006-evidence-summary.json Final T007 gate summary: t007-evidence-summary.json

All voice runs used Primary Node audio endpoints over the private address and recorded mini_model_free_assertion.passed=true; Mini was not used to host STT, TTS, or LLM models. Final runtime restoration evidence is recorded in runtime-restoration.md: Heavy used the same running container with restart count 0 and health 200 observations during the matrix window and final handoff; production Fast was restored to vllm-qwen36 with health 200; and all experimental candidate serves were stopped or absent.

Voice artifacts in this report are stage-latency evidence. They measure STT, LLM, and TTS timing with TTS first-audio observed, but the STT hypothesis field is empty with WER 1.0, so these artifacts must not be read as semantic STT accuracy evidence.

Candidate/config Engine / quant / context Voice total / LLM stage TTFT / E2E Approx decode tok/s Tool / session / intelligence Evidence Result
nvidia/Qwen3.6-27B-NVFP4 control vLLM / NVFP4 / 32K 1130.21 ms / 814.83 ms 6203.94 ms / 9041.91 ms 67.65 pass / pass / 0.50 qwen36-27b-baseline-vllm-32k.voice.json, qwen36-27b-baseline-vllm-32k.bakeoff.json Control rerun succeeded
nvidia/Qwen3.6-35B-A3B-NVFP4 vLLM / NVFP4 / 32K 377.52 ms / 165.40 ms 1489.36 ms / 2302.37 ms 236.16 pass / pass / 0.50 qwen36-35b-a3b-vllm-nvfp4-32k.voice.json, qwen36-35b-a3b-vllm-nvfp4-32k.bakeoff.json Best promotion candidate
nvidia/Gemma-4-31B-IT-NVFP4 vLLM / NVFP4 / 32K then 8K retry no successful voice run no benchmark artifact n/a no score load-failures.md Rejected: does not fit 5090 recipe
zai-org/GLM-4.7-Flash SGLang / BF16 / 32K no successful SGLang voice run no SGLang benchmark artifact n/a no score load-failures.md Rejected for SGLang: no KV headroom
zai-org/GLM-4.7-Flash llama.cpp / UD-Q4_K_XL / 32K 2376.21 ms / 961.49 ms 6196.05 ms / 7417.46 ms 157.20 pass / pass / 0.00 glm47-flash-llamacpp-q4-32k.voice.json, glm47-flash-llamacpp-q4-32k.bakeoff.json Verified but not competitive
mistralai/Devstral-Small-2-24B-Instruct-2512 vLLM / FP8 / reduced 8K 923.98 ms / 433.12 ms 742.46 ms / 3755.56 ms 57.75 pass / pass / 1.00 devstral-small2-vllm-fp8-8k.voice.json, devstral-small2-vllm-fp8-8k.bakeoff.json, load-failures.md Promising fallback, context-limited

Approx decode tok/s is derived from the evidence JSON as output_tokens * 1000 / (e2e_ms - ttft_ms) because the timing fields are recorded in milliseconds.

Rubric Score

Scores below apply the 100-point rubric to the measured configuration, not to the model family in general. They are promotion guidance, not an automatic router policy change.

Candidate/config Voice Intelligence/tool Context Agent reliability Ops fit Total Notes
nvidia/Qwen3.6-35B-A3B-NVFP4, vLLM NVFP4 32K 30 22 15 14 8 89 Fastest measured stage-latency path, 32K context, same deterministic intelligence miss as baseline
nvidia/Qwen3.6-27B-NVFP4, production control 20 22 15 13 10 80 Stable baseline, but much slower than 35B-A3B
mistralai/Devstral-Small-2-24B-Instruct-2512, vLLM FP8 8K 24 29 5 14 6 78 Passed the small intelligence suite, but only after reducing context to 8K and adding Mistral request fallback
zai-org/GLM-4.7-Flash, llama.cpp Q4 32K 8 10 15 10 5 48 Tool/session pass, but voice latency and deterministic intelligence failures make it unsuitable for Fast voice
nvidia/Gemma-4-31B-IT-NVFP4 0 0 0 0 0 0 No viable loaded endpoint on the 32 GB 5090 recipe

Determination

Recommendation: promote nvidia/Qwen3.6-35B-A3B-NVFP4 as the next human-gated Fast-tier promotion candidate, but do not auto-promote it in this task. It beat the production nvidia/Qwen3.6-27B-NVFP4 control on the measured voice path and loaded-endpoint bakeoff while preserving the same 32K context target, tool-call pass, and session-recall pass. It should not be auto-promoted by this task because router policy promotion remains a human gate and the deterministic intelligence suite still has one shared failure with the current baseline.

Promotion addendum: the human gate was later approved, and the repo promotion changes are tracked in 2026-07-08-fast-tier-promotion.md. The original task still did not auto-promote; the later promotion step updates the production Fast serve and router config explicitly.

Historical close-out note: keep the qwen36-27b production Fast baseline until the promotion task explicitly updates the deployed route. That condition is satisfied by the follow-up promotion addendum above. Devstral-Small-2 is worth keeping as a reduced-context fallback candidate for agent/code behavior, but it is not a default Fast voice replacement because the successful evidence is 8K, not 32K. GLM-4.7-Flash via llama.cpp is verified but too slow for this use case. Gemma-4-31B-IT-NVFP4 is rejected for the RTX 5090 Fast role under the tested vLLM recipe.

Recipe Table

Config Serve recipe Reproduction command Notes
Qwen3.6-27B baseline configs/serve-recipes.toml#nvidia-qwen36-27b-nvfp4 anvil-serving serves --manifest examples/primary-node/serves.toml up fast Production Fast control on port 30003
Qwen3.6-35B-A3B configs/serve-recipes.toml#nvidia-qwen36-35b-a3b-nvfp4 anvil-serving serves --manifest examples/primary-node/serves.toml up fast-qwen36-35b-a3b Best measured candidate, vLLM NVFP4 32K
Gemma-4-31B configs/serve-recipes.toml#nvidia-gemma4-31b-it-nvfp4 anvil-serving serves --manifest examples/primary-node/serves.toml up fast-gemma4-31b Requires Gemma parser/template flags; failed memory gate
GLM-4.7 SGLang configs/serve-recipes.toml#zai-glm47-flash anvil-serving serves --manifest examples/primary-node/serves.toml up fast-glm47-flash-sglang BF16 path failed KV allocation
GLM-4.7 llama.cpp configs/serve-recipes.toml#zai-glm47-flash anvil-serving serves --manifest examples/primary-node/serves.toml up fast-glm47-flash-llamacpp Uses cached Unsloth UD-Q4_K_XL GGUF by explicit -m path
Devstral Small 2 configs/serve-recipes.toml#mistralai-devstral-small2-24b-instruct-2512 FAST_DEVSTRAL_SMALL2_MAX_MODEL_LEN=8192 FAST_DEVSTRAL_SMALL2_MAX_NUM_SEQS=2 anvil-serving serves --manifest examples/primary-node/serves.toml up fast-devstral-small2 32K recipe exceeded practical 5090 KV headroom; reduced 8K run succeeded

Adversarial Reviews

Two T007 adversarial review packets were run and resolved before final submit:

Packet Reviewer Objections Resolution
t007-review-docs-fixtures.json gpt-5.5 high reasoning SERVES rerun command would skip the bakeoff voice subsection without latency flags; decode-rate formula omitted the ms-to-sec conversion Added the voice latency flags to the SERVES rerun command and corrected the formula to output_tokens * 1000 / (e2e_ms - ttft_ms)
t007-review-evidence-gate.json gpt-5.5 high reasoning Review acceptance was prose-backed only; report linked only the T006 typed evidence summary Added durable T007 review packets and the T007 final gate summary, then linked them from this report

Three earlier review passes were run before T006 was applied:

Reviewer Focus Objections Resolution
Evidence audit, small model Candidate and artifact coverage Heavy health and Fast restoration were initially thin Added runtime-restoration.md and t006-evidence-summary.json; final status showed Fast and Heavy health 200
Code/docs adversarial review, gpt-5.5 high reasoning Diff, policy, and Mistral fallback Low concern that Compose-style env placeholders in serve-recipes.toml weakened copyability Restored literal default values in configs/serve-recipes.toml; kept env overrides only in Compose
Evidence/gates adversarial review, gpt-5.5 high reasoning Acceptance gates and overclaims Heavy before/after proof was incomplete; voice artifacts were being described too broadly Added Heavy container continuity, health timestamps, typed summary, and the stage-latency caveat; re-review reported no remaining blocker

Rerun Workflow

Use Anvil Serving surfaces only; do not run one-off lifecycle scripts as the operational path.

  1. Confirm the topology and keep Mini model-free:
anvil-serving serves --manifest examples/primary-node/serves.toml status
  1. Start exactly one candidate on Primary Node:
anvil-serving serves --manifest examples/primary-node/serves.toml up fast-qwen36-35b-a3b
  1. Gate the loaded endpoint before benchmarking:
anvil-serving preflight \
  --base-url http://127.0.0.1:39010/v1 \
  --model qwen36-35b-a3b-nvfp4 \
  --needle-ctx 32768 \
  --tool-batch 5 \
  --no-thinking
  1. Capture staged voice latency against the same loaded endpoint:
anvil-serving voice benchmark \
  --config examples/voice/openclaw-anvil-voice.toml \
  --profile dark-audio \
  --candidate-base-url http://127.0.0.1:39010/v1 \
  --candidate-model qwen36-35b-a3b-nvfp4 \
  --candidate qwen36-35b-a3b-vllm-nvfp4-32k \
  --evidence-out docs/findings/fast-tier-bakeoff-evidence/qwen36-35b-a3b-vllm-nvfp4-32k.voice.json
  1. Capture the loaded-endpoint bakeoff artifact:
anvil-serving benchmark \
  --bakeoff \
  --base-url http://127.0.0.1:39010/v1 \
  --model qwen36-35b-a3b-nvfp4 \
  --candidate-id qwen36-35b-a3b \
  --config-id vllm-nvfp4-32k \
  --context-targets 32768 \
  --suite chat,context,tool,session,intelligence,voice \
  --thinking-mode disabled \
  --voice-latency-ms 377.52 \
  --stt-latency-ms 68.65 \
  --tts-latency-ms 143.46 \
  --source-recipe configs/serve-recipes.toml#nvidia-qwen36-35b-a3b-nvfp4 \
  --serve-command "anvil-serving serves --manifest examples/primary-node/serves.toml up fast-qwen36-35b-a3b" \
  --evidence-out docs/findings/fast-tier-bakeoff-evidence/qwen36-35b-a3b-vllm-nvfp4-32k.bakeoff.json
  1. Restore production Fast and verify Heavy:
anvil-serving serves --manifest examples/primary-node/serves.toml down fast-qwen36-35b-a3b
anvil-serving serves --manifest examples/primary-node/serves.toml up fast
anvil-serving serves --manifest examples/primary-node/serves.toml status

Source Classes

Use these source labels in configs/serve-recipes.toml and final evidence:

  • official: Hugging Face model card, vendor page, or engine recipe docs.
  • community-prior: Reddit, blog, forum, model discussion, or third-party recipe that still requires local verification.
  • locally-verified: Primary Node evidence captured by anvil-serving benchmark, voice, preflight, serve status, and log artifacts.

Sources