Skip to content

STT Model Benchmark And Candidate Recommendation

Publication redaction: Operator-specific GPU UUIDs and network hostnames were replaced with stable labels or public-safe placeholders. Hardware class, measurements, and event ordering are unchanged.

Status: executed on Primary Node, 2026-07-08. This report covers speech-to-text candidates for the OpenClaw/Anvil Voice path. Companion Node was kept model-free. Heavy was not disrupted. Fast was stopped only while the RTX 5090 hosted STT candidates, then restored to healthy.

Scope

Goal: find the best local STT model for the speech-to-speech path, with latency weighted above accuracy once a candidate clears a basic transcript correctness gate.

Command/resource topology:

  • Command host and model host: Primary Node.
  • Candidate GPU: RTX 5090, UUID REDACTED_GPU_UUID_REMOVED_GPU.
  • Heavy GPU: RTX PRO 6000, left running and untouched.
  • Mini: OpenClaw Gateway / Anvil Voice Realtime role only; no STT/TTS/LLM model serve was started on Mini.
  • Test audio: .anvil/evidence/voice-samples/quick-brown-fox-kokoro-16k.wav, 16 kHz mono PCM, reference the quick brown fox jumps over the lazy dog.

Research Inputs

Primary references used before candidate selection:

No credible Gemma-branded STT candidate surfaced in this research pass. Gemma belongs in the LLM leg of the voice cascade, not the STT model slot.

Rubric

Hard gates:

  • Must run from the anvil-serving operational surface: models pull, serves, and voice stt-benchmark.
  • Must keep Mini model-free and must not use the Heavy GPU.
  • Must expose an OpenAI-compatible /v1/audio/transcriptions endpoint or have a clear adapter path to one.
  • Must avoid transcript hallucination/repetition on the smoke sample.
  • Must reach wer_normalized <= 0.05 on the smoke sample before being considered a viable Talk candidate.

Scoring priorities after the hard gates:

Dimension Weight Notes
Warm STT latency 40% Target < 200 ms; acceptable experimental band < 300 ms.
Accuracy 30% Use normalized WER for ASR content; track exact WER only for punctuation/case drift.
Operational fit 20% Managed compose entry, predictable health, no custom scripts, low startup surprises.
Resource/cost 10% 5090 VRAM/opportunity cost, cold-start tax, and whether Fast must be displaced.

Historical Benchmark Harness

The experiment branch used the following STT-only benchmark command. This is a historical reproduction record, not current CLI guidance; the branch's command and candidate serve changes were not carried forward during publication.

anvil-serving voice stt-benchmark \
  --config examples/voice/openclaw-anvil-voice.toml \
  --profile dark-audio \
  --sample .anvil/evidence/voice-samples/quick-brown-fox-kokoro-16k.wav \
  --reference-text "the quick brown fox jumps over the lazy dog"

The experiment branch also used candidate overlays and temporary managed serve entries for:

  • voice-stt-qwen3-asr-0-6b
  • voice-stt-qwen3-asr-1-7b
  • voice-stt-whisper-large-v3-turbo-fp8
  • voice-stt-whisper-large-v3-turbo

Qwen3-ASR required post-processing because the served output can include language English<asr_text>.... The experiment branch's STT client supported provider-neutral request fields for OpenAI/vLLM transcription forms, including language and max_completion_tokens; that implementation was not carried forward with this evidence-only publication.

Live Results

Warm latency excludes the first request after a cold model start when a clear warm repeat exists.

Candidate Endpoint model Runtime First request ms Warm ms Normalized WER Result
Baseline Parakeet tdt_ctc-110m existing Dark audio endpoint 82.25 82.62, 89.30 0.0 Pass; fastest and simplest
Qwen3-ASR 0.6B qwen3-asr-0.6b qwenllm/qwen3-asr:latest / vLLM 33210.86 196.36, 278.98 0.0 on warm runs after postprocess Pass after provider-prefix postprocess; viable fallback
Qwen3-ASR 1.7B qwen3-asr-1.7b qwenllm/qwen3-asr:latest / vLLM 34538.35 290.59, 223.96 0.0 Pass; not justified over 0.6B on this smoke
Whisper Turbo FP8 whisper-large-v3-turbo-fp8 vLLM in Qwen ASR image 44007.92 730.69, 669.57 1.0 Fail; repeated hallucinated phrase
Whisper Turbo BF16 whisper-large-v3-turbo vLLM in Qwen ASR image 45784.58 642.42, 643.63 1.0 Fail; repeated hallucinated phrase
Whisper Turbo BF16 bounded same same, with max_completion_tokens=64 n/a 235.39, 141.52 1.0 Fail; faster but still wrong

The 17 raw artifacts are published in the 2026-07-08-stt-model-benchmark-evidence/ directory. They contain three runs each for Parakeet, Qwen3-ASR 0.6B, Qwen3-ASR 1.7B, Whisper Turbo FP8, and Whisper Turbo BF16, plus two bounded Whisper Turbo BF16 runs. The first Qwen3-ASR 0.6B artifact preserves the raw provider-prefixed transcript and its resulting normalized WER of 0.4444; the two warm artifacts record the corrected post-processing path and normalized WER of 0.0.

Candidate Notes

Parakeet remains the production default. It won the local latency test by a wide margin and had zero normalized WER on the smoke sample. It also keeps the operational surface simple because it is already the Dark-host audio endpoint used by OpenClaw Talk.

Qwen3-ASR 0.6B is the best second choice. It cleared the smoke accuracy gate and stays within the experimental latency band when warm. Its drawbacks are the cold-start tax, the need to displace Fast on the 5090 during isolated tests, and the Qwen image quirk where CUDA_VISIBLE_DEVICES had to use ordinal 0 even while the Docker device reservation remains pinned to the 5090 UUID.

Qwen3-ASR 1.7B is not rejected, but it did not show a reason to prefer it over 0.6B on this sample. Keep it for a larger accuracy corpus with noisy audio, numbers, accents, and tool-call phrases.

Whisper Turbo is rejected for the current vLLM/OpenAI transcription path. Both the FP8 and base variants reached health and exposed /v1/audio/transcriptions, but they produced repeated hallucinated text. Adding bounded decode fields reduced latency but did not fix correctness. This does not prove Whisper itself is bad; it means this specific vLLM serve path is not acceptable for OpenClaw Talk without further debugging or a different adapter such as faster-whisper.

SenseVoice, Moonshine, Distil-Whisper, and Canary-Qwen remain research candidates. They need an adapter or runtime decision before they can be compared fairly through the same OpenAI-compatible benchmark command.

Recommendation

Keep the current Parakeet/tdt_ctc-110m STT endpoint as the OpenClaw Talk default.

Keep Qwen3-ASR 0.6B as the next candidate for a future managed evaluation against real voice samples. It is the only new candidate that cleared the correctness gate while staying close enough to the latency target to justify more work.

Do not promote Qwen3-ASR 1.7B or Whisper Turbo from this smoke. Qwen3-ASR 1.7B needs a harder accuracy corpus before its larger runtime footprint is justified. Whisper Turbo needs a separate compatibility/debug task before it can be judged as an ASR model rather than as a failed serving recipe.

Next benchmark step: build a small local corpus with at least 20 short samples: clean speech, conversational filler, zip codes/numbers, weather/tool requests, fast speech, mild background noise, and one long utterance. Run Parakeet, Qwen3-ASR 0.6B, and Qwen3-ASR 1.7B through the current canonical benchmark surface, or first restore an equivalent reviewed STT-only operator command, and report p50/p95 latency plus normalized WER by category.