Voice on a 16GB Mini: local STT+TTS, LLM routed to primary-node¶
Publication redaction: The operator-specific network hostname was replaced with a public-safe placeholder. The observed workflow and result are unchanged.
STATUS: T016 LIVE PROOF SUPPORTED. The harness is
scripts/voice/mini_validation.py. The acceptance command writes a JSON report, appends one session row below, and returns nonzero unless the verdict issupportedso a negative-control or partial proof cannot satisfy the task. Missing target hardware, primary-node route/auth, endpoint model identity, post-benchmark per-serve memory, nonblank STT/LLM text, or first-audio proof isunsupported, not acceptance evidence:SUPERSEDED FOR NORMAL OPERATIONS (2026-07-08). This remains a valid optional Mini-local audio proof, but it is no longer the reference OpenClaw Talk or candidate benchmark topology. Companion Node's 16 GB RAM is reserved for OpenClaw Gateway, Anvil Voice Realtime/proxy, Claude Code, and Codex during normal validation. Do not use this note to justify running STT/TTS/LLM model serves on Mini outside an explicit same-host/local-audio test.
Related: docs/findings/2026-07-04-hf-speech-to-speech-review.md s8 (VRAM/RAM
math: STT ~1-4GB, TTS ~0.5-7GB — comfortably small even on a 16GB box) · the
saved Mini<->router tailnet-binding note (router publishes its tailnet IP,
not loopback) · scripts/voice/mini_validation.py
Known gaps (flagged, not hidden)¶
- Driver-process RSS is not the number that matters. The script's own
resource.getrusage-based memory reading reflects its driver process, not the STT/TTS serves. The report also records host memory before/after load and requires post-benchmark per-serve memory. Managed container serves usedocker stats; native Mini serves use macOSlsofpluspsto attribute RSS to the process listening on the configured loopback port, so lazy model load is included. - A non-Mini run is a negative control, not a pass. The harness records
host_is_16gb_class,host_matches_expected_mini, andhost_hw_model_matches_expected; runs on a workstation, generic 16GB VM, or GPU host must be read asunsupportedunless the report proves a 16GB-class macOS host, a Companion Node host identity, and the expected Mini hardware model (Mac16,10by default). - The Mini manifest uses native STT/TTS lifecycle for optional local-audio
validation.
examples/voice/companion-node.tomldeclareslifecycle = "native"for both audio endpoints.anvil-serving voice up/downnow starts and stops the MLX Audio processes with manifest-declared commands, PID files, and logs under/tmp/anvil-voice-mini; the harness still requires them to be ready on127.0.0.1:30010/30011, complete the live benchmark, and produce endpoint-attributed process RSS plus a matching/v1/modelsmodel id after the benchmark. Managed container serves may still be validated with a customserves.toml. - Router auth is checked both ways. Manifests that name
ANVIL_ROUTER_TOKENmust prove the token is present for the positive route probe and that a no-Authorization/v1/routeprobe is rejected with 401/403.
How to run¶
The default manifest is examples/voice/companion-node.toml when present: STT and
TTS are native-managed loopback endpoints on the Mini, while the LLM base URL
points at primary-node over the tailnet and declares the expected
route/provider/model and expected endpoint host. The shell running the command
must have ANVIL_ROUTER_TOKEN set when the manifest names that auth env var.
For a custom Mini serve manifest:
python scripts/voice/mini_validation.py \
--config examples/voice/companion-node.toml \
--serves-manifest ./serves.toml \
--report /tmp/mini-run1.json
Exploratory negative-control runs may opt into a zero exit for diagnostics, but that mode is not acceptable as T016 evidence:
Measurement template¶
| metric | value | notes |
|---|---|---|
| host total memory | TBD | must be 16GB-class for a target-hardware pass |
| host memory before load | TBD | available/used GB |
| host memory after serves ready | TBD | available/used GB |
| host memory after benchmark | TBD | available/used GB; verdict uses this value |
| expected Mini host match | TBD | default pattern matches Companion Node/mini-host.example; override with --target-host-pattern only for renamed target hardware |
| expected Mini hardware model | TBD | default Mac16,10; override with --target-hw-model-pattern only for approved target hardware changes |
| STT startup (s) | TBD | |
| STT memory proof after benchmark | TBD | docker_stats for managed containers, or macos_process_rss attributed to the 127.0.0.1:30010 listener |
| TTS startup (s) | TBD | |
| TTS memory proof after benchmark | TBD | docker_stats for managed containers, or macos_process_rss attributed to the 127.0.0.1:30011 listener |
| TTFA (ms), LLM on primary-node | TBD | via anvil_serving.voice.benchmark |
| turn latency (ms) | TBD | |
| STT/LLM text and TTS audio | TBD | STT hypothesis and LLM reply must be nonblank; TTS output must include >=0.25s of audio |
| LLM endpoint host / route | TBD | must be primary-node tailnet host and expected route provider/model/tier |
| LLM auth env present | TBD | ANVIL_ROUTER_TOKEN expected for primary-node |
| driver process peak RSS (MB) | TBD | informational only — see gap #1 |
| failure mode(s) observed | TBD | e.g. OOM-kill, tailnet timeout, cold-start stall |
| verdict | TBD | supported is the only accepting verdict |
Session log¶
| timestamp (UTC) | host | host memory | verdict | STT | TTS | TTFA / latency ms | host used / available GB | failure modes | report path |
|---|---|---|---|---|---|---|---|---|---|
| 2026-07-06T08:08:08Z | Mac | 16.0 GB; 16gb_class=True | supported | ready; rss=64.5MB pid=10877 | ready; rss=92.11MB pid=11459 | 1083.18 / 1083.75 | 12.42 / 3.58 | all required Mini validation checks passed | docs/findings/2026-07-voice-16gb-mini.json |
| 2026-07-06T08:17:40Z | Mac | 16.0 GB; 16gb_class=True | supported | ready; rss=64.66MB pid=10877 | ready; rss=82.22MB pid=11459 | 1015.0 / 1015.44 | 12.83 / 3.17 | all required Mini validation checks passed | docs/findings/2026-07-voice-16gb-mini.json |
(mini_validation.py appends a row here automatically — see
append_finding_row in that script.)
Findings¶
The 2026-07-06 target-hardware run on Companion Node (Mac16,10, 16.0 GB RAM)
passed the required split topology: STT and TTS were local native-managed
loopback endpoints, while the LLM call routed over the tailnet to primary-node.
The report recorded 3.17 GB available after load, post-benchmark listener RSS
for both local audio serves, a nonblank STT hypothesis, a nonblank LLM reply,
first synthesized audio, TTFA 1015.0 ms, turn latency 1015.44 ms, and TTS RTF
0.1127. The positive route proof returned fast-local / qwen36-27b /
local, and the no-Authorization route probe was rejected with HTTP 401.
Decision¶
supported for T016 on the measured 16GB Companion Node. The accepting evidence
is docs/findings/2026-07-voice-16gb-mini.json; any non-target run,
all-local Mini LLM run, missing primary-node route/auth proof, missing
post-benchmark audio-serve memory, or missing first-audio benchmark remains
unsupported.