Skip to content

Findings index

Eight-image follow-up, 2026-09-18: GLM EXL3 r9 passed four eight-image comparison requests, one eight-image high-resolution request. Limit now eight/request; no concurrent eight-image soak.

2026-09-18: GLM EXL3 vision enablement — 10/10 image attempts, routed acceptance and a matching Pi output transcript; 260K retrieval passed on retry; initial refusal retained.

2026-09-19: GLM-5.3-Flash r10 APC promotion follow-up — explicit human approval and fresh acceptance followed the retained campaign. Shared-prefix visible TTFT improved 67.84%; fresh-prefix request throughput fell 2.39%, within the frozen 5% gate. Exact r9 remains the documented rollback.

Dated evidence snapshots — benchmarks, live validations, and lab notebooks — that ground the decisions recorded in docs/adr/ and the PRD task history. Each file is a point-in-time record, accurate as of its date and not maintained afterwards; treat the ADRs and the main docs as the current source of truth. .json files are the machine-readable raw evidence backing a companion .md narrative. For current conclusions, enter through the benchmark portal, choose RTX PRO 6000 or RTX 5090, then use the model dossiers or run catalog. This page remains the complete chronological evidence index. Newest first.

Policy

Latest runtime decision: GLM-5.3-Flash r10 APC promotion follow-up — human-approved bounded APC configuration with fresh acceptance. The historical r9 campaign close and its exact rollback remain linked evidence.

Latest bounded run: GLM output-budget diagnosis, 2026-09-11 — retain 4K; 16 valid tasks, eight invalid coding runs.

Latest media bringup: managed Anvil Media bringup, 2026-09-15 — fixed image smoke and Wan v1/v2 functional diagnostics; Wan remains unavailable pending quality.

These findings remain public as durable evidence under ADR-0027. A later benchmark or ADR can supersede a recommendation, but it does not erase the historical observation or move its load-bearing evidence into a private repository.

Address redaction (2026-07-29): private tailnet addresses in published findings and raw evidence files were replaced with the generic placeholder 100.64.0.10. This is a publication-safety edit only — one host, one placeholder, applied uniformly — and changes no measurement, configuration, or conclusion.

  • Add and index a sanitized dated narrative with the exact revision/configuration, topology, method, evidence type, result, failures, and caveats.
  • For local functional, capacity, or quality results, add the result card and companion publication summary defined by the finding format. Platform copy remains derivative; it never replaces the finding, raw evidence, or promotion gate.
  • Keep only the bounded raw JSON/CSV/text needed to audit the claim. Each new raw file must be at most 1 MiB, and all checked-in raw evidence from one experiment, qualification, promotion, or evidence packet must total at most 5 MiB. Splitting paths or narratives does not reset the limit. Exceptions must identify the files/total bytes, justify why bounded or external evidence is insufficient, and name the approving reviewer.
  • Absent an approved exception, put larger, binary, or high-volume evidence at an anonymously downloadable, immutable, non-expiring versioned/content-addressed HTTPS URL retained at least as long as the citation. Record its retention owner/term, byte size, SHA-256 digest, and provenance. Expiring CI artifacts, private buckets, and mutable latest URLs do not qualify.
  • Sanitize both checked-in and externally stored evidence before publication. Never include secrets, credentials, private prompts, personal data, machine-local tokens, or unrelated logs in either location.
  • Retain evidence while any public doc, ADR, benchmark table, or release note depends on it; there is no age-based pruning. Publish a linked erratum or superseding finding instead of silently overwriting a merged measurement. Corrections use a new artifact path; sensitive/legal removals leave a public tombstone with nonsensitive provenance when safe and lawful.
  • Private notes may preserve planning history, but private-only citations cannot ground a public claim. Restate the claim and its auditable support publicly.

Only existing artifacts' size and format are grandfathered; sanitization, correction, and public citation requirements still apply. Any future size cleanup is a separate reviewed migration that must preserve public access, provenance, and content hashes. The legacy corpus is not retroactively certified: issue #175 published or gap-recorded the private-only grounding it identified, while the broader machine-local path and public-artifact audit remains tracked by issue #290.

Date File Subject
2026-09-19 Qwen3.8 context-envelope investigation 128K direct-I/O gates passed; final 160K baked profile has startup, direct preflight 8/8, and repeated 150K retrieval 9/9. The 64K results remain a separate comparison; bounded native client compatibility passed.
2026-09-19 Huihui Qwen3.8 promotion addendum Baked compatible-runtime qualification and bounded router preflight passed. Earlier rejected findings remain historical evidence.
2026-09-19 Huihui Qwen3.8 NInfer compatible-runtime follow-up Compatible v2 runtime recovers smoke, core, protocol, and vision; matched NInfer MTP3/no-spec strict 32-word 8K/C1 cells pass 12/12 at 176.2/71.4 mean decode; workload-matched cross-profile GGUF is 107.7 tok/s and uses less GPU memory; 32K MTP is unpaired, managed restoration verified, no promotion
2026-09-19 Huihui Qwen3.8 27B NInfer NVFP4 RTX 5090 scout Pinned no-spec 8K/C1 candidate passed direct protocol gates and corrected image-only 12/12, but strict tools failed 0/3 at default and low-effort diagnostics; incumbent control passed 3/3, no speed or footprint comparison, managed restoration verified, rejected/no-promotion
2026-09-19 VoiceChat Apple feasibility stop Whole tool-capable replacement blocked before candidate acquisition; disk arithmetic passes but memory is unresolved; protected audio diagnostics are not candidate latency or quality
2026-09-19 GLM-5.3-Flash r10 APC promotion follow-up Explicit human approval and fresh bounded acceptance after the retained r9/r10 APC campaign; shared-prefix TTFT −67.84%, unique-prefix request throughput −2.39% within the frozen gate, and exact r9 rollback retained
2026-09-19 Swift / stock Qwen3.8 Apple artifact feasibility stop Read-only Apple M4 Max storage-policy stop for uncached Swift/stock Qwen3.8 GGUF pairs; memory containment unresolved; no download, launch, benchmark, cleanup, or promotion
2026-09-17 GLM mixed 3.5-bpw startup Host RAM exhaustion during loading shut down the desktop session; zero candidate inference requests; exact 4-bpw baseline restored; retry requires containment
2026-09-17 GLM mixed 3.5-bpw bounded fix-forward Loader recovery passed C1/C4 functional and repeated quality checks plus a 216K actual-token needle; strict capacity was inconsistent and the baseline remains selected
2026-09-13 GLM Flash EXL3 327K intelligence and context qualification Selected text-only primary after fixed quality, 258K context, agentic, SWE, promotion, and bounded client acceptance; Qwen rollback declared, fresh boot/reboot tests unrun
2026-09-12 Qwen3.8 efficient variants on RTX 5090 Four pinned fine-tuned/pruned variants plus no-spec controls; all functional gates pass, strict output and bounded quality failures retained; Signal/Swift9/10 versus incumbent8/10 on ten questions, but no qualified replacement; exact262K incumbent restored
2026-09-11 GLM v0.4.3 qualification Historical native 524K/C4 qualification and promotion acceptance; 14/14 Pi tasks at the then-current 4K output cap
2026-09-08 M4 Max local voice refresh Historical same-host MLX candidate evidence published September 12: strict output/tool/JSON failures retained, no LLM promotion, separate Kokoro runtime-update receipt; no current deployment claim
2026-09-09 GLM-5.3-Flash ormandj v0.4.2 runtime qualification Pinned candidate runtime has faster bounded 4K/120K matched cells and direct gates, but strict 60-request turnover remains 58/60 then 57/60; user selected retain-baseline/no-promotion and the exact baseline was restored
2026-09-09 GLM-5.3-Flash native NCCL P2P transport A/B P2P enabled on the pinned native TP2/C1 profile after bounded 4K/120K latency evidence and routed restoration; both 380K strict capacity cells and the raw PowerShell diagnostic marker failure remain explicit limitations
2026-09-08 GLM-5.3-Flash native Linux versus Windows/WSL Same pinned model/image and physical dual-PRO pair; historical-style C1 decode +20.5–33.0% across 4K–380K targets, endurance 60/60 at +39.6%; coding 15/15, images 12/12, agentic 30/30, official SWE smoke 1/1; extended context 128/150 through 376,484 actual tokens with 9 empty and 13 incorrect answers; strict output and canary failures excluded, routed context/reasoning limits retained; whole-stack comparison, no-promotion
2026-09-04 Qwen3.8 27B RTX PRO 6000 comprehensive optimization campaign Current SGLang K/chunk/compile/Mamba/target/topology matrix plus kelnei/vLLM MTP2/no-spec: matched sustained-output N100 measured one TP1 at 764.3, TP2 at 587.9, and two TP1 replicas at 1,401.8–1,423.4 aggregate tok/s with 100/100 canaries; TP2 strict JSON failed twice and is rejected; RadixArk is the lower-TTFT tradeoff; kelnei MTP2 improved 59.7% over no-spec; unique 82K/C8 remained non-interactive; exact GLM service and route restored, no-promotion
2026-09-03 Qwen3.8 27B RTX 5090 quant bakeoff Fourteen-arm managed quant/speculation comparison after a process-free idle baseline; Gittensor target-only SGLang won TTFT, CometKim won decode but failed tools, new Unsloth Dynamic V3 MTP3 led clean 64K speculation, exact incumbent restored, no-promotion
2026-09-03 2026-09-03-qwen38-ninfer-nvfp4-rtx5090.md Matched Qwen3.8 27B NInfer NVFP4 qualification on one RTX 5090: MTP3 raised median decode from 75.3 to 165.9 tok/s with near-flat 0.421/0.430-second TTFT, returned the exact marker from a 201,746-token prompt with 8,192-token completion reserve, passed bounded text/tool quality, retained explicit 17/20 C1 overload and 2,354 MiB-free policy caveats, restored the GGUF incumbent, and made no promotion or route change
2026-09-03 2026-09-03-qwen38-ninfer-nvfp4-rtx5090-feasibility.md Preregistered interval-based RTX 5090 feasibility screen for the Qwen3.8 27B NInfer NVFP4 252,928-token/C1 profile: benchmark-survivor under an explicit 1 GiB model-only reserve, with physical demand bounded at 23,564,681,216–33,117,175,808 bytes and every load-bearing source and uncertainty retained
2026-09-03 2026-09-03-glm53-concurrency-capacity-interpretation.md Cross-run GLM-5.3-Flash concurrency interpretation on dual RTX PRO 6000: separates scheduler ceilings, shared/reported KV-token pools, short-request batches, and measured long-context concurrency; distinguishes the current 393K/C1 profile, immediate 524K EXL3/DFlash2 rollback, and historical BrandonMusic lanes; records bounded C2/C4 planning math without claiming a new qualification
2026-09-02 2026-09-02-glm53-sglang-sm120-swe-smoke.md GLM-5.3-Flash repository-agent smoke from an isolated macOS worker: one fixed SWE-bench Verified instance attempted, officially graded, and resolved; 11 routed requests; revision-bound isolated harness environment; ambient mini-SWE-agent and Docker launch defects fixed forward
2026-09-02 2026-09-02-glm53-sglang-sm120-393k-promotion.md Human-approved GLM-5.3-Flash SGLang 393K/C1 promotion on dual RTX PRO 6000: model-only zero-reserve waiver with post-workload telemetry retained, promotion-manifest image/OCR fixture fix-forward, complete direct and routed gates, real Pi/OpenClaw/Hermes acceptance, 2,543 MiB/card after client work, exact 524K rollback preserved, and Mid Mod Pi left unchanged pending a host-local router credential
2026-09-02 2026-09-02-glm53-sglang-sm120-qualification.md Pinned SGLang rc.14 GLM-5.3-Flash W4A16 qualification on dual RTX PRO 6000 under WSL2: four hash-gated fix-forwards, matched adaptive-MTP A/B, selected 245,760-token/C1 profile at 108.57/93.35/95.00 tok/s for nominal 4K/120K/230K, tools 20/20, coding 15/15, image/OCR 12/12, endurance 60/60, 3,487 MiB post-workload reserve per card, larger-envelope rejections, and exact incumbent restoration with no promotion
2026-08-31 2026-08-31-glm53-xgrammar-524k-qualification.md GLM-5.3-Flash fix-forward replacement on dual RTX PRO 6000: digest-pinned xgrammar correction, matched no-spec/DFlash2 K5 A/B at 524K, 83.08 tok/s at 4K and pooled 69.99 at 240K, 2/2 concurrent 250K-class prompts, strict structured/tool/quality/media gates, exact 1M rollback drill, and real OpenClaw/Hermes/Pi forward acceptance
2026-08-30 2026-08-30-anvil-serving-1.0.0-release-readiness.md Major-release readiness for the six-family Anvil Serving umbrella, executable product journeys, first-class Voice and Media boundaries, Fleet JSON diagnostic preservation, package/artifact gates, and publication-only closure with no serving or fleet deployment
2026-08-30 2026-08-30-media-envelope-parsing-corrections.md Client-side parsing corrections from a live Hermes Media run: nested job envelopes, native MCP image content blocks, authenticated artifact retrieval, and preservation of the server contract
2026-08-30 2026-08-30-glm53-k3-dflash2-1m-optimization.md Human-authorized one-week GLM-5.3-Flash default on dual RTX PRO 6000: exact EXL3 K3 target plus DFlash2 K5 at 1M/c16 with image/OCR, exact 950K-target retrieval, 82.1/67.4/67.9 tok/s at 4K/131K/240K, K3 and batch-chunk A/Bs, real Hermes/Pi/OpenClaw acceptance, and bounded exact cache cleanup
2026-08-29 2026-08-29-glm53-cardillo-purtell-qualification.md Exact GLM-5.3-Flash TR3/EXL3 4 bpw qualification on dual RTX PRO 6000 under WSL2: vision/OCR fixed-K5 at 72.8/55.7 tok/s for 4K/128K, 250K-target / 206,296-actual retrieval, tools 20/20, bounded coding 15/15, near-500K no-spec context, adaptive-tool corruption rejection, 0xSero 3.0-bpw watch decision, verified recipes, direct hands-on service, and no promotion
2026-08-28 2026-08-28-hermes-image-quality-production.md Production enablement of the exact FLUX.2 Klein image workflow through real Hermes and the authenticated MCP gateway: fixed draft/standard/high profiles, eight independently reviewed images with six strict passes and two retained draft failures, layered latency telemetry, exact cold resume/teardown, twenty-two defects fixed forward, fail-closed overrides, split controller credentials, native artifact verification, Qwen preservation, and video still quality-blocked
2026-08-28 2026-08-28-media-gateway-live-validation.md Exact merged media gateway live validation across Primary Node, Mid Mod, and Mini: real Hermes MCP image/video jobs, A2A replay, authenticated artifact checks, cold approval/unavailable controls, two prompt-adherent FLUX.2 images, one decodable but perceptually failed Wan2.2 video, complete teardown, and no promotion
2026-08-28 2026-08-28-media-gateway-release-readiness.md Source-merge readiness for the unified MCP/A2A media gateway, managed worker, pinned image/video workflows, and Hermes skill: complete source/package gates passed, isolated worker rollback passed, live enablement remains human-required with no route, workflow availability, promotion, client, or serving-state change
2026-08-28 2026-08-28-comfyui-media-qualification.md Exact managed ComfyUI qualification on one RTX 5090: FLUX.2 Klein 4B FP8 produced a decodable 512×512 PNG at 12,919 MiB peak; Wan2.2 TI2V 5B produced a 17-frame H.264 MP4 at 18,263 MiB peak; clean rollback, quality human-required, no route or promotion
2026-08-26 2026-08-26-qwen38-flash-next-vision-promotion.md Qwen3.8 Flash Next vision promotion and comprehensive dual-PRO TP=2 benchmark: direct media 30/30, live routed repeats 57/60 strict with retained literal-rubric misses, admission/SSE/tool edges 8/8, six-size context/throughput sweep through 245K actual prompt tokens, 516,032-token KV pool, and Hermes/Pi/OpenClaw vision convergence
2026-08-26 2026-08-26-qwen38-flash-next-qsa-fast-mtp3-promotion.md Fix-forward RadixArk Qwen3.8 Flash Next QSA-fast/MTP3 text Primary promotion: exact SM120 patch provenance, matched no-spec A/B, 154.9 tok/s at 4K and 134.1 at 128K, full-reserve capacity, bounded quality, and fresh Hermes/Pi/OpenClaw 262K acceptance
2026-08-26 2026-08-26-qwen38-flash-next-promotion.md Human-authorized RadixArk Qwen3.8 Flash Next NVFP4 text Primary promotion on dual RTX PRO 6000 TP=2: exact revision/image pins, 253,325-token direct and routed retrieval with 8,192-token reserve, bounded quality, tools 20/20, Responses, real OpenClaw/Hermes/Pi acceptance, and the SM120/WSL2 fix-forward record
2026-08-24 2026-08-24-anvil-serving-0.35.1-release-readiness.md Release-candidate verification for inference-owned model metadata, the explicit capability-meta-router product decision, synchronized 0.35.1 source/local-image defaults, package artifacts, and publication-only closure with no route, promotion, client, or live fleet change
2026-08-22 2026-08-22-anvil-serving-0.35.0-release-readiness.md Release-candidate verification for routed model evaluation, context and long-tool qualification, recipe feasibility screening, serving-agnostic routing guidance, and the Qwen3.8 llm.secondary evidence; package publication only, with no route, promotion, or fleet deployment change
2026-08-21 2026-08-21-qwen38-27b-gguf-250k-rtx5090.md Managed llama.cpp Qwen3.8 27B Q4_0/no-spec versus Q4_0/MTP3 qualification on one RTX 5090: exact retrieval through 253,822 actual prompt tokens with 8,192-token reserve, tools after 110K, 18/18 images, 3/3 101-turn endurance, +50.7% short decode, mathematical Q6_K+MTP disqualification, exact baseline restoration, and promotion deferred
2026-08-21 2026-08-21-deepseek-v4-flash-0731-infernal-r18-1m-promotion.md Human-approved Infernal Invocation r18 1M Primary promotion on dual RTX PRO 6000 TP=2: exact digest/source pins, 1,040,063-token retrieval, agentic and 160/160 tool gates, matched K5/no-spec A/B, authenticated routed acceptance, Mini generation-2 context/compaction convergence, real Hermes/Pi/OpenClaw passes, and explicit automatic-rollback P1 caveat
2026-08-21 2026-08-21-qwen38-27b-rtx5090-recipe-research.md Extensive current-source RTX 5090 recipe research plus matched MTP3/ReplaySSM qualification: decode +80.5% at 4K and +67.9% at 64K, but only 70,231 KV tokens and 1.9% slower 64K end-to-end; EXL3, NInfer, TurboQuant, alternative NVFP4, GGUF, Reddit, X, and independent-site leads normalized; exact 128K baseline restored, no route or promotion change
2026-08-21 2026-08-21-qwen38-27b-radixark-nvfp4-dflash2-rtx5090.md RTX 5090 NVFP4 DFlash2 diagnosis: exact high-throughput/float32 arm exposed 24,347 KV tokens; BF16, 0.945 memory, disabled radix/prefill graph, and one Mamba slot raised the safe ceiling to 70,262 and passed a 49,549-token retrieval plus tools 20/20, but still failed the 128K contract; exact stock baseline restored, no route or promotion change
2026-08-20 2026-08-20-qwen38-sharp-template-ab.md Same-weight RTX 5090 Qwen3.8 27B stock-versus-Sharp v22.1 chat-template A/B: Sharp functional pass, unchanged 24/30 MMLU-Pro diagnostic with +10.8% completion tokens and +10.7% latency, -5.1% tokens on a thinking-disabled behavior lane but one failed ambiguity check, exact stock restoration, and rejected/no-promotion decision
2026-08-17 2026-08-17-qwen38-27b-radixark-nvfp4-rtx5090-128k.md Retained direct 128K RadixArk Qwen3.8 27B NVFP4 qualification on one RTX 5090: 119,675-token retrieval, tools 20/20, multimodal 30/30, eight-image/two-video boundary 4/4, corrected invalid corpus expectations retained, 64K rollback, and no route or promotion
2026-08-17 2026-08-17-qwen38-27b-radixark-nvfp4-rtx5090.md Single-RTX-5090 RadixArk Qwen3.8 27B NVFP4 qualification: exact digest-pinned SGLang recipe, 60K retrieval and tools 20/20, direct image/OCR/video pass, deterministic multimodal 30/30 including temporal and mixed-media cases, LFM2.5-VL auxiliary comparison, FP8-KV warning, and no-promotion decision
2026-08-16 2026-08-16-deepseek-v4-flash-0731-infernal-r15-393k-promotion.md Human-approved Infernal Invocation r15 393K Primary promotion on dual RTX PRO 6000 TP=2: matched K5/no-spec A/B, 351,118-token direct and 340,119-token routed retrieval, tools and repeated agentic passes, fixed-port r33 rollback, retained estimator/reasoning/client caveats, and full upstream author credit
2026-08-16 2026-08-16-qwen38-27b-video-router.md Current official-FP8 SGLang Qwen3.8 video qualification and managed router-only expansion: direct 30/30 with video 14/14, live admitted 28/28, fail-closed two-image/one-video limits, SSE/tool/error edges, Primary regression, and no model restart
2026-08-15 2026-08-15-qwen38-27b-agentic-swe-scout.md Router-only AI-MBP25 qualification of the current Qwen3.8 27B SGLang service: agentic smoke 2/2, agentic scout 16/18 with both failures isolated to debug-loop repetitions, and fixed five-instance SWE-bench Verified scout resolved 5/5 under the official grader; bounded evidence, no promotion change
2026-08-15 2026-08-15-qwen38-27b-sglang-fp8-single-promotion.md Human-approved single-card official-FP8 SGLang promotion at TP=1/393K/MTP 3/1/4 with FP8 E4M3 KV and CPU multimodal transport: guarded 108K/tools gate, direct and routed 18/18 media, routed Responses correction, Hermes/OpenClaw acceptance, and the second RTX PRO 6000 left empty
2026-08-15 2026-08-15-qwen38-27b-sglang-consolidation-ab.md Matched SGLang BF16, official-FP8, and Inferact-NVFP4 consolidation A/B at TP=1/393K/MTP=3: all three passed a repeated 18-request image/OCR/chart/UI/spatial/two-image corpus, quantized candidates materially improved multimodal and text latency, official FP8 selected as the preferred single-service challenger, exact current split restored, and no promotion
2026-08-15 2026-08-15-qwen38-27b-sglang-mtp-multimodal-qualification.md Matched SGLang MTP=3 qualification for official FP8 and Inferact NVFP4 on both RTX PRO 6000 placements: 111.3 versus 98.1 decode tok/s, complete functional and repeated quality gates, 389K retrieval, and bounded image/OCR recovery through explicit CPU feature transport; exact current split restored, no promotion
2026-08-15 2026-08-15-qwen38-27b-sglang-nvfp4-qualification.md Matched SGLang TP=1/393K official-FP8 versus Inferact NVFP4 qualification on two RTX PRO 6000 lanes: audited Safetensors, full functional and bounded-quality passes, 388,979-token retrieval, five-run and cross-card speed comparison, retained WSL2 multimodal limitation, exact current-split restoration, and no-promotion result
2026-08-15 2026-08-15-qwen38-27b-mtp-depth-qualification.md Official-FP8 Qwen3.8 27B MTP=4/5 qualification on both RTX PRO 6000 lanes: complete functional and near-393K passes, repeated deterministic quality, cross-card swap proving lane variance, exact MTP=3 split restoration, and no-promotion result
2026-08-15 2026-08-15-qwen38-27b-external-recipe-refresh.md Current vLLM, SGLang, Hugging Face, Reddit, X, and community recipe refresh for Qwen3.8 27B on RTX PRO 6000: dormant official-FP8 MTP=4/5 A/B recipes, official-weight SGLang compatibility plan, third-party artifact exclusions, and no live or promotion change
2026-08-14 2026-08-14-qwen38-27b-split-promotion.md Human-approved official Qwen3.8 27B split promotion: FP8 MTP=3 text Primary plus BF16 MTP=3 multimodal/OCR at matched 393K, routed 30/30 media, one 32-image request, relay thinking-control fix, and Hermes/OpenClaw client acceptance without fallback
2026-08-14 2026-08-14-qwen38-27b-tp-mtp-context-matrix.md Official Qwen3.8 27B BF16/FP8 TP=1 versus TP=2 matrix on dual RTX PRO 6000: matched MTP=3 controls, 10-request 4K latency/decode, cold retrieval at 388,979, 598,729, and 985,107 actual prompt tokens, engine KV accounting, PCIe/P2P caveat, exact split restoration, and no-promotion boundary
2026-08-14 2026-08-14-qwen38-27b-1m-context.md Official Qwen3.8 27B FP8 1M-context continuation on one RTX PRO 6000: repeated retrieval through 825,049 actual prompt tokens, steep cold-prefill latency, post-stress gate, exact 262K restoration, current official recipe comparison, and no-promotion boundary
2026-08-14 2026-08-14-qwen38-27b-official-qualification.md Official Qwen3.8 27B BF16 multimodal and FP8 text qualification on co-resident single-card RTX PRO 6000 lanes: artifact-safety hashes, full functional/reasoning/quality/context/capacity gates, BF16 media 30/30, FP8 c5 30/30, MTP=3, prefix-cache, and KV-precision A/Bs, retained warnings, and no-promotion boundary
2026-08-11 2026-08-11-deepseek-v4-flash-0731-r33-393k-promotion.md Human-approved r33 393K Primary promotion on dual RTX PRO 6000 TP=2: direct 359,900-token capacity, exact managed routing, OpenClaw and Hermes aligned to 393K context/32K output/high reasoning with client-path passes, retained routed-needle calibration failure, and unsubmitted SWE smoke due to missing wheel profiles
2026-08-10 2026-08-10-deepseek-v4-flash-0731-r33-batch-token-ab.md Matched r33 batch-token A/B on dual RTX PRO 6000 TP=2: 8,192 to 4,096 reduced peak activation 34.1%, raised minimum-rank KV allocation 0.72 GiB, retained 6/6 functional and 119,503-token capacity passes, reproduced the 8,192 baseline, and confirmed the 4,096 candidate healthy/direct-only at campaign close with a reported-token accounting caveat
2026-08-10 2026-08-10-deepseek-v4-flash-0731-r33-quality-control.md Digest-pinned r33 target-only quality control on dual RTX PRO 6000 TP=2: complete functional and repeated high-reasoning gates, 119,503 actual prompt tokens at 73.86 decode tok/s, context-generator calibration caveat, measured 283,917-token GPU KV ceiling, and a translated-not-loaded 393K FP8-KV plus 16 GiB host-offload candidate
2026-08-10 2026-08-10-deepseek-v4-flash-0731-community-config-refresh.md Two-pass community configuration research for DeepSeek V4 Flash 0731: 11 pinned or gap-labeled dual-RTX-PRO and dual-DGX-Spark candidates, 28 sources, 13 Reddit channel groupings, a quality-first 393K FP8-KV arm, NVFP4 weight/cache separation, runtime and agent-protocol gates, and no-promotion scope
2026-08-07 2026-08-07-deepseek-0731-vision-nvfp4-sglang-first-load.md WebBrain DeepSeek V4 Flash 0731 Vision (NVFP4) first GPU load on SGLang TP=2: marlin/marlin MoE JIT fix, grounded image conditioning, confabulated OCR/GUI quality failures, no chat template, and no-promotion fail-back to the 650K Primary
2026-08-07 2026-08-07-deepseek-0731-vision-nvfp4-recipe-intake.md WebBrain DeepSeek V4 Flash 0731 Vision (NVFP4) source preservation, B200-to-dual-PRO recipe translation, starting-state snapshot, field-report priors, and the bounded qualification decision bar that preceded the same-day first load
2026-08-03 2026-08-03-deepseek-context-agentic-swe-smoke.md Remote AI-MBP25 benchmark-worker qualification against DeepSeek 0731: 8K context pass, retained agentic final-answer failure after successful tool-error recovery, and one officially graded SWE-bench Verified resolution
2026-08-02 2026-08-02-deepseek-v4-flash-0731-primary-promotion.md Human-approved DeepSeek 0731 650K Primary promotion: Dark/Mini Pi and Mini OpenClaw high-reasoning smokes, generic per-tier output clamp with warning, exclusive TP=2 safety, and retained 1M client-shaped B12X workspace failures
2026-08-02 2026-08-02-deepseek-v4-flash-0731-650k-1m-pi-qualification.md DeepSeek 0731 GPU-only 650K/1M qualification after moving display output to the iGPU: 640K and 985K retrieval, matched 32K performance, maxseq1 B12X workspace crash, maxseq4 recovery, maxseq16 1M qualification, memory caveat, and no-promotion decision
2026-08-02 2026-08-02-deepseek-v4-flash-0731-native-kv-offload-256k.md DeepSeek 0731 r16 native KV offload on WSL2: mmap-only pinning translation, 128K store/replay, 256K context through 249,573 prompt tokens, 16 GiB CPU-to-GPU reload, stale-tmpfs attribution, and ownership-aware Anvil lifecycle cleanup
2026-08-01 2026-08-01-deepseek-v4-flash-0731-r16-dspark-qualification.md Official DeepSeek V4 Flash 0731 on the pinned r16 B12X runtime: DSpark K5 versus same-image no-spec A/B, low/high/max reasoning, 128K timing and per-card telemetry, 27/27 coding-agent attempts, durable WSL2 translation, and no-promotion reserve decision
2026-08-01 2026-08-01-deepseek-v4-flash-0731-research-update.md DeepSeek V4 Flash 0731 identity, intelligence evidence, reasoning/tool protocol, vLLM/SGLang and DSpark status, 0731 NVFP4/GGUF conversions, reconciliation with the local dual-PRO TP=2 run, and priority no-promotion qualification plan
2026-08-01 2026-08-01-dual-pro-tp2-model-campaign.md Five-model exclusive TP=2 qualification on two RTX PRO 6000 cards: DeepSeek V4 Flash 0731, Inkling Small, Qwen3.5, Nemotron 3 Super, and Laguna S; pinned recipes, reasoning-aware capacity, NVFP8 search, failures, and no-promotion decision
2026-07-29 2026-07-29-agents-a1-primary-promotion.md Agents-A1 official FP8 passes the complete 262K protocol-v3 gate and becomes the thinking-disabled Primary; Qwen3.5 becomes the immediate managed rollback
2026-07-29 2026-07-29-agents-a1-qwen-262k-head-to-head.md Same-GPU, same-context Agents-A1 official FP8 versus current Qwen3.5 122B NVFP4: matched 8K/240K telemetry, unchanged image/video corpus, runtime video failure, restoration, and no-promotion verdict
2026-07-28 2026-07-28-agents-a1-multimodal-qualification.md Agents-A1 BF16, official FP8, and ProtoLabs NVFP4 qualification on the RTX PRO 6000: text/image/video gates, context/capacity, memory, kernel tuning, routed media admission, and no-promotion decision
2026-07-28 2026-07-28-qwen35-122b-primary-qualification.md Official NVIDIA Qwen3.5 122B NVFP4 at its native 262,144-token window: pinned recipe, single-PRO-6000 gates, current loading research, caveats, and Laguna rollback
2026-07-28 2026-07-28-nemotron35-asr-qualification.md Shared 30-case English STT qualification: Nemotron 3.5 ASR not qualified, Qwen3-ASR 0.6B qualified as an unpromoted replacement candidate, reusable corpus/evidence CLI, and restored protected services
2026-07-27 2026-07-27-omni-voice-stack-qualification.md Co-resident Qwen2.5-Omni-3B, Parakeet STT, and Kokoro TTS on the RTX 5090; measured memory, multimodal gates, Gemma license blocker, and no-promotion caveat
2026-07-27 2026-07-27-omni-stack-qualification.md Exclusive RTX 5090 Omni tier for auxiliary text, image understanding, and OCR; managed lifecycle, capacity, routing, and caveats
2026-07-27 2026-07-27-anvil-serving-release-readiness-sweep.md Full CLI/parser sweep, Agents-A1 and Laguna qualification, purpose-service, voice, ComfyUI, stack, cache, and resolved release fixes
2026-07-26 2026-07-26-laguna-s-heavy-qualification.md Laguna S 2.1 NVFP4 thinking-control diagnosis, repeated 240K quality gate, capacity evidence, and guarded Heavy promotion
2026-07-22 2026-07-22-private-evidence-publication-audit.md #175 inventory, sanitized public publication, offline rerun, and exact missing-mirror record
2026-07-22 2026-07-22-adr-0008-evidence-gap.md Public provenance correction for ADR-0008 raw logs that were never committed and could not be found
2026-07-18 2026-07-18-lifecycle-aware-wsl-cache-reclaim.md Primary Node managed Puzzle Heavy load: 49.9 GiB cache-growth attribution, page-cache-only reclaim, retained VRAM/health/identity/inference, and exact stopped-state restoration
2026-07-18 2026-07-18-gpt-oss-puzzle-heavy-promotion.md Pinned GPT-OSS Puzzle 88B Anvil vLLM fix, RTX PRO 6000 functional and benchmark evidence, default Heavy transition, and Gemma 4 rollback
2026-07-17 2026-07-17-gemma4-31b-optimization.md Current Google 31B QAT template, 128K baseline, native-MTP compatibility failure, and WSL2 implications
2026-07-17 2026-07-17-gpt-oss-puzzle-qualification.md GPT-OSS Puzzle 88B Anvil vLLM port and RTX PRO 6000 qualification evidence without promotion
2026-07-16 2026-07-16-gemma4-vllm0251-wsl2-c128.md vLLM 0.25.1 WSL2 pinned-memory upgrade, V1/V2 Gemma 4 c1/c8/c128 retest, larger-model sweep, and corrected high-concurrency NVFP4 conclusion
2026-07-16 2026-07-16-gemma4-unsloth-nvfp4-follow-up.md Unsloth Gemma 4 NVFP4 12B/26B-A4B/31B Fast/Heavy matrix, direct QAT speed A/B, template/tool regression, and no-promotion result
2026-07-16 2026-07-16-gemma4-chat-template-bakeoff.md July 15 Gemma 4 template matrix on RTX 5090 and PRO 6000, Fast hold, Heavy 12B W4A16 promotion, rollback proof, and raw evidence
2026-07-13 2026-07-13-q36-pro6000-container-recipe.md First physical RTX PRO 6000 build and characterization of the q36 engine: pinned container recipe, context matrix, MTP A/B, smoke, reasoning, and repeated MMLU-Pro evidence
2026-07-13 2026-07-13-e4b-fast-router-promotion.md Gemma 4 E4B fast-tier router promotion, profile reseed (calibration pending), and OpenClaw harness lockstep (gpu-reservations:T007)
2026-07-13 Gemma 4 E4B promotion evidence README Live RTX 5090 preflight, reservation sizing, and promotion evidence inventory
2026-07-13 2026-07-13-e4b-voice-consult-benchmark.md E4B-backed chat-fast voice-consult latency regression that blocked retiring the 35B baseline
2026-07-13 2026-07-13-t011-ocr-rebalance.md OCR bring-up and RTX 5090 resident-set rebalance with routed validation
2026-07-13 2026-07-13-t013-vision.md Vision serve/preset bring-up, first evictable reservation, routed proof, and eviction validation
2026-07-13 2026-07-13-t015-resident-set.md Live RTX 5090 full resident-set, ledger, health, and eviction-drain validation
2026-07-12 2026-07-12-thinkingcap-heavy-promotion.md ThinkingCap FP8 model-aware functional/quality gates and guarded Heavy promotion with GPT-OSS rollback
2026-07-12 2026-07-12-green-context-mps-capability.md Read-only Green Context/MPS inspector, successful Docker Desktop prerequisite probe on the RTX 5090, and unexecuted creation plan
2026-07-12 docker-desktop-rtx5090-prerequisite.json Raw Docker Desktop CUDA 13.1 prerequisite evidence for the UUID-selected RTX 5090; no context or workload created
2026-07-12 2026-07-12-qwen36-protocol-v2-comparison.md Repeated protocol-v2 Qwen3.6 comparison, budget audit, Unsloth NVFP4 v0.25 recipe, five-session validation, and selected resident Heavy quality challenger
2026-07-12 2026-07-12-rtx-pro-6000-heavy-eval-v2.md Repaired repeated ARC/MMLU-Pro Heavy comparison and Laguna NVFP4 vLLM/SGLang sm_120 rejection
2026-07-12 2026-07-12-heavy-intelligence-challengers.md Mistral Small 4 and Nemotron 3 Super single-PRO-6000 Heavy gates, five-session comparison, and selected resident experiment
2026-07-12 2026-07-12-qwen36-27b-heavy-variation-bakeoff.md Qwen3.6-27B NVFP4, official FP8, and ThinkingCap FP8 Heavy validation, five-session capacity, and selected resident candidate
2026-07-12 2026-07-12-qwen36-27b-eval-baseline.md Qwen3.6-27B NVFP4+MTP current built-in eval baseline and invalid-for-ranking session-derived suite control
2026-07-12 2026-07-12-gpt-oss-120b-deterministic-recheck.md GPT-OSS-120B conventional benchmark and deterministic-eval token-budget control
2026-07-12 2026-07-12-nemotron-puzzle-recheck.md Nemotron Puzzle 75B Heavy-candidate preflight, standard benchmark, and deterministic session-eval recheck
2026-07-12 2026-07-12-qwen35-122b-mxfp4-benchmark.md Single-RTX-PRO-6000 Qwen3.5-122B MXFP4/Marlin throughput and deterministic session-eval result (do not promote)
2026-07-11 2026-07-11-system-observability-overhead.md Strict observability overhead and benchmark-effect gate
2026-07-11 2026-07-11-system-observability-artifact-contract.md Synthetic contract validation for external raw telemetry and a sanitized manifest
2026-07-10 2026-07-10-blackwell-local-model-bakeoff.md RTX PRO 6000 and RTX 5090 local-model bakeoff vs production baselines: Nemotron text/Omni, Gemma 4 31B, Ornith 35B, MiniMax M2.7 REAP, DeepSeek V4 Flash — plus the 2026-07-11 extension (Nemotron Puzzle 75B + Qwen3.6-27B with verified MTP speedups, Qwen3.5-35B and Gemma E4B on llama.cpp)
2026-07-10 scorecard.csv Machine-readable bakeoff scorecard (per-candidate config, gates, throughput, verdict)
2026-07-10 2026-07-10-qwen35-122b-heavy-candidate.md Qwen3.5-122B-A10B-NVFP4 heavy-tier candidate (primary-node)
2026-07-10 heavy-tier-bakeoff-evidence/qwen35-122b-a10b-vllm-nvfp4-131k.bakeoff.json Raw heavy-tier bakeoff evidence — Qwen3.5-122B-A10B-NVFP4
2026-07-08 2026-07-08-voice-latency-final-recommendation.md Voice latency final recommendation (voice-latency-model-ab:T007)
2026-07-08 2026-07-08-fast-tier-llm-bakeoff.md RTX 5090 Fast-tier candidate registry, source-backed priors, scoring rubric, and local-gate plan
2026-07-08 2026-07-08-fast-tier-promotion.md Human-gated Qwen3.6-35B-A3B-NVFP4 Fast-tier promotion and validation record
2026-07-08 2026-07-08-stt-model-benchmark.md Dark-host STT benchmark: Parakeet, Qwen3-ASR, and Whisper Turbo
2026-07-08 2026-07-08-voice-latency-ab-final-report.md OpenClaw Talk voice latency candidate A/B status report (evidence synthesis)
2026-07-08 2026-07-08-voice-latency-candidate-matrix.md Voice latency candidate benchmark matrix (T005)
2026-07-08 2026-07-08-openclaw-talk-live-validation.md OpenClaw Talk live validation evidence (T006)
2026-07-07 2026-07-07-voice-latency-model-shortlist.md Voice LLM candidate shortlist for OpenClaw Talk latency (T002)
2026-07-07 2026-07-07-voice-latency-baseline.md Anvil Voice latency baseline for OpenClaw Talk (T001)
2026-07-07 2026-07-07-openclaw-colo-interaction-benchmark.md OpenClaw COLO interaction benchmark — live pass from the Companion Node gateway
2026-07-07 2026-07-07-anvil-score-prd-scope-gap.md Anvil score --prd scope gap, confirmed in Anvil 0.4.2
2026-07-06 2026-07-06-openclaw-workbench-skill-smoke.md Live Companion Node smoke check for the workbench skill
2026-07-06 2026-07-openclaw-anvil-voice-option-live.md OpenClaw Anvil Voice live validation
2026-07-06 2026-07-openclaw-anvil-voice-option-live.json Raw T008 live-validation result record (pass, Companion Node)
2026-07-06 2026-07-openclaw-anvil-voice-option.md OpenClaw Anvil Voice option discovery
2026-07-06 2026-07-openclaw-anvil-voice-gateway-smoke.json Sanitized T008 gateway smoke result
2026-07-06 2026-07-openclaw-anvil-voice-gateway-status.json OpenClaw gateway/service status snapshot (Companion Node)
2026-07-06 2026-07-openclaw-anvil-voice-mini-validation.json Companion Node host-identity validation snapshot
2026-07-06 2026-07-openclaw-anvil-voice-plugin-inspect.json Anvil Voice plugin runtime inspect output
2026-07-06 2026-07-openclaw-anvil-voice-realtime-process.json Mini realtime/audio server process listing
2026-07-06 2026-07-openclaw-anvil-voice-talk-catalog.json OpenClaw Talk modes/transports/providers capability catalog
2026-07-06 2026-07-openclaw-anvil-voice-talk-config.json OpenClaw Talk config snapshot (anvil realtime provider)
2026-07-06 2026-07-voice-tts-ab.md Voice TTS candidate preflight: Kokoro-82M, Orpheus-3B, Qwen3-TTS (T009)
2026-07-06 2026-07-voice-tts-ab.json Raw TTS A/B measurements
2026-07-06 tts-ab-kokoro-5090-20260706.json Kokoro TTS benchmark run on the RTX 5090
2026-07-06 tts-ab-kokoro-current-20260706.json Kokoro TTS benchmark run, current serve config
2026-07-06 tts-ab-orpheus-current-20260706.json Orpheus-3B TTS benchmark run
2026-07-06 tts-ab-qwen3-current-20260706.json Qwen3-TTS benchmark run
2026-07-05 2026-07-voice-stt-ab.md Voice STT A/B: parakeet.cpp vs vLLM-served Whisper (primary-node)
2026-07-05 stt-ab-live-20260705.json Raw STT A/B live run (cold)
2026-07-05 stt-ab-live-warm-20260705.json Raw STT A/B live run (warm)
2026-07-04 2026-07-04-openclaw-keyless-failover.md OpenClaw keyless failover: does the exhaustion-503 hand off to the native subscription? (T005)
2026-07-04 2026-07-04-hf-speech-to-speech-review.md Architecture review of huggingface/speech-to-speech (voice-pipeline PRD input)
2026-07-04 2026-07-04-voice-pipeline-v1-status.md voice-pipeline v1 build status and pre-bring-up punch list
2026-07 2026-07-voice-independent-verification.md Voice pipeline independent verification gate (T017, passed)
2026-07 2026-07-voice-local-loop-proof.md Voice local loop proof: mic → VAD → STT → anvil LLM → TTS → speakers (T010)
2026-07 2026-07-voice-realtime-proof.md Voice Realtime proof: official openai SDK client against the anvil Realtime server
2026-07 2026-07-voice-16gb-mini.md Voice on a 16 GB Mini: local STT+TTS with the LLM routed to primary-node (T016)
2026-07 2026-07-voice-16gb-mini.json Raw evidence for the 16 GB Mini proof
2026-06-29 2026-06-29-harness-intent-routing.md Dated multi-harness feasibility research for model-name-as-intent routing, with version-dependent limits
2026-06-28 2026-06-28-planning-capability-eval.md Historical Anvil PRD-to-tasks evaluation with complete bounded prompts, outputs, judge records, and reproducible offline aggregates
2026-06-28 2026-06-28-anvil-integration-audit.md Pinned Anvil integration audit: one planning endpoint, no fleet or two-endpoint router
(running) blackwell-sm120-lab-notebook.md Blackwell sm_120 lab notebook: which models serve (and how) on primary-node

2026-09-19: Huihui 64K C1 promotion

The earlier Huihui NInfer MTP3 qualification used 65,536 context tokens at C1. Direct preflight passed 7/7, routed preflight 6/6, vision 12/12, and native Hermes 9/9 with the exact local alias. Descriptive, canary-free C1 short-input capacity passed 12/12 at 182.1 mean decode tokens/s; the long-input cell passed 6/6 at 167.1, with 60,769-60,776 actual prompt tokens. Post-workload GPU use was 24,224 MiB. The 32K results remain historical evidence. Endurance, interactive browser acceptance, and a matched 64K no-spec control remain unmeasured. See the 64K finding.