| 2026-09-05T00:49:00Z |
ambiguous-output |
installed CLI resolution |
anvil-serving --version resolved to the public worktree package, but the initially attempted eval benchmark run action is not part of the current CLI; the correct action is eval benchmark capacity |
inspected focused help and recorded the current action |
none; documentation and examples already use the current action |
| 2026-09-05T00:58:00Z |
unsafe-default |
pre-mutation live snapshot |
both target GPUs are owned by an existing exclusive TP=2 serve, so a candidate cannot be loaded without an interruption |
stopped before mutation and required a managed preview plus explicit human gate |
retain the transition and rollback proof in private evidence |
| 2026-09-05T01:03:00Z |
missing-identity |
current SGLang image tag |
the cookbook exposes a mutable tag, which is insufficient for a reproducible recipe |
resolved the OCI index, amd64 manifest, creation time, FlashInfer/CUDA versions, and source label without pulling |
all executable recipes use the immutable digest |
| 2026-09-05T01:06:00Z |
ambiguous-output |
RadixArk BF16-LM-head metadata query |
an abbreviated revision was rejected and a stale PowerShell value could have been misread as a valid response |
requeried the model endpoint, captured the full revision, and repeated the tree calculation |
never reuse an API response variable after a failed request |
| 2026-09-05T01:17:00Z |
unsafe-default |
managed mode leave preview |
a guessed split restore group was correctly rejected because it did not match the active profile's rollback router model; the valid rollback group is another TP=2 serve and frees no GPU |
retained the valid exact restore preview and stopped before mutation |
a TP1 campaign requires an explicit Primary-interruption gate or a separately authorized reroute design |
| 2026-09-05T01:31:00Z |
ambiguous-output |
target-only startup admission |
16 configured Mamba state slots looked compatible with C8, but five slots were consumed per request and the runtime capped active requests at C3 |
added an otherwise matched 96-slot managed recipe and reran C8 |
treat the runtime's startup admission calculation as authoritative; configured max-running alone is not effective concurrency |
| 2026-09-05T01:58:00Z |
ambiguous-output |
82K/C8 shared-prefix sweep |
a warm repeat was dramatically faster than unique-prefix and first-repeat workloads |
retained all three conditions as separate artifacts and graph categories |
benchmark graphs and summaries must never collapse warm shared-prefix and unique-prefix traffic |
| 2026-09-05T02:07:00Z |
failure |
preserved baseline container restart |
the image's startup patch verifies and modifies a file, so restarting the already-patched preserved container failed its checksum |
unloaded the failed preserved container and used a clean managed recreation |
paired private recipe now has strict pristine/already-patched hash states plus a regression test; candidate-operations skill requires proven restart idempotency before --keep-container |
| 2026-09-05T02:12:00Z |
failure |
managed baseline recreation and router readmission |
the service reached health, but the router transaction returned HTTP 401 because the declared credential was not loaded into that command process |
loaded only the required user-local credential into the process and retried through the managed mode surface |
restoration preview must include router credential presence when readmission is in scope |
| 2026-09-05T02:17:00Z |
failure |
direct recipe restoration |
the exact original service was healthy, but direct recreation left the router expecting the rollback identity and the routed alias returned 503 |
re-entered the exact original mode with authenticated readmission |
direct model health is not restoration; verify mode ledger, expected/observed route identity, and routed acceptance |
| 2026-09-05T02:22:00Z |
ambiguous-output |
routed restoration smoke |
the correct short coding response ended with length, failing the harness's default finish-reason policy |
preserved the policy-failure artifact and reran the restoration-only probe with stop,length allowed |
keep strict finish-reason policy for qualifications; document explicit restoration exceptions |
| 2026-09-05T03:05:00Z |
measurement-shape |
SGLang optimization screens |
natural short completions made aggregate throughput useful for within-screen selection but not comparable with sustained-output or external headline workloads |
added a 100-request, unique-canary, 256-word/512-token sustained-output workload for finalists |
publication separates screening and headline workloads and records their completion-distribution boundary |
| 2026-09-05T03:25:00Z |
negative-result |
K12 torch.compile arm |
compile measured 361.6 aggregate tok/s versus 421.7 for the matched uncompiled K12/2K screen, a 14.2% regression |
retained the result and stopped advancing compile |
keep compile as an exact-runtime A/B, never a presumed optimization; cold compile duration was not retained and remains a telemetry gap |
| 2026-09-05T03:42:00Z |
ambiguous-output |
Mamba state allocation |
the official C8 minimum of 40 state slots tied the 96-slot short screen but did not improve the retained 82K/C8 result |
kept 96 slots for the finalist so short and long evidence use the better observed envelope |
future recipes calculate active minimum explicitly, then require a divergent-prefix long-context gate before reducing slots |
| 2026-09-05T04:10:00Z |
compatibility |
TP2 startup |
the pristine current SGLang path entered an incompatible symmetric-memory/logits-gather route under WSL2 |
used the recipe's checksum-gated ordinary-NCCL fallback, disabled custom all-reduce, and retained conservative NCCL controls |
TP2 evidence is specific to this patched WSL2 path; a native-Linux runtime remains a separate qualification |
| 2026-09-05T04:28:00Z |
correctness-failure |
TP2 structured JSON |
full preflight emitted a valid JSON object, a literal </think> delimiter, and a duplicate object; an isolated repeat reproduced the failure |
rejected TP2 even though tools, Responses, and 100/100 performance canaries completed |
retain both failed JSON artifacts; any revisit must fix the parser/output path and pass repeated structured-output gates before performance matters |
| 2026-09-05T04:55:00Z |
topology-result |
TP2 versus two TP1 replicas |
TP2 delivered 587.9 aggregate tok/s; two synchronized TP1 replicas delivered 1,401.8–1,423.4 at aggregate C16 with per-request latency close to TP1 |
selected DP2 as the bounded throughput winner |
DP2 still needs a managed load-balancer, admission, health, and failover contract before any client-facing use |
| 2026-09-05T05:30:00Z |
tradeoff |
RadixArk target |
RadixArk K8 cut median sustained-load TTFT 46.9% versus Inferact but reduced aggregate throughput 2.3% and increased median E2E 6.8% |
retained RadixArk as the lower-TTFT tradeoff and Inferact as the sustained decode/E2E selection |
publish target choice as a workload tradeoff, not a universal checkpoint winner |
| 2026-09-05T06:05:00Z |
alternate-runtime |
kelnei vLLM MTP2 |
MTP2 improved matched throughput 59.7% over no-spec and counters proved the drafter active, but the result remained behind the SGLang finalist |
retained both matched artifacts and the counter delta |
keep alternate-runtime claims paired with their exact no-spec control; do not inherit publisher performance numbers |
| 2026-09-05T06:30:00Z |
missing-telemetry |
cross-arm hardware telemetry |
power, clocks, temperature, host-memory pressure, cold startup, compile duration, and energy/token were not captured consistently for every arm |
excluded those metrics from comparative claims |
add a synchronized telemetry collector to a future runner revision before making efficiency claims |
| 2026-09-05T08:10:00Z |
restoration-validation |
managed mode entry |
output validation was blocked because two unrelated adjacent promotion evidence directories expected by the validator were absent |
created only the two expected empty directories, then repeated the managed validation; no unrelated evidence was copied or changed |
a new ticket records that mode-entry validation must not depend on unrelated adjacent promotion output directories |
| 2026-09-05 |
measurement-correction |
independent DP2 review |
legacy replica timestamps have whole-second precision; the longest local duration alone does not prove the union window |
retained raw inputs unchanged and published the 1,401.8–1,423.4 tok/s timing bound |
combiner retains conservative bounds, validates precise clock intervals for new runs, and rejects duplicate input populations; adversarial regression tests cover these contracts |
| 2026-09-05 |
evidence-safety |
independent publication-tool review |
derived-output paths could collide with source evidence, and permissive inputs could mislabel or inflate measurements |
held merge until source/output, schema, identity, and numeric validation were hardened |
report, graph, finalizer, and replica-combiner regression suites cover malformed inputs and source preservation |
| 2026-09-05 |
rendered-UX |
recipe report browser check |
raw HTML evidence links resolved incorrectly under MkDocs directory URLs and metric text wrapped poorly |
corrected Markdown-aware link containers and responsive metric layout |
desktop/mobile browser navigation, local-link status, console, and overflow checks pass; report freshness is checked in docs CI |