Skip to content

Qwen3.8 27B PRO 6000 possibility campaign friction log

Time Category Command or stage Observation Workaround Durable follow-up
2026-09-05T00:49:00Z ambiguous-output installed CLI resolution anvil-serving --version resolved to the public worktree package, but the initially attempted eval benchmark run action is not part of the current CLI; the correct action is eval benchmark capacity inspected focused help and recorded the current action none; documentation and examples already use the current action
2026-09-05T00:58:00Z unsafe-default pre-mutation live snapshot both target GPUs are owned by an existing exclusive TP=2 serve, so a candidate cannot be loaded without an interruption stopped before mutation and required a managed preview plus explicit human gate retain the transition and rollback proof in private evidence
2026-09-05T01:03:00Z missing-identity current SGLang image tag the cookbook exposes a mutable tag, which is insufficient for a reproducible recipe resolved the OCI index, amd64 manifest, creation time, FlashInfer/CUDA versions, and source label without pulling all executable recipes use the immutable digest
2026-09-05T01:06:00Z ambiguous-output RadixArk BF16-LM-head metadata query an abbreviated revision was rejected and a stale PowerShell value could have been misread as a valid response requeried the model endpoint, captured the full revision, and repeated the tree calculation never reuse an API response variable after a failed request
2026-09-05T01:17:00Z unsafe-default managed mode leave preview a guessed split restore group was correctly rejected because it did not match the active profile's rollback router model; the valid rollback group is another TP=2 serve and frees no GPU retained the valid exact restore preview and stopped before mutation a TP1 campaign requires an explicit Primary-interruption gate or a separately authorized reroute design
2026-09-05T01:31:00Z ambiguous-output target-only startup admission 16 configured Mamba state slots looked compatible with C8, but five slots were consumed per request and the runtime capped active requests at C3 added an otherwise matched 96-slot managed recipe and reran C8 treat the runtime's startup admission calculation as authoritative; configured max-running alone is not effective concurrency
2026-09-05T01:58:00Z ambiguous-output 82K/C8 shared-prefix sweep a warm repeat was dramatically faster than unique-prefix and first-repeat workloads retained all three conditions as separate artifacts and graph categories benchmark graphs and summaries must never collapse warm shared-prefix and unique-prefix traffic
2026-09-05T02:07:00Z failure preserved baseline container restart the image's startup patch verifies and modifies a file, so restarting the already-patched preserved container failed its checksum unloaded the failed preserved container and used a clean managed recreation paired private recipe now has strict pristine/already-patched hash states plus a regression test; candidate-operations skill requires proven restart idempotency before --keep-container
2026-09-05T02:12:00Z failure managed baseline recreation and router readmission the service reached health, but the router transaction returned HTTP 401 because the declared credential was not loaded into that command process loaded only the required user-local credential into the process and retried through the managed mode surface restoration preview must include router credential presence when readmission is in scope
2026-09-05T02:17:00Z failure direct recipe restoration the exact original service was healthy, but direct recreation left the router expecting the rollback identity and the routed alias returned 503 re-entered the exact original mode with authenticated readmission direct model health is not restoration; verify mode ledger, expected/observed route identity, and routed acceptance
2026-09-05T02:22:00Z ambiguous-output routed restoration smoke the correct short coding response ended with length, failing the harness's default finish-reason policy preserved the policy-failure artifact and reran the restoration-only probe with stop,length allowed keep strict finish-reason policy for qualifications; document explicit restoration exceptions
2026-09-05T03:05:00Z measurement-shape SGLang optimization screens natural short completions made aggregate throughput useful for within-screen selection but not comparable with sustained-output or external headline workloads added a 100-request, unique-canary, 256-word/512-token sustained-output workload for finalists publication separates screening and headline workloads and records their completion-distribution boundary
2026-09-05T03:25:00Z negative-result K12 torch.compile arm compile measured 361.6 aggregate tok/s versus 421.7 for the matched uncompiled K12/2K screen, a 14.2% regression retained the result and stopped advancing compile keep compile as an exact-runtime A/B, never a presumed optimization; cold compile duration was not retained and remains a telemetry gap
2026-09-05T03:42:00Z ambiguous-output Mamba state allocation the official C8 minimum of 40 state slots tied the 96-slot short screen but did not improve the retained 82K/C8 result kept 96 slots for the finalist so short and long evidence use the better observed envelope future recipes calculate active minimum explicitly, then require a divergent-prefix long-context gate before reducing slots
2026-09-05T04:10:00Z compatibility TP2 startup the pristine current SGLang path entered an incompatible symmetric-memory/logits-gather route under WSL2 used the recipe's checksum-gated ordinary-NCCL fallback, disabled custom all-reduce, and retained conservative NCCL controls TP2 evidence is specific to this patched WSL2 path; a native-Linux runtime remains a separate qualification
2026-09-05T04:28:00Z correctness-failure TP2 structured JSON full preflight emitted a valid JSON object, a literal </think> delimiter, and a duplicate object; an isolated repeat reproduced the failure rejected TP2 even though tools, Responses, and 100/100 performance canaries completed retain both failed JSON artifacts; any revisit must fix the parser/output path and pass repeated structured-output gates before performance matters
2026-09-05T04:55:00Z topology-result TP2 versus two TP1 replicas TP2 delivered 587.9 aggregate tok/s; two synchronized TP1 replicas delivered 1,401.8–1,423.4 at aggregate C16 with per-request latency close to TP1 selected DP2 as the bounded throughput winner DP2 still needs a managed load-balancer, admission, health, and failover contract before any client-facing use
2026-09-05T05:30:00Z tradeoff RadixArk target RadixArk K8 cut median sustained-load TTFT 46.9% versus Inferact but reduced aggregate throughput 2.3% and increased median E2E 6.8% retained RadixArk as the lower-TTFT tradeoff and Inferact as the sustained decode/E2E selection publish target choice as a workload tradeoff, not a universal checkpoint winner
2026-09-05T06:05:00Z alternate-runtime kelnei vLLM MTP2 MTP2 improved matched throughput 59.7% over no-spec and counters proved the drafter active, but the result remained behind the SGLang finalist retained both matched artifacts and the counter delta keep alternate-runtime claims paired with their exact no-spec control; do not inherit publisher performance numbers
2026-09-05T06:30:00Z missing-telemetry cross-arm hardware telemetry power, clocks, temperature, host-memory pressure, cold startup, compile duration, and energy/token were not captured consistently for every arm excluded those metrics from comparative claims add a synchronized telemetry collector to a future runner revision before making efficiency claims
2026-09-05T08:10:00Z restoration-validation managed mode entry output validation was blocked because two unrelated adjacent promotion evidence directories expected by the validator were absent created only the two expected empty directories, then repeated the managed validation; no unrelated evidence was copied or changed a new ticket records that mode-entry validation must not depend on unrelated adjacent promotion output directories
2026-09-05 measurement-correction independent DP2 review legacy replica timestamps have whole-second precision; the longest local duration alone does not prove the union window retained raw inputs unchanged and published the 1,401.8–1,423.4 tok/s timing bound combiner retains conservative bounds, validates precise clock intervals for new runs, and rejects duplicate input populations; adversarial regression tests cover these contracts
2026-09-05 evidence-safety independent publication-tool review derived-output paths could collide with source evidence, and permissive inputs could mislabel or inflate measurements held merge until source/output, schema, identity, and numeric validation were hardened report, graph, finalizer, and replica-combiner regression suites cover malformed inputs and source preservation
2026-09-05 rendered-UX recipe report browser check raw HTML evidence links resolved incorrectly under MkDocs directory URLs and metric text wrapped poorly corrected Markdown-aware link containers and responsive metric layout desktop/mobile browser navigation, local-link status, console, and overflow checks pass; report freshness is checked in docs CI