Skip to content

Friction log

  • Skill review initially held on owner/source identity and CLI/input-test coverage. Fix: scoped instructions and real CLI subprocess plus invalid/duplicate/path tests. Independent skill recheck accepted; packet retained.
  • New command declaration initially left generated CLI manifest stale. Focused command-tree test caught it. Fix: native write_manifest() followed by CLI-reference regeneration; command-tree tests and full CLI audit passed. No live request was made on that source.
  • No callable serving operation-contract/recipe MCP wrappers discovered in this tool context. Managed Anvil Serving CLI used with explicit owner/endpoint and retained status instead. No raw Docker mutation.

Offline validation corrections

Independent runner review blocked live use until managed identity checks, true client-observation overlap boundaries, tokenizer/input/stream allocation caps, complete request budgets, hard client process deadlines, failure classes and evidence validation were implemented. Final targeted suite: 271 passed; final full-suite identity/results are retained in development-source.json. Shell Python lacked Markdown dependencies; reran with the existing development environment. New skill source link needed an absolute published URL for strict MkDocs. Both documentation failures resolved. Full evidence remains diagnostic and no-promotion.

Crash log capture bound

First capture requested tail=6000 and was rejected before reading logs (supported maximum 5000). Retried at 5000, preserving the full bounded owner log in crash-175k-overlap-model.log. No lifecycle action occurred before successful capture.

Live evidence fixes

  • HTTP200 streamed engine errors and stream IDs were omitted by the initial runner. The parent native artifact remains unchanged; separate owner/kernel correlation binds its failure. Stream-error and exact-ID retention plus real-transport malformed-ID regressions now pass independent review (270 targeted tests).
  • A DCP1 small-case output cap was mistakenly 512 rather than matched 2048. Preserve its no-visible-answer failure; the runtime stayed healthy. The completed result was lost when visible-answer validation raised. Fix-forward: retain bounded returned result before validation; a reasoning-only real-CLI regression asserts usage and finish remain present. Independent review also corrected its taxonomy to semantic_output_absent; valid terminal streams are not malformed protocol. The 271-test suite and Ruff pass. Corrected small scenario restores baseline 2048/1024. No larger request advanced before this correction passed.
  • Managed log filtering applies to stdout while stderr can bypass it. Always capture both streams into a private file, then inspect a bounded excerpt. Keep this operational invariant in the notebook; no raw Docker fallback.
  • Container-specific status and registry inventory expose different field sets. Use registry inventory for recipe/registry identity and container-specific status for containment; both raw captures retained. The stability runner already uses the identity-bearing inventory.
  • One test command used nonexistent test filenames and ran zero tests. Discovered actual files; the subsequent correctly scoped suite passed 271 tests. This error did not authorize live traffic.

Final process dispositions

The qualification skill now requires a mechanical normalized scenario diff, including output caps and reasoning policy, before traffic. The first 512-cap attempt remains failed/non-comparable; the corrected small run has a new scenario identity. The independent follow-up accepted that transfer control. The declared time guard refused the larger 201K case before requests and the campaign restored the baseline. No silent budget extension or unrecorded probe.

Status selectors accept model IDs or an explicit --container, not a container name as MODEL. Two final read-only selector attempts returned no match/help; the corrected model-ID query retained the identity-bearing status. JSON command envelopes must be unwrapped before comparison; a failed offline assertion was corrected before any restoration receipt was written. These are command/data handling invariants, not model failures. No live action depended on those errors.

The CUDA fault remains unresolved at the kernel level. The durable disposition is the retained crash correlation and the experimental DCP1 trial, not a claim that recovery fixed it. Further source attestation and first-fault investigation remain explicit follow-up gates in coverage-and-gaps.md.

Publication checks caught a missing hardware mention-audit entry and private container IDs before upload. The audit entry is now explicit. Public container IDs use consistent one-way redaction tokens with receipts; the read-only evidence reader labels sanitized provenance and preserves strict live identity checks. Real transport and live-rejection regressions passed in the 272-test suite. The first finalizer invocation used unsupported size-policy wording; the documented exact grammar is now retained in manifest source. No evidence producers ran after finalization. Build tools were absent from the shared test environment; isolated uv run --no-project --with build --with twine produced the wheel and clean-install smoke, without changing the host installation.

Independent publication review found that the read-only summary called absent identity observations “recorded.” The final source uses explicit sanitized, recorded and unverified states, with missing-identity and malformed/mixed incomplete-identity regressions. Each claimed provenance now requires all observations to match its exact token form. The final 272-test gate passed; the preceding full local suite passed 9,596 tests with 53 skips. Revision boundaries are explicit in development-source.json; PR-head CI remains the final release check.