Friction log¶
- CLI discovery: documented
eval benchmark runalias is absent in this checkout. Use the registeredeval benchmark capacitycommand. No live request occurred. Documentation discrepancy retained for publication. -
Managed Serving MCP wrappers are not exposed in this session. Use the product CLI from the verified worktree; no raw lifecycle commands.
-
Direct preflight general-image check failed its unchanged exact phrase assertions (STATUS READY and GPU 42 PERCENT); OCR and all other checks passed. Native failure retained in direct-preflight-disabled.json. Inspect output and owner logs, then repeat the unchanged image check and run the independent image corpus; do not count a retry as erasing the failure.
-
Image failure diagnosis: the retained answer contains STATUS and READY, and GPU and 42 PERCENT in separate Markdown table cells. Exact contiguous-string assertions fail. OCR passed with the identical fixture hash. This is an exact-output compliance failure, not evidence of wrong visual values or an engine crash. Preserve the original check unchanged and use the independent image corpus for broader interpretation.
- The native server reports version 0.0.0.dev1+g033446bb05; historical modality metadata labels an engine source revision 4c2c169b53dbf362f0cd95111f4ae275cd0167c1. Image digest identity matches exactly, but do not equate these different revision fields or invent source attestation. Retain image digest as immutable image identity; the multimodal build-ref field requires the separately verified OCI source revision.
-
MacBook noninteractive SSH PATH omits anvil-serving. The installed executable at the recorded user-local tool path works and reports version 1.0.0. Use its exact Python module launcher for the matched SWE job; no PATH or installation change.
-
Managed recipe status lacks restart count, OOMKilled, and start-time fields. Used a narrow read-only Docker inspection for exactly those lifecycle diagnostics and immutable container/image IDs; retained privately. Existing product gap is documented in the Windows promotion finding. No raw Docker lifecycle operation was used.
-
Strict 128-word controlled-output scout failed 3/3; all performance metrics were correctly withheld. Retained controlled-scout-4k-r3.json. No finalist run using that workload is authorized by the scout gate. Diagnose count/canary capture before any alternate predeclared diagnostic workload; never pool retries or failed samples.
-
Unique natural-answer C1/4K diagnostic returned a request-canary failure; the launcher stopped before its dependent context cells. Retain native incomplete population and withhold comparison metrics. This is a model exact-prefix compliance limit, not a reason to disable the canary validator.
-
Quality launcher rejected a native preflight file used directly as control evidence (schema mismatch), before requests. Fixed by retaining a control-evidence/v1 index linked to both reasoning-off/on native proofs. Invalid launch transcript retained; quality retry independently validates the declared schema.
-
Multimodal launch rejected a digest in engine-build-ref before sending requests: this field requires an exact 40-character source ref. A narrow read-only OCI revision-label inspection verified a547c90c74f1363920287eb80adc88a16d1e7005. Retry uses that build ref and separately retains the image digest. Original argument rejection retained in image-corpus.log.
-
Generic benchmark evidence inspector rejects the native multimodal-benchmark-evidence/v1 artifact. Raw native JSON was independently read: 12/12 pass with exact corpus/model identity. The capacity inspector succeeds. Record this schema-coverage gap in the campaign ticket; do not convert the modality artifact.
-
Routed 380K-target needle fails HTTP413 before upstream dispatch; direct identical input succeeds. Authoritative router logs identify over-context admission. Current source estimates the filler at 534375 tokens from 2137500 bytes, above 393216; the model reports fewer actual tokens. Preserve failed routed gate; use a separately labeled 260K target for routed acceptance and direct endpoint for the extended context-quality suite. Investigate installed router context-admission parity without changing it during this OS baseline.
-
Routed context cause confirmed: live installed config has upstream context admission, but enabled media admission checks estimated text+visual+output against the same limit even for text-only input. Bounded read-only installed-config inspection and installed source establish this second gate. 260K routed retrieval passes. Durable gap ticket records the competing policies.
- Routed thinking-enabled checks produce correct answers but no retained reasoning channel; required-reasoning evidence fails. Installed protocol projection lacks reasoning_content preservation. Preserve direct-pass/routed-fail distinction; tracked in the same gap ticket.
Disposition ledger¶
- Open, deferred to the linked campaign ticket: exact output/canary compliance, image phrase sensitivity, routed text admission and reasoning projection, missing CLI alias documentation, multimodal inspector schema coverage, and lifecycle diagnostic fields. The benchmark preserves the starting deployment to keep the migration baseline interpretable.
- Closed for this campaign: exact MacBook launcher selected without changing PATH; control-evidence index validated; multimodal build-ref corrected to the inspected OCI revision. Invalid launch records remain retained.
-
Native context controls: the adapter uses its default thinking behavior despite the job parameter. Context quality is labeled separately; no reasoning-disabled performance claim is made from it.
-
Repository verification initially stopped on the new finding missing from the measured-hardware mention audit; publication adds the explicit classification. A subsequent native trusted-file test rejected the newly created worktree root because its inherited mode was group-writable (0775). Tightened only that disposable worktree root to0755; focused trusted-file tests pass16 with5 platform skips. No shared parent permissions or serving state changed. Full verification reruns against the corrected test workspace.
-
Repository full run reached7211 passes with one launcher failure because the checkout MCP registration invokes
python, absent from the ambient PATH. Adding the prepared validation virtualenv bin directory only to the test command PATH resolves the focused contract6/6. No global provider or MCP configuration was modified; the full suite reruns with this explicit environment. -
Documentation build caught four repository-relative links outside the site docs tree. Replaced recipe links with the exact repository commit URL and kept the local ticket as a literal path beside the public failure ledger; Markdown link validation passes after staging new owned artifacts.
-
Open, deferred: durable context status emits no per-case progress or partial native observations during the150-case sweep. Bounded owner logs prove ongoing prefill, not completed benchmark counts. Added a managed progress/partial-evidence acceptance section to the campaign ticket; do not infer a progress percentage from telemetry.
Completed context and closing checks¶
The native sweep completed 128/150, with 9 empty length-terminated and 13 incorrect visible answers. Its non-monotonic threshold-derived 8K field is not a physical limit. No matched Windows arm exists. Post-workload routed smoke/JSON passed; configuration fingerprints, model identities and container state were unchanged. The direct catalog dynamic created timestamp changed; all other fields matched. Bounded owner-log filtering returned no matching output; this does not assert complete historical log coverage.
Chart publication review¶
Visual inspection found overlapping axis/legend text and near-equal series labels in the shared renderer. The chart renderer now reserves separate footer rows, offsets paired value labels and centers a single category. Both retained chart packs were regenerated twice, independently reconciled with native metric values and hashes, and visually inspected. Existing renderer regression gates pass; no model requests were needed.