Shadow Jev evidence pilot¶
Run from the directory containing a reviewed jev-pilot.json:
Each run requires a new output directory. The command retains one receipt per packet immediately, including the full-context baseline, stable word-overlap baseline and Jev proposal. It never executes diagnostics, benchmarks, grades, route changes, cancellation or promotion. It is outside benchmark timing and has no automatic-consumption mode.
Install Anvil separately. Use its protected TYPESAFE_API_KEY mechanism;
credentials never belong in this config. jev uses the existing Serving policy
contract, including an absolute trusted anvil_binary, pinned jev-1.13.0,
enabled, allow_api, allow_export, and capabilities context_ranking and
incident_triage. API and export permission plus --allow-export are all
required. --no-jev makes no provider calls.
{
"schema": "anvil-serving.jev-pilot/v1",
"packets": ["selected-case.json"],
"output": "shadow-run-001",
"jev": {
"enabled": false,
"capabilities": ["context_ranking", "incident_triage"],
"allow_api": false,
"allow_export": false,
"anvil_binary": "/opt/anvil/bin/anvil",
"model": "jev-1.13.0"
}
}
Paths resolve relative to the config. Select 1–64 unique packets. Packets have
exact fields schema: anvil-serving.jev-decision-packet/v1, id, intent,
observation, optional_limit (1–24), and evidence. Each evidence row has
id, selected sanitized text, kind, artifact and sha256. Artifact
references are never opened or exported by the command. The owner must verify
them and retain access to every original; the digest binds the referenced
artifact, not the truth of its contents. A missing_capture row alone may
have a null digest. Missing originals must remain explicit notices.
Owner-classified mandatory, contradiction, failure, and missing_capture
rows are always pinned. Only optional rows may be ranked. Incorrect owner
classification remains a risk: independent labels and review must check it.
There are at most 24 optional rows; overlarge input fails instead of truncating.
The stable deterministic baseline ranks by distinct word overlap with intent.
Jev scores reorder optional IDs; tied scores at the cutoff keep all evidence
and flag ambiguity. Known unambiguous error signatures avoid incident calls.
Unknown or conflicting signatures use the closed incident vocabulary and
reviewed diagnostic templates from the existing Anvil bridge. No confidence
threshold is treated as calibrated proof.
The command binds packet and policy digests and checks both around calls. Revoked or changed inputs invalidate recommendations. Provider failures restore the full view and escalation path; disabled calls use the deterministic comparator. Unfamiliar incidents remain escalated even with a semantic category suggestion; every attempt and fallback remains in the receipt. Anvil validates typed responses; Serving independently validates capability IDs and provenance. Source instructions never become tool authority.
Compare before enabling anything¶
Freeze independently labeled tuning and held-out cases before calls. Replay all three retained prompts with the same downstream model, settings, tools, context budget and outcome oracle; retain actual provider usage and failures. Count preprocessing, all Jev calls, downstream calls and fallbacks. The command reports selection latency, not overall decision latency or token savings. Missing cost/usage is unknown, never zero. A valid failed-call usage receipt counts even when its answer was rejected. Report total tokens, cost where known, median/p95 decision latency, escalation, critical recall, selection and category errors and downstream correctness. Do not infer benefit from Jev's latency alone or from shorter character counts.
Issue #242's initial target is 20% lower median downstream input tokens versus the stronger baseline, with no total cost or median latency regression and no held-out critical miss. Unknown billing prevents cost acceptance. Keep shadow mode and the simpler baseline unless all measured criteria pass. A future automatic consumer needs its own reviewed restore path and sampled audits.