Skip to content

Campaign plan v1

User priority: useful Pi coding/reasoning at preserved intelligence, tool and visual reliability, voice/laptop stability. Six-bit matched GGUF/llama.cpp Metal first; matched Q4 only scouting if Q6 unsafe. Shared runtime/quant/projector policy/template/parser/sampling/context/seed, speculation disabled. Exact identities are unresolved pending research.

No weights or inference until disk and memory/load containment pass. Disk policy initially 10 GiB free reserve plus 1 GiB evidence/runtime allowance; downloader temporary behavior must be established rather than assumed. Unified-memory policy reserve 16 GiB unchanged absent measured startup and representative workload evidence. No cache deletion, service interruption, launchd changes, route changes or reboot authorized. Isolated candidates only; one downloader and one candidate at a time.

Stages: research, feasibility, scout, finalist, quality, restoration, publication. Investigation uses one smallest evidence-supported configuration change per version; retain every failure. No artificial overall time cutoff: stop only at user approval boundary or a documented physical/runtime/evidence barrier after supported alternatives assessed. Reserve closure work before asking approval. Build concurrency ceiling two workers, no full-core compilation; do not build before storage gate.

At frozen finalists use 3 repetitions of identical 12 text/tool and 15 visual tasks, matched run order and repair allowance (one retry, all attempts retained). Critical coding/tool/identity/parser/isolation/cancel/stability success 100%. Overall accepted Swift >=95% stock, no >5% relative quality loss; visual gate separately >=95% with zero critical false-confident numeric errors. Discrete suite counts and uncertainty reported; repeated tasks are not independent unique tasks.

Efficiency: >=20% fewer generated reasoning+visible output tokens and >=20% lower median task E2E OR >=1.20x successful tasks/hour. Include repair cost and failures in denominator. Reasoning/visible splits require authoritative tokenization or explicit missing status, no text-character proxy. Report truncation avoidance separately from intelligence.

Capacity first 24000 input +8192 reserve C1; larger 64000+8192 only if first passes with safe measured headroom. Finalist timing cells: response_words=128, max_tokens=8192, strict output policy, seed=42, explicit unique/shared cache populations, request canaries for unique cells; identical supported native reasoning/sampling policy chosen from exact upstream/runtime evidence. Cold and warm remain separate. No p99 claim below 100 requests/population.

Voice baseline must measure actual STT/TTS/voice-LLM request latency before candidate load and continuously during a bounded 15-minute finalist soak. Default ceiling <=20% p95 degradation per service versus same-workload baseline, zero failures, zero swap growth, no restart. Baseline sample count and nearest-rank method recorded. Health/catalog alone never qualify voice or candidate.

All validators independent from tested models. No real user repository or account data in corpus. Promotion remains human gated and cannot replace primary, voice, STT or TTS. No suitable auxiliary alias is assumed. Client changes only in exact final promotion packet.