Skip to content

GLM-5.3-Flash native Linux migration evidence

Campaign 2026-09-08-glm53-linux-wsl-comparison measures the existing model on two RTX PRO 6000 Blackwell Max-Q GPUs against retained Windows/WSL evidence. Launcher revision: 8abcc5dc75a1b2389210b553120abb76e5864d3a, initially clean tracked worktree; campaign changes are documentation and evidence only. The dated finding explains the bounded functional, capacity and quality results and no-promotion decision.

Campaign controls

Native measurements

Population Retained native evidence Interpretation
Direct disabled/off Full gate, smoke General-image exact phrase failed; other nine checks passed
Direct enabled/on Functional gate, control index Reasoning evidence present with separate diagnostic budget
Historical-style C1 n3 4K, 120K, 262K, 380K Warm online direct, shared prefixes, variable short output; descriptive medians
Retained WSL C1 n3 4K, 120K, 262K, 380K Matching original native bytes; no new Windows arm
Strict output Scout 0/3 exact128-word compliance; all performance withheld
Unique natural answer 4K diagnostic 1/10 first-position canary compliance; dependent depths not run
Coding Quality Five deterministic cases x3, 15/15
Images Corpus Six cases x2, 12/12; native multimodal schema retained
Endurance Linux, WSL 60/60 each; same variable-output/shared-prefix limitations
Routed acceptance Disabled, 260K retrieval, enabled 380K estimate rejects413; reasoning projection evidence fails; visible answers otherwise correct
Deep agentic Native job, 30 observations 30/30 endpoint-only co-resident cases
Extended context Native job, summary 128/150; 9 empty and 13 incorrect; default thinking, no Windows control
SWE smoke Native job, official grader stage, comparison Same one instance resolves1/1; new trajectory is slower and longer

Derived views and boundaries

Comparison data, prompt equivalence, graph manifest, graph data, and context comparison chart trace the historical comparison to native request artifacts. The separate endurance chart, manifest and data compare the n=60 populations. Both chart packs are embedded in the finding; repeated graph renders are byte-identical.

Capacity artifacts retain anvil-serving.benchmark/v1; durable jobs retain anvil-serving.benchmark-evidence/v1; image evidence retains multimodal-benchmark-evidence/v1. Preflight and coding retain their native contracts. Failed and partial populations remain visible. Console logs in this bundle record CLI failures and summaries; native JSON is authoritative.

The migration also changes the driver, transport and cache history. Decode improvements are whole-stack descriptive results, not an OS-only causal result or a strict controlled-output qualification. p99 at n3/n60 is descriptive only. No speculation A/B, video qualification, higher engine concurrency, full SWE score, model promotion or client-catalog change is claimed.