Publication summary: GLM-5.3-Flash and bounded dual-PRO qualification¶
This file is derivative publishing copy. The linked dated finding and raw artifacts are authoritative.
Canonical facts¶
- Model identity:
brandonmusic/GLM-5.3-Flash-tr3-4bpw@5ab363a8dcf6405955fd5f99671e01a1c9fb124b; preferred served nameglm53-flash-tr3-4bpw-tp2-262k-fixed-mtp5-vision - Runtime identity:
vLLM 0.1.dev20051+g487ecf187; pinned Purtell image digestsha256:da5cec95778bf6996660b52e28a6e51737fec69cfc3d508bf298c8a89f273ac5 - Local setup: 2x RTX PRO 6000 Blackwell Max-Q 96 GB, TP=2 over PCIe, Docker Desktop/WSL2, vision fixed K5, 262K, c1 performance
- Recipe: managed vision fixed-K5 262K recipe
- Measurement path: warm direct OpenAI-compatible API; no router or client path
- Headline result: vision fixed K5 decoded at 72.8 tok/s at 4K and 55.7 tok/s at 128K, each from three c1 low-reasoning repetitions
- Capability result: image understanding, verbatim OCR, tools 20/20, bounded high-reasoning coding 15/15, and exact retrieval at a 250K target / 206,296 actual prompt tokens passed; the no-spec companion recovered an exact needle at 495,045 prompt tokens
- Important caveat: video is disabled, only one-image behavior is proven, adaptive K1-K5 plus ReplaySSM corrupted repeated tools, and 0xSero 3.0-bpw has no complete serve path and failed its publisher's held-out quality gate
- Decision:
challenger,no-promotion; no route, client, or promotion state changed - Canonical evidence: https://github.com/fakoli/anvil-serving/blob/main/docs/findings/2026-08-29-glm53-cardillo-purtell-qualification.md
X / short post¶
Dual-PRO GLM-5.3-Flash: vision/OCR, tools 20/20, 262K context, 72.8 tok/s. 3-bpw watch-only. No promotion. https://github.com/fakoli/anvil-serving/blob/main/docs/findings/2026-08-29-glm53-cardillo-purtell-qualification.md
Reddit¶
I reproduced the Cardillo/Purtell GLM-5.3-Flash recipe on two RTX PRO
6000 Blackwell Max-Q cards under WSL2. Vision fixed K5 passed semantic image
understanding, verbatim OCR, 20/20 tools, exact retrieval at a 250K target /
206,296 actual prompt tokens,
and a bounded 15/15 high-reasoning coding suite. It reached 72.8 tok/s at 4K
and 55.7 at 128K. Video remains disabled. The no-spec 524K profile recovered
an exact needle at 495,045 prompt tokens and retained 1,603,111 reported KV
tokens. Adaptive MTP corrupted repeated tool output. The 0xSero 3.0-bpw lead
was not downloaded because its own release ledger lacks a complete server and
records a failed held-out quality gate. GLM remains unpromoted pending hands-on
and routed client acceptance. Full immutable identities, recipes, failures,
raw artifacts, and runtime-state proof are in the canonical finding:
https://github.com/fakoli/anvil-serving/blob/main/docs/findings/2026-08-29-glm53-cardillo-purtell-qualification.md
What matched or differed on your hardware?
Screenshot alt text¶
Benchmark result card for GLM-5.3-Flash TR3/EXL3 4 bpw on two 96 GB RTX PRO 6000 Blackwell Max-Q GPUs. Vision fixed five-token MTP passed image, OCR, tools, bounded coding, and a 250K-target / 206,296-actual retrieval while reaching about 73 output tokens per second at 4K. A no-speculation profile passed exact retrieval near 495K prompt tokens with three full windows of reported KV capacity. Adaptive MTP is marked rejected for tool corruption, 0xSero 3.0-bpw is watch-only, and the overall model remains no-promotion.