GLM EXL3 vision enablement on dual RTX PRO 6000 Max-Q¶
Later same-day update: Eight-image acceptance passed and the configured limit is now eight images per request. The original one-image results below remain retained.
Date: 2026-09-18. Evidence: functional. Decision: current, user-authorized vision enablement; preceding one-image vision configuration retained for rollback.
Result card¶
GLM-5.3-Flash 4-bpw EXL3 passed bounded image and routed client acceptance on two RTX PRO 6000 Blackwell Max-Q 96 GB GPUs under native Linux.
| Setup | Value |
|---|---|
| Model | brandonmusic/GLM-5.3-Flash-tr3-4bpw@a5fee929cf4888b1824323e33e8a19b60129e025 |
| Runtime | Pinned v84 vLLM image; FP8 DS-MLA KV; speculation disabled; TP2/EP2/DCP2 |
| Contract | Configured 327,680 tokens, C4 scheduler, eight images/request, no video |
| Memory containment | 52 GiB host RAM ceiling, no container swap, 12 GiB admission reserve |
| Recipe | Current eight-image reconstruction |
| Measurement | Direct and routed online acceptance after cold load; warmed kernels; mixed image-cache state |
| Measurement | Result | Boundary |
|---|---|---|
| Eight-image comparison | 4/4 requests | Eight cards, forward/reverse order, two repetitions, C1 |
| Eight high-resolution images | 1/1 request | 2520×2520 each; 63,456 input tokens; 28.17 seconds |
| Initial one-image corpus | 10/10 | r8: five synthetic cases, two repetitions, C1 |
| Initial r8 direct preflight | 7/7 checks | Includes 20/20 tool calls, streaming, continuation, image and OCR |
| Initial r8 routed preflight | 4/4 checks | Text, streaming tools, image and OCR |
| Pi output transcript | Eight cards correct | Request attachments, route and served identity not independently correlated |
| Current r9 preflight | 6/6 checks | Coding, JSON, four tool calls, streaming, continuation, image |
| Initial r8 long context | 231,045-token retrieval and 110,760-token tool call passed | 259,922-token retrieval refused initially, then passed on retry |
Why it matters: the existing local text/tool model can now consume an image, without changing weights, KV precision or speculation.
Important caveat: this is bounded functional acceptance, not renewed broad quality or speed qualification. The 327K maximum remains configured; the 260K semantic probe passed on retry after an initial refusal; no full-window soak was run.
Artifact manifest · Evidence index · Publication summary
Initial one-image configuration and method¶
Served identity: glm53-flash-exl3-4bpw-v84-fp8-tp2-c4-327k-r8-vision-nospec.
Image digest: sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692.
Build ref inherited from the pinned baseline metadata: 6dc2f516688fe6f84c6994dcd20fddf296853a6c; runtime reports 0.1.dev20111+g7f1e92bec.d20260827. No new source-to-image attestation was performed.
The recipe removes --language-model-only, selects the bundled multimodal chat template and limits images to one, video to zero. Other inference flags retain the baseline: FP8 KV, 2048 batched tokens, C4, 0.970 GPU allocation, chunked prefill, prefix cache disabled, GLM reasoning/tool parsers and no speculation. Router admission estimates 8192 tokens/image, above the startup encoder profile's 7921-token budget. Startup completed with zero container OOM events. The boot owner was updated; reboot was not tested.
The single-image corpus excludes the original two-image comparison. Its five cases cover a chart, text, scene labels, UI screenshot and spatial counting. Native assertions score visible output independently. Image and long-context requests overlapped; timings are retained for audit but are not comparative performance measurements. No prior intelligence score was rerun or attributed to this new configuration.
Initial one-image results and limitations¶
Direct preflight, image attempts, routed preflight establish bounded image/text/tool acceptance. The Pi response is an output-only transcript; request attachments and routing are not independently correlated.
Long-context checks passed at 231,045 and 110,760 actual input tokens. The additional 259,922-token test returned a refusal identifying the synthetic launch-code request as prompt injection. It completed normally, without a memory failure. The user reported an overlapping request and requested a retry. The same 259,922-token fixture passed in 58.9 seconds. No matched baseline or concurrency-controlled pair establishes the cause of the initial refusal. Do not equate the configured maximum with validated answer quality or simultaneous full-window capacity.
Companion-client catalog synchronization covered OpenClaw, all managed Hermes profiles and Pi; a repeat made no changes. A matching local Pi output was observed, but full local catalog refresh was blocked by its pre-existing 16K compaction reserve versus the router's 64K output ceiling. Its compaction settings were preserved. Controller recovery and packaged-CLI memory-limit parity are recorded in friction.
Final state¶
Vision routes and image capabilities were activated under the user's explicit request. After the eight-image follow-up, the preceding one-image vision recipe is the immediate rollback. Final-state evidence records healthy admission, an active/enabled boot owner and zero OOM events. This is dated evidence; active operator assignments remain private.
Eight-image follow-up¶
The user requested testing eight images first. The r9 recipe changes the per-request image limit from one to eight, retaining pinned weights/runtime, FP8 KV, no speculation, configured 327,680 context, C4 and the 52 GiB host-RAM/no-container-swap boundary. The former one-image recipe is the rollback.
- Four requests with eight distinct 768×512 numbered cards passed, including forward and reverse image order, each repeated twice. Independent assertions verified every transcribed value in image order, the minimum, maximum and total.
- One request with eight 2520×2520 cards passed at 63,456 actual prompt tokens, completing in 28.17 seconds. This is a single observed functional timing, not a throughput comparison.
- Six direct preflight checks passed, including coding, JSON, four tool calls, streaming tools, tool-result continuation and a single image.
- The retained Pi output matches all eight cards. It has no correlated request envelope or router record, so it does not independently establish attachment count, route or r9 served identity. Companion catalog synchronization reported no changes; the existing image capability records remain compatible.
- The router limit and active/enabled boot owner now select the eight-image recipe. Final managed status recorded zero OOM/OOM-kill events; the RAM ceiling triggered reclaim events. No full-window or simultaneous eight-image soak, broad visual-quality benchmark or reboot was tested.
Eight-image recipe · Trial plan · Four normal-size requests · Large-image request · Pi response · Summary