Skip to content

Publication summary: GLM-5.3-Flash native NCCL P2P transport A/B

Canonical facts

  • Model identity: ormandj/GLM-5.3-Flash-W4A16-NVFP4-K32-Experts-FP8-WO@c3cbb9891b67c741bcbf6b176dd7af9265b069db
  • Runtime identity: digest-pinned SGLang image 0c063795; TP=2, 393,216 tokens, C1.
  • Local setup: two RTX PRO 6000 Blackwell Max-Q GPUs; only NCCL_P2P_DISABLE=1 to 0 changed.
  • Headline result: 4K n=12: median TTFT -9.0%, E2E -7.4%, decode +2.1%. 120K n=3: TTFT -8.9%, E2E -8.9%, decode +1.4%.
  • Important caveat: fixed 32-word output, small populations, partial cache metadata, and no high-context performance result.
  • Decision: user-authorized native GLM default retains P2P enabled; no global NCCL or client-contract change.
  • Canonical evidence: dated finding and artifact manifest.

Screenshot alt text

Chart comparing P2P-disabled and P2P-enabled native GLM-5.3-Flash TP=2 C1 cells at 4K and 120K. P2P enabled has lower median TTFT and E2E; the chart does not include the failed 380K strict-output cells.

Claim ledger

Public claim Conditions Evidence
P2P enabled lowered median latency in the eligible cells 4K n=12 and 120K n=3, strict 32 words, unique-prefix C1 comparison
The native default retains P2P enabled authorized bounded recipe decision after routed checks restoration
No high-context performance claim both 380K strict cells returned 33 rather than 32 words coverage

Publication validation

The 19 focused benchmark-docs tests passed, and all relative links passed the 606-file tracked Markdown check. Graph rendering and manifest finalization were byte-identical on repeated runs. A strict MkDocs build passed using the repository's requirements-docs.txt in an isolated tool environment; no runtime serving dependencies were changed.