- "Search space mirrors glm5.2-fp4-b300-sglang-agentic-mtp exactly so the two SKUs are comparable: one TP8 + HiCache arm at conc [1, 4, 8, 12, 16]. Steps of at least 2 because single-step sampling cannot separate configurations by more than run-to-run noise on the agentic corpus, and a hard stop at conc 16. TP8-only for memory as well: the ~433 GB NVFP4 checkpoint needs ~54 GB/GPU across 8 B200s and does not fit below 8. The DEP (attention-DP + EP8) throughput arm is not wired up for the same reason as on B300 -- its frontier peak sits well above the conc-16 cap -- but the branch, including the --speculative-moe-a2a-backend none / --speculative-moe-runner-backend triton pair that GLM-5.2-NVFP4's unquantized bf16 nextn layer needs under expert parallelism, is kept in the script."
0 commit comments