Skip to content

[AMD] update model in case qwen3.5-fp4-mi355x-sglang - #2465

Draft
yuzho-amd wants to merge 2 commits into
SemiAnalysisAI:mainfrom
yuzho-amd:main
Draft

[AMD] update model in case qwen3.5-fp4-mi355x-sglang#2465
yuzho-amd wants to merge 2 commits into
SemiAnalysisAI:mainfrom
yuzho-amd:main

Conversation

@yuzho-amd

@yuzho-amd yuzho-amd commented Aug 3, 2026

Copy link
Copy Markdown

Update the model in case qwen3.5-fp4-mi355x-sglang from amd/Qwen3.5-397B-A17B-MXFP4 to amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2.
This model has quantized attention to PTPC FP8 and shared experts to MXFP4.

This is the end-to-end performance results on MI355:
image

And it needs these PRs from aiter and sglang:

repo PR summary status
aiter ROCm/aiter#4346 fix allreduce acc issue merged
aiter ROCm/aiter#4396 add ptpc tuned config approved
sglang sgl-project/sglang#29723 fuse allreduce+rmsnorm+inputquant waiting for review
aiter ROCm/aiter#4001 update fusedmoe tuned config merged
aiter ROCm/aiter#4025 update fusedmoe tuned config merged

@yuzho-amd yuzho-amd changed the title update model in case qwen3.5-fp4-mi355x-sglang [AMD] update model in case qwen3.5-fp4-mi355x-sglang Aug 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant