Skip to content

[Perf Regression] 34 config(s) regressed @ 83a22aa7 #2162

Description

@github-actions

Performance Regression Detected

Commit: 83a22aa7
Run: https://github.com/ROCm/ATOM/actions/runs/34143052526
Date: 2026-09-08T08:44:58.221467+00:00

Regressed Configurations

Model ISL/OSL Conc Tput (cur) Tput (base) Δ% TPOT (cur) TPOT (base) Δ%
DeepSeek-R1-0528-MXFP4 1024/1024 4 419.6 422.3 -0.6% 9.27 9.24 0.3%
DeepSeek-R1-0528-MXFP4 1024/1024 8 723.3 727.7 -0.6% 10.60 10.65 -0.5%
DeepSeek-R1-0528-MXFP4 8192/1024 4 364.2 373.4 -2.5% 10.36 10.22 1.3%
DeepSeek-R1-0528-MXFP4 8192/1024 32 1221.0 1229.8 -0.7% 24.61 24.52 0.3%
DeepSeek-R1-0528-MXFP4 MTP3 1024/1024 8 955.9 979.0 -2.4% 7.79 7.80 -0.0%
DeepSeek-R1-0528-MXFP4 MTP3 1024/1024 16 1549.1 1573.7 -1.6% 9.71 9.59 1.3%
DeepSeek-V4-Pro DPA 8192/1024 1024 5166.0 5183.8 -0.3% 179.85 133.10 35.1%
DeepSeek-V4-Pro EPLB r0 MegaMoE MegaMoE 8192/1024 4096 4752.5 5631.5 -15.6% 653.75 509.98 28.2%
DeepSeek-V4-Pro MTP3 1024/1024 8 733.1 806.0 -9.1% 10.59 9.39 12.9%
DeepSeek-V4-Pro MTP3 1024/1024 128 3616.9 3611.0 0.2% 33.92 34.17 -0.8%
DeepSeek-V4-Pro TBO 8192/1024 256 1946.7 2541.2 -23.4% 122.18 95.12 28.4%
GLM-5.2-FP8 1024/1024 16 743.9 743.4 0.1% 20.94 20.99 -0.2%
GLM-5.2-FP8 MTP3 1024/1024 16 1167.4 1179.7 -1.0% 13.04 13.02 0.1%
GLM-5.2-MXFP4 1024/1024 8 597.3 584.4 2.2% 12.92 13.26 -2.6%
GLM-5.2-MXFP4 8192/1024 8 495.5 499.9 -0.9% 15.03 15.04 -0.1%
GLM-5.2-MXFP4 MTP3 1024/1024 32 2250.6 2372.3 -5.1% 13.56 12.90 5.1%
Kimi-K2.7-Code-MXFP4 1024/1024 64 2583.8 2586.4 -0.1% 24.02 24.03 -0.1%
Kimi-K3 1024/1024 16 613.0 607.6 0.9% 25.06 25.40 -1.4%
Kimi-K3 8192/1024 16 503.3 503.9 -0.1% 29.93 30.08 -0.5%
Kimi-K3 8192/1024 128 1185.7 1207.9 -1.8% 101.88 100.42 1.4%
Kimi-K3 DSpark 1024/1024 8 564.7 550.4 2.6% 13.22 13.10 0.9%
MiniMax-M3-MXFP4 1024/1024 64 3537.3 3623.8 -2.4% 17.16 16.92 1.4%
MiniMax-M3-MXFP4 8192/1024 8 816.9 859.8 -5.0% 9.14 8.81 3.8%
MiniMax-M3-MXFP4 EAGLE3 1024/1024 8 1572.2 1563.1 0.6% 4.68 4.82 -2.9%
MiniMax-M3-MXFP4 EAGLE3 1024/1024 64 5310.6 5321.6 -0.2% 11.25 11.37 -1.0%
MiniMax-M3-MXFP8 1024/1024 8 732.2 798.8 -8.3% 10.34 9.49 9.1%
MiniMax-M3-MXFP8 1024/1024 128 4332.9 4321.7 0.3% 28.33 28.49 -0.6%
MiniMax-M3-MXFP8 1024/1024 256 6377.5 6479.3 -1.6% 38.34 38.13 0.6%
MiniMax-M3-MXFP8 EAGLE3 1024/1024 8 1314.0 1291.7 1.7% 5.69 5.87 -3.0%
Qwen3.5-397B-A17B-MXFP4 8192/1024 4 370.4 371.2 -0.2% 10.35 10.37 -0.2%
gpt-oss-120b 1024/1024 8 1556.4 1525.2 2.0% 4.96 5.07 -2.1%
gpt-oss-120b 1024/1024 32 4098.8 4163.6 -1.6% 7.47 7.37 1.4%
gpt-oss-120b 8192/1024 4 852.3 832.6 2.4% 4.51 4.63 -2.6%
gpt-oss-120b 8192/1024 8 1430.0 1443.5 -0.9% 5.32 5.29 0.6%

Performance Summary

Summary not available

Profiler Traces

Download from workflow artifacts.
Open in Perfetto UI or Chrome chrome://tracing for analysis.

Next Steps

  1. Download profiler-analysis-34143052526 artifact
  2. Open trace files in Perfetto UI
  3. Compare kernel durations against previous traces
  4. Identify bottleneck changes

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions