Skip to content

Clarify all-sequence policy training for Qwen3.5 GRPO - #597

Merged
ShriyaRishab merged 3 commits into
mlcommons:masterfrom
al-rigazzi:arigazzi/qwen35_397b_grpo_zero_advantage_filtering
Sep 30, 2026
Merged

ShriyaRishab merged 3 commits into
mlcommons:masterfrom
al-rigazzi:arigazzi/qwen35_397b_grpo_zero_advantage_filtering

Conversation

@al-rigazzi

@al-rigazzi al-rigazzi commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Clarify the existing requirement that every completed Qwen3.5 GRPO training sequence that is otherwise eligible under the benchmark rules must be included in the policy-training batch and processed by the policy forward pass and loss computation.
  • Prohibit skipping, removing, replacing, or resampling sequences based on their actual or expected loss or gradient contribution, including zero advantage, equal rewards within a rollout group, trajectory-importance-sampling masks, or equivalent mechanisms.
  • Clarify that grpo.use_dynamic_sampling = false retains zero-reward-variance groups rather than discarding and replacing them.

Clarification of existing rules

This change makes explicit the intended scope of the existing Qwen3.5 GRPO requirements—use_dynamic_sampling = false, a fixed 16 generations per prompt, and the reference GRPO training method—by requiring every otherwise eligible generated sequence to remain in the policy-training computation, including when its advantage or TIS-masked loss contribution is zero. It does not change the permitted masking of gradient contributions by the reference loss.

Permitted advantage or importance-sampling masks may reduce a sequence's gradient contribution to zero, but they may not be used to remove the sequence before policy training or reduce the number of sequences presented to the policy-training computation.

Rationale

Predictable zero-contribution sequences must not be exploited to reduce measured training work. In particular, the frequency of zero-advantage groups is an artifact of the benchmark's fixed dataset and limited rollout count and does not reflect typical training conditions.

Scope

This change is intentionally separate from #596, which covers Qwen3.5 GRPO offline evaluation. If #596 merges first, this branch will be rebased so that both rules share one benchmark-specific appendix entry.

Validation

  • git diff --check
  • AsciiDoc-to-HTML render
  • Repository-wide audit for conflicting Qwen3.5 sequence-filtering rules

@al-rigazzi
al-rigazzi requested review from a team as code owners September 17, 2026 14:55
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@al-rigazzi al-rigazzi changed the title Prohibit zero-advantage sample filtering for Qwen3.5 GRPO Clarify all-sequence policy training for Qwen3.5 GRPO Sep 17, 2026
…b_grpo_zero_advantage_filtering

# Conflicts:
#	training_rules.adoc
@RissyRan

Copy link
Copy Markdown

Thanks! Checking with team and see if any questions on this update.

@RissyRan RissyRan left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Look good to us. Thanks!

@ShriyaRishab
ShriyaRishab merged commit 67f2c26 into mlcommons:master Sep 30, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants