Skip to content

Add Qwen3.5 GRPO offline evaluation rule - #596

Merged
ShriyaRishab merged 2 commits into
mlcommons:masterfrom
al-rigazzi:al-rigazzi/qwen35_397b_grpo_offline_eval
Sep 21, 2026
Merged

ShriyaRishab merged 2 commits into
mlcommons:masterfrom
al-rigazzi:al-rigazzi/qwen35_397b_grpo_offline_eval

Conversation

@al-rigazzi

@al-rigazzi al-rigazzi commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add benchmark-specific offline evaluation rules for qwen35_397b_grpo and link them from the quality-measure table.
  • Starting at the validation step CEIL(2.5 + 3840 / global_batch_size), require a checkpoint after every training step until training stops, while maintaining full-capacity sample generation and processing.
  • Permit post-run evaluation of saved checkpoints in any order using BF16 or the rollout-generation quantization format.
  • Define run_stop using the passing checkpoint's latest weight-update timestamp, excluding the time required to write that checkpoint, and add an illustrative example.

Validation

  • git diff --check
  • AsciiDoc-to-HTML render

@al-rigazzi
al-rigazzi requested review from a team as code owners September 8, 2026 18:59
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@ShriyaRishab

Copy link
Copy Markdown
Contributor

Meeting notes 9/10/26:

What exactly is included in the timed section? Is it checkpoint saving time for 2 checkpoints or all checkpoints? How frequently should checkpoints be stored (because its very vague now and we want it fixed)?

Comment thread training_rules.adoc Outdated
| |LLM finetuning |Llama2_70B_LoRA| Every 384 sequences; `CEIL(384 / global_batch_size) * global_batch_size` steps if 384 is not divisible by GBS. Skipping first `FLOOR(0.125*global_batch_size+2)` evaluations |v6.1
| |LLM MoE pretraining |deepseekv3_671b| First evaluation at `GBS * FLOOR(42 + 24576 / GBS)` samples; every step thereafter. |v6.1
| |LLM post training |qwen35_397b_grpo| First evaluation at `global_batch_size * CEIL(2.5 + 3840 / global_batch_size)` samples; every training step (i.e., every `global_batch_size` samples) thereafter. |v6.1
| |LLM post training* |qwen35_397b_grpo| First evaluation at `global_batch_size * CEIL(2.5 + 3840 / global_batch_size)` samples; at least one additional evaluation thereafter. |v6.1

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Am I understanding it correctly that we could either 1) eval each step after global_batch_size * CEIL(2.5 + 3840 / global_batch_size) OR 2) at least one additional eval & offline evaluation with 2 checkpoints? Thanks!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@al-rigazzi please address

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have updated the rule and it should be clearer now.

Signed-off-by: Jeremi Piotrowski <jpiotrowski@nvidia.com>
@jepio
jepio force-pushed the al-rigazzi/qwen35_397b_grpo_offline_eval branch from a359205 to e40a25c Compare September 11, 2026 16:23
jepio added a commit to NVIDIA-NeMo/RL that referenced this pull request Sep 14, 2026
Implements the offline-evaluation rules for this benchmark from
mlcommons/training_policies#596:

- DEFERRED_OFFLINE_EVAL=1 disables inline validation and checkpoints every
  step from grpo.val_start_at (plus the final step), weights only.
- Each checkpoint's training_info.json records its weight-update end time
  (training_step_end_time_ms).
- After training stops, a second driver in the same allocation/Ray cluster
  restores each checkpoint without optimizer state, refits vLLM, and runs
  the configured validation in step order, stopping at the first checkpoint
  that reaches the target accuracy.
- run_stop is emitted once, backdated to the passing checkpoint's
  weight-update timestamp (aborted at the final checkpoint's timestamp when
  no checkpoint crosses); operational failures fail the job.

Also fixes the dropped final-step train metrics: the trainers log a step's
validation before its train metrics, and a target-reaching eval emitted
run_stop first, suppressing the final step's tracked_stats. Evals started
while a train block is open are now held and flushed after the step's train
metrics land.

Bumps mlperf-logging to the qwen35-ruleset pin and the default ruleset to
6.1.0.
jepio added a commit to NVIDIA-NeMo/RL that referenced this pull request Sep 14, 2026
Implements the offline-evaluation rules for this benchmark from
mlcommons/training_policies#596:

- DEFERRED_OFFLINE_EVAL=1 disables inline validation and checkpoints every
  step from grpo.val_start_at (plus the final step), weights only.
- Each checkpoint's training_info.json records its weight-update end time
  (training_step_end_time_ms).
- After training stops, a second driver in the same allocation/Ray cluster
  restores each checkpoint without optimizer state, refits vLLM, and runs
  the configured validation in step order, stopping at the first checkpoint
  that reaches the target accuracy.
- run_stop is emitted once, backdated to the passing checkpoint's
  weight-update timestamp (aborted at the final checkpoint's timestamp when
  no checkpoint crosses); operational failures fail the job.

Also fixes the dropped final-step train metrics: the trainers log a step's
validation before its train metrics, and a target-reaching eval emitted
run_stop first, suppressing the final step's tracked_stats. Evals started
while a train block is open are now held and flushed after the step's train
metrics land.

Bumps mlperf-logging to the qwen35-ruleset pin and the default ruleset to
6.1.0.
Comment thread training_rules.adoc
* qwen35_397b_grpo (Qwen3.5-397B-A17B)
** Evaluation is performed offline based on checkpoints. Offline validation requires creating a sequence of checkpoints near the end of training that can be evaluated post-run. Beginning at the `<validation start step>` computed using the benchmark formula `CEIL(2.5 + 3840 / global_batch_size)`, the implementation must save a checkpoint after every individual step until training stops. Each checkpoint must store the timestamp of its latest weight update. Submissions must save at least one checkpoint and have the flexibility to select when to halt training, though stopping at `<validation start step> + 1` is recommended. To maintain consistent system performance and avoid artificial tapering of the workload, training samples must continue to be generated and processed at full capacity right up to the final stopping step.

** Once training halts, saved checkpoints may be evaluated offline in any order until the one that achieves target evaluation accuracy is found. Checkpoint weights must be evaluated using either BF16 or the exact quantization format used during previous rollout generation. To finalize the run, emit an MLLOGGER run_stop event using the weight update timestamp of the passing checkpoint. Because the final score relies on that weight update timestamp, the time spent writing the final checkpoint itself to disk is excluded from the score.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since we run training and evaluation separately, should we maintain two separate mllog files for each, or continue using the same one as https://github.com/mlcommons/training/blob/master/llm_post_training/rcp_logs/GBS512/seed_1.out?

@susanbao susanbao Sep 14, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we meet the accuracy target on step k, we used the weight update timestamp of the passing checkpoint, ignoring the checkpoint saving time for step k, while the checkpoint saving time for step k-1 is accounted for. Is that correct?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we meet the accuracy target on step k, we used the weight update timestamp of the passing checkpoint, ignoring the checkpoint saving time for step k, while the checkpoint saving time for step k-1 is accounted for. Is that correct?

Yes, that's the correct interpretation (unless storing the checkpoint can be done asynchronously, thus hiding its cost).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since we run training and evaluation separately, should we maintain two separate mllog files for each, or continue using the same one as https://github.com/mlcommons/training/blob/master/llm_post_training/rcp_logs/GBS512/seed_1.out?

It should be in a single mllog file for scoring. run_start is emitted from the training part, and run_stop from the evaluation part.
See here: mlcommons/training#906 (comment)

Comment thread training_rules.adoc

** Once training halts, saved checkpoints may be evaluated offline in any order until the one that achieves target evaluation accuracy is found. Checkpoint weights must be evaluated using either BF16 or the exact quantization format used during previous rollout generation. To finalize the run, emit an MLLOGGER run_stop event using the weight update timestamp of the passing checkpoint. Because the final score relies on that weight update timestamp, the time spent writing the final checkpoint itself to disk is excluded from the score.

** For example, if `<validation start step>` is step 18 and submitter continues training through step 20 while maintaining full hardware load, the submission must write distinct checkpoints with update timestamps for steps 18, 19 and 20. During offline evaluation, if step 19 is evaluated first and it meets the accuracy target, the run_stop event using step 19's weight update timestamp is used to compute the score. Instead, if step 18 meets the accuracy target, the run_stop event using step 18's weight update timestamp is used and no checkpoint saving time in included in the final score.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The RCL_Logging will not be changed with the offline eval. We will have block_start -> post-training -> block_stop -> eval_start -> tracked_stats -> eval_accuracy -> eval_stop -> run_stop. The timestamp of run_stop should be the corresponding saved weight update timestamp in the checkpoint. Am I right?

How should we set the timestamp of eval_start, tracked_stats, eval_accuracy and eval_stop in the eval process? If directly from the computer time, the timestamp of eval will be larger than the timestamp of run_stop.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@susanbao

For these events you should emit the computer time: "block_start -> post-training -> block_stop -> eval_start -> tracked_stats -> eval_accuracy -> eval_stop".

The only event that you need to adjust the timestamp for is run_stop. You may emit that one not with the computer time, but with the timestamp of the training step (right after the optimizer.step() is finished) that passes the threshold. This timestamp may be stored together with the checkpoint. This is the code in the updated reference that emits the run_stop with modified timestamp:
https://github.com/NVIDIA-NeMo/RL/blob/f987c0596af2bd6cfb9958e768deb8df951d7504/nemo_rl/algorithms/mlperf_grpo_deferred.py#L372-L381

            mllogger.end(
                key="run_stop",
                metadata={
                    "status": "success" if passed else "aborted",
                    "samples_count": samples,
                },
                time_ms=int(info["training_step_end_time_ms"]),
            )

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The timestamp of run_stop should be the corresponding saved weight update timestamp in the checkpoint. Am I right?

Yes

If directly from the computer time, the timestamp of eval will be larger than the timestamp of run_stop.

Yes, that is fine and correct.

@ShriyaRishab
ShriyaRishab merged commit 895832c into mlcommons:master Sep 21, 2026
2 checks passed
@al-rigazzi
al-rigazzi deleted the al-rigazzi/qwen35_397b_grpo_offline_eval branch September 21, 2026 14:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants