Repository navigation
Add Qwen3.5 GRPO offline evaluation rule - #596
ShriyaRishab merged 2 commits into
Conversation
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
|
Meeting notes 9/10/26: What exactly is included in the timed section? Is it checkpoint saving time for 2 checkpoints or all checkpoints? How frequently should checkpoints be stored (because its very vague now and we want it fixed)? |
| | |LLM finetuning |Llama2_70B_LoRA| Every 384 sequences; `CEIL(384 / global_batch_size) * global_batch_size` steps if 384 is not divisible by GBS. Skipping first `FLOOR(0.125*global_batch_size+2)` evaluations |v6.1 | ||
| | |LLM MoE pretraining |deepseekv3_671b| First evaluation at `GBS * FLOOR(42 + 24576 / GBS)` samples; every step thereafter. |v6.1 | ||
| | |LLM post training |qwen35_397b_grpo| First evaluation at `global_batch_size * CEIL(2.5 + 3840 / global_batch_size)` samples; every training step (i.e., every `global_batch_size` samples) thereafter. |v6.1 | ||
| | |LLM post training* |qwen35_397b_grpo| First evaluation at `global_batch_size * CEIL(2.5 + 3840 / global_batch_size)` samples; at least one additional evaluation thereafter. |v6.1 |
There was a problem hiding this comment.
Am I understanding it correctly that we could either 1) eval each step after global_batch_size * CEIL(2.5 + 3840 / global_batch_size) OR 2) at least one additional eval & offline evaluation with 2 checkpoints? Thanks!
There was a problem hiding this comment.
We have updated the rule and it should be clearer now.
Signed-off-by: Jeremi Piotrowski <jpiotrowski@nvidia.com>
a359205 to
e40a25c
Compare
Implements the offline-evaluation rules for this benchmark from mlcommons/training_policies#596: - DEFERRED_OFFLINE_EVAL=1 disables inline validation and checkpoints every step from grpo.val_start_at (plus the final step), weights only. - Each checkpoint's training_info.json records its weight-update end time (training_step_end_time_ms). - After training stops, a second driver in the same allocation/Ray cluster restores each checkpoint without optimizer state, refits vLLM, and runs the configured validation in step order, stopping at the first checkpoint that reaches the target accuracy. - run_stop is emitted once, backdated to the passing checkpoint's weight-update timestamp (aborted at the final checkpoint's timestamp when no checkpoint crosses); operational failures fail the job. Also fixes the dropped final-step train metrics: the trainers log a step's validation before its train metrics, and a target-reaching eval emitted run_stop first, suppressing the final step's tracked_stats. Evals started while a train block is open are now held and flushed after the step's train metrics land. Bumps mlperf-logging to the qwen35-ruleset pin and the default ruleset to 6.1.0.
Implements the offline-evaluation rules for this benchmark from mlcommons/training_policies#596: - DEFERRED_OFFLINE_EVAL=1 disables inline validation and checkpoints every step from grpo.val_start_at (plus the final step), weights only. - Each checkpoint's training_info.json records its weight-update end time (training_step_end_time_ms). - After training stops, a second driver in the same allocation/Ray cluster restores each checkpoint without optimizer state, refits vLLM, and runs the configured validation in step order, stopping at the first checkpoint that reaches the target accuracy. - run_stop is emitted once, backdated to the passing checkpoint's weight-update timestamp (aborted at the final checkpoint's timestamp when no checkpoint crosses); operational failures fail the job. Also fixes the dropped final-step train metrics: the trainers log a step's validation before its train metrics, and a target-reaching eval emitted run_stop first, suppressing the final step's tracked_stats. Evals started while a train block is open are now held and flushed after the step's train metrics land. Bumps mlperf-logging to the qwen35-ruleset pin and the default ruleset to 6.1.0.
| * qwen35_397b_grpo (Qwen3.5-397B-A17B) | ||
| ** Evaluation is performed offline based on checkpoints. Offline validation requires creating a sequence of checkpoints near the end of training that can be evaluated post-run. Beginning at the `<validation start step>` computed using the benchmark formula `CEIL(2.5 + 3840 / global_batch_size)`, the implementation must save a checkpoint after every individual step until training stops. Each checkpoint must store the timestamp of its latest weight update. Submissions must save at least one checkpoint and have the flexibility to select when to halt training, though stopping at `<validation start step> + 1` is recommended. To maintain consistent system performance and avoid artificial tapering of the workload, training samples must continue to be generated and processed at full capacity right up to the final stopping step. | ||
|
|
||
| ** Once training halts, saved checkpoints may be evaluated offline in any order until the one that achieves target evaluation accuracy is found. Checkpoint weights must be evaluated using either BF16 or the exact quantization format used during previous rollout generation. To finalize the run, emit an MLLOGGER run_stop event using the weight update timestamp of the passing checkpoint. Because the final score relies on that weight update timestamp, the time spent writing the final checkpoint itself to disk is excluded from the score. |
There was a problem hiding this comment.
Since we run training and evaluation separately, should we maintain two separate mllog files for each, or continue using the same one as https://github.com/mlcommons/training/blob/master/llm_post_training/rcp_logs/GBS512/seed_1.out?
There was a problem hiding this comment.
If we meet the accuracy target on step k, we used the weight update timestamp of the passing checkpoint, ignoring the checkpoint saving time for step k, while the checkpoint saving time for step k-1 is accounted for. Is that correct?
There was a problem hiding this comment.
If we meet the accuracy target on step k, we used the weight update timestamp of the passing checkpoint, ignoring the checkpoint saving time for step k, while the checkpoint saving time for step k-1 is accounted for. Is that correct?
Yes, that's the correct interpretation (unless storing the checkpoint can be done asynchronously, thus hiding its cost).
There was a problem hiding this comment.
Since we run training and evaluation separately, should we maintain two separate mllog files for each, or continue using the same one as https://github.com/mlcommons/training/blob/master/llm_post_training/rcp_logs/GBS512/seed_1.out?
It should be in a single mllog file for scoring. run_start is emitted from the training part, and run_stop from the evaluation part.
See here: mlcommons/training#906 (comment)
|
|
||
| ** Once training halts, saved checkpoints may be evaluated offline in any order until the one that achieves target evaluation accuracy is found. Checkpoint weights must be evaluated using either BF16 or the exact quantization format used during previous rollout generation. To finalize the run, emit an MLLOGGER run_stop event using the weight update timestamp of the passing checkpoint. Because the final score relies on that weight update timestamp, the time spent writing the final checkpoint itself to disk is excluded from the score. | ||
|
|
||
| ** For example, if `<validation start step>` is step 18 and submitter continues training through step 20 while maintaining full hardware load, the submission must write distinct checkpoints with update timestamps for steps 18, 19 and 20. During offline evaluation, if step 19 is evaluated first and it meets the accuracy target, the run_stop event using step 19's weight update timestamp is used to compute the score. Instead, if step 18 meets the accuracy target, the run_stop event using step 18's weight update timestamp is used and no checkpoint saving time in included in the final score. |
There was a problem hiding this comment.
The RCL_Logging will not be changed with the offline eval. We will have block_start -> post-training -> block_stop -> eval_start -> tracked_stats -> eval_accuracy -> eval_stop -> run_stop. The timestamp of run_stop should be the corresponding saved weight update timestamp in the checkpoint. Am I right?
How should we set the timestamp of eval_start, tracked_stats, eval_accuracy and eval_stop in the eval process? If directly from the computer time, the timestamp of eval will be larger than the timestamp of run_stop.
There was a problem hiding this comment.
For these events you should emit the computer time: "block_start -> post-training -> block_stop -> eval_start -> tracked_stats -> eval_accuracy -> eval_stop".
The only event that you need to adjust the timestamp for is run_stop. You may emit that one not with the computer time, but with the timestamp of the training step (right after the optimizer.step() is finished) that passes the threshold. This timestamp may be stored together with the checkpoint. This is the code in the updated reference that emits the run_stop with modified timestamp:
https://github.com/NVIDIA-NeMo/RL/blob/f987c0596af2bd6cfb9958e768deb8df951d7504/nemo_rl/algorithms/mlperf_grpo_deferred.py#L372-L381
mllogger.end(
key="run_stop",
metadata={
"status": "success" if passed else "aborted",
"samples_count": samples,
},
time_ms=int(info["training_step_end_time_ms"]),
)There was a problem hiding this comment.
The timestamp of run_stop should be the corresponding saved weight update timestamp in the checkpoint. Am I right?
Yes
If directly from the computer time, the timestamp of eval will be larger than the timestamp of run_stop.
Yes, that is fine and correct.
Summary
qwen35_397b_grpoand link them from the quality-measure table.CEIL(2.5 + 3840 / global_batch_size), require a checkpoint after every training step until training stops, while maintaining full-capacity sample generation and processing.run_stopusing the passing checkpoint's latest weight-update timestamp, excluding the time required to write that checkpoint, and add an illustrative example.Validation
git diff --check