Skip to content

dlrmv4: replace the GBS 8192 RCP with a GB300 rerun of the reference - #478

Open
mmarcinkiewicz wants to merge 1 commit into
mlcommons:masterfrom
mmarcinkiewicz:michalm/dlrmv4-rcp-8192-gb300-reference
Open

mmarcinkiewicz wants to merge 1 commit into
mlcommons:masterfrom
mmarcinkiewicz:michalm/dlrmv4-rcp-8192-gb300-reference

Conversation

@mmarcinkiewicz

Copy link
Copy Markdown
Contributor

Follow up to mlcommons/training#908

GBS16k and GBS32k look alright. The proposed RCPs represent the true distribution very well

The AMD set for GBS 8192 came from three campaigns; the 3 h campaign's
stalled runs were replaced by new seeds, so the 20 runs carry one run above
80 M samples while 20-seed reruns of the reference code show five. The new
entry is 20 seeds chosen as ventile midpoints of a 179-run distribution and
run with the unmodified reference (mlcommons/training b5e8e2a) on 2x4 GB300,
uncapped, each run to the 0.75 AUC target.

Olympic mean 73.02 M (was 69.07 M), std 6.69 M (was 3.81 M), checker floor
68.16 M (was 66.30 M). Other batch sizes unchanged.
@mmarcinkiewicz
mmarcinkiewicz requested review from a team as code owners September 30, 2026 07:15
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@ShriyaRishab

Copy link
Copy Markdown
Contributor

WG meeting notes (10/1/26):

  1. AMD concern is that a set of submission runs like the previous RCPs fails the current RCPs?
  2. Changing very close to the submission is a concern from AMD
  3. Keeping new RCPs saves us from cherry-picking and unlucky/unfairness

WG decided to require N=12 runs (drop 2 slowest and 2 fastest) instead of N=10 to work around this. Apply this for all GBS. Need to apply logging and rules. @matthew-frank will make these PRs.

@mrmhodak

mrmhodak commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

@mmarcinkiewicz and @ShriyaRishab:

We have had an extensive internal discussion on this:

  1. Changing RCPs this close is disruptive, our optimizations have assumed the RCP convergence
  2. MI355X GPUs is the reference HW for this model and thus the RCP runs should come from MI355X and not from GB300.

Based on that, we are against changing the RCPs on a short notice.

However, the longer runs are a concern. Our offer to improve this is to do a 1/3 removal out of 12: Remove 1 shortest run and 3 longest.

@mmarcinkiewicz

Copy link
Copy Markdown
Contributor Author

Miro,

Changing RCPs this close is disruptive, our optimizations have assumed the RCP convergence

this is actually you shooting yourself in the foot - the submission is supposed to be run on random seeds, so using fixed, you would miss (as we have till now) that there are bad seeds till you actually see them in your submission - what would you do then? Just rerun with a different seed? This is something we'd like to avoid for everyone.

However, what we can do one time, and on a non-precedent setting basis, to allow you (and everyone else) to run with the same seeds the submission is using, so everyone converges more or less the same. What do you think about that?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants