Repository navigation
Update rcp_checker.py to handle llama31_8b epochs correctly - #434
gyulaz-htec wants to merge 1 commit into
Conversation
Right now the RCP check fails for llama3 using the reference implementation. It throws an error because it trries to get the epch from epoch_nums instead of sample_count:
```
File "/home/gyuzakor/actions-runner/_work/_tool/Python/3.12.11/x64/lib/python3.12/site-packages/mlperf_logging/rcp_checker/rcp_checker.py", line 496, in check_directory
bs, subm_epochs, benchmark = get_submission_epochs(result_files, version, bert_train_samples)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/gyuzakor/actions-runner/_work/_tool/Python/3.12.11/x64/lib/python3.12/site-packages/mlperf_logging/rcp_checker/rcp_checker.py", line 136, in get_submission_epochs
curr_not_converged, curr_subm_epochs, curr_bs, curr_benchmark = read_submission_file(
^^^^^^^^^^^^^^^^^^^^^
File "/home/gyuzakor/actions-runner/_work/_tool/Python/3.12.11/x64/lib/python3.12/site-packages/mlperf_logging/rcp_checker/rcp_checker.py", line 93, in read_submission_file
conv_epoch = json.loads(eval_accuracy_str)["metadata"]["epoch_num"]
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^
```
Processing llama31_8b logs the same way as llama31_405b logs should solve the problem.
|
MLCommons CLA bot: |
|
recheck |
|
@gyulaz-htec - can you please fix the CLA error so we can merge this? |
|
Closing as its a duplicate of #435 |
Right now the RCP check fails for llama3 using the reference implementation. It throws an error because it trries to get the epch from epoch_nums instead of sample_count:
Processing llama31_8b logs the same way as llama31_405b logs should solve the problem.