Skip to content

Export prediction artifacts in TabArena's result format - #42

Merged
adrian-prior merged 7 commits into
mainfrom
adrian/prediction-artifact-runner
Sep 23, 2026
Merged

adrian-prior merged 7 commits into
mainfrom
adrian/prediction-artifact-runner

Conversation

@adrian-prior

@adrian-prior adrian-prior commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

Generated by Codex

Export model validation and test predictions with --predictions-dir in TabArena's raw result format for offline evaluation. Add --refit-all-configs to produce test artifacts for successful configs beyond the selected/default configs. Keep extra fit/predict timings in artifact metadata, separate from benchmark result timings.

Behavior
  • Normal export mode saves existing tuning and selected/default test predictions without extra fits or predictions.
  • Separate benchmark and extra-refit loops preserve the benchmark timing policy; recorded extra timings cover fit/predict calls, excluding scoring and artifact I/O.
  • Completed artifacts contain validation/test labels and predictions. Configs without test predictions retain validation-only artifacts.
  • Test labels are retrieved for export after prediction. Models must preserve the supplied evaluation-table row order.
  • An explicitly requested all-config export fails if an extra refit or artifact write fails; exceptions propagate without a separate error summary.
  • Partial artifacts are replaced atomically, so interrupted validation exports can be retried. Existing completed artifacts remain protected from overwrite.
  • Native system artifacts and cloud transfer integration are outside this change.
Format and scope
  • Uses data/<model_config>/<dataset__task>/<repeat>_<fold>/results.pkl, with the seed identifying a repeat of temporal fold 0.
  • Labels are included in each config artifact. Loading requires trusted pickle files.
  • Validation-only outputs use validation.partial until test predictions are available, so pickle discovery does not treat them as completed results.
  • Preserves temporal holdout semantics and existing run fingerprints without hashing input tables again.
  • Checks expected row order and array shape; models must return predictions in input-table order. Plain prediction arrays cannot expose a model-side permutation.
  • Local artifact I/O only. Upload/download integration and resumable prediction caching are outside this change.
Testing
  • Rebased onto the workspace package refactor; shared contracts are imported from relarena_core.
  • All three package builds and required-package-data checks pass. The three isolated wheel-installation checks also passed during rebase validation.
  • Latest full suite: 490 passed, 10 skipped. All configured pre-commit hooks passed, including pinned Ruff 0.15.13.
  • The 15 prediction tests cover binary/regression round trips, normal/all-config modes, both final-fit regimes, reproduced metrics, row/class/split mismatches, label corruption, fit counts and timing separation, plus rejection of unsupported task types.
  • Added regression tests for propagated optional-refit/export errors, repeated validation exports, recovery from interrupted partial files, and preservation of existing files after failed writes.
  • Ran three cheap LightGBM configs (1, 2 and 3 estimators) on rel-f1/driver-dnf, seed 7, with all-config refits. Verified 566 validation rows, 702 test rows and one additional nonwinning refit.
  • TabArena's actual result reader at revision de1be372095e23e72e2f42992eceea036b97dea6 loaded all three LightGBM artifacts and 20 synthetic binary/regression artifacts directly. Recomputed validation/test metric errors and checked config metadata and simulation inputs.
  • The reader check is repeatable with workflows/verify_tabarena_predictions.py in an environment with TabArena installed.
  • No successful cloud upload/download round trip is claimed. These tests do not detect a model returning shuffled predictions without row IDs.

@adrian-prior
adrian-prior added this pull request to stack #43 September 11, 2026 15:37
@adrian-prior
adrian-prior force-pushed the adrian/prediction-artifact-runner branch from 3a6a8a9 to 71ab8de Compare September 16, 2026 12:35
@adrian-prior
adrian-prior removed this pull request from stack #43 September 16, 2026 12:40
@adrian-prior adrian-prior changed the title Export model predictions from benchmark runs Export prediction artifacts in TabArena's result format Sep 16, 2026
@adrian-prior
adrian-prior changed the base branch from adrian/prediction-artifacts to main September 16, 2026 12:40
@adrian-prior
adrian-prior marked this pull request as ready for review September 16, 2026 12:41

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/relarena/src/relarena/runner.py
Comment thread src/relarena/predictions.py Outdated
@adrian-prior
adrian-prior force-pushed the adrian/prediction-artifact-runner branch from 71ab8de to fc1da91 Compare September 16, 2026 13:43

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit fc1da91. Configure here.

Comment thread workflows/verify_tabarena_predictions.py

@Innixma Innixma left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Awesome!

@adrian-prior
adrian-prior merged commit c45538c into main Sep 23, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants