Skip to content

[MLPerf 6.1][GRPO] Define Qwen3.5 OpenHands agent harness contract - #599

Merged
ShriyaRishab merged 7 commits into
mlcommons:masterfrom
al-rigazzi:arigazzi/qwen35_openhands_harness_contract
Sep 30, 2026
Merged

ShriyaRishab merged 7 commits into
mlcommons:masterfrom
al-rigazzi:arigazzi/qwen35_openhands_harness_contract

Conversation

@al-rigazzi

@al-rigazzi al-rigazzi commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Define a benchmark-specific OpenHands CodeActAgent harness contract for qwen35_397b_grpo.
  • Allow a submitter-selected immutable OpenHands repository and commit, with narrow documented deviations, while preserving model inputs, tool behavior, turn accounting, termination semantics, and scoring.
  • Make security_risk optional for execute_bash and str_replace_editor, and require disclosure of the harness repository, commit, and all reference deviations.

Motivation

The current rule pins the agent harness to a single OpenHands repository and commit. Integrating the harness with different training, inference, serving, or environment frameworks may require narrow implementation changes for compatibility, reliability and failure handling, resource isolation, or security. An exact commit requirement makes behaviorally equivalent implementations noncompliant even when they preserve benchmark-relevant behavior. This change instead fixes the model-visible inputs, tool-call contract, action and result semantics, turn accounting, termination behavior, and scoring contract, while allowing a submitter-selected immutable implementation and requiring every deviation from the reference to be disclosed.

We propose making security_risk optional because it is auxiliary, model-generated metadata that is not required for tool execution or task scoring in this benchmark's noninteractive configuration. The reference OpenHands tool schemas declare security_risk required for execute_bash and str_replace_editor, but the reference runtime does not enforce its presence: it checks command for execute_bash, checks command and path for str_replace_editor, and sets security_risk only when supplied. In the benchmark's noninteractive configuration, confirmation_mode defaults to false, risk handling is conditional on confirmation mode, and task reward depends only on whether the task is resolved. Its surrounding syntax accounts for a substantial fraction of observed rollout–trainer logprob outliers, which can cause entire successful trajectories to be discarded. Making the annotation optional removes the requirement to generate those problematic contexts while preserving task semantics, tool capabilities, execution controls, and training filters. Preliminary multi-seed results support the mitigation.

Validation

  • Documentation-only policy change.
  • git diff --check origin/master...HEAD passes.

@al-rigazzi
al-rigazzi requested review from a team as code owners September 21, 2026 13:36
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

Resolve conflict in the benchmark-specific rules: both master (offline
evaluation rule) and this branch (OpenHands agent harness contract) add a
qwen35_397b_grpo entry at the same location. Keep both texts under a
single qwen35_397b_grpo heading.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
ShriyaRishab
ShriyaRishab previously approved these changes Sep 30, 2026
@ShriyaRishab
ShriyaRishab merged commit 14bc89f into mlcommons:master Sep 30, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants