Skip to content

Commit ab863b6

Browse files
author
Yakup Koray Budanaz
committed
X9 folds -clean into its arm instead of dropping the earlier wave: an owed rerun of 1-7 kernels erased the other ~40 kernels of the arm in every figure (2026-09-18 rule: pool, latest run per kernel wins)
1 parent 43af124 commit ab863b6

3 files changed

Lines changed: 46 additions & 143 deletions

File tree

‎docs/DESIGN_data_collection_and_scoring.md‎

Lines changed: 6 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -189,13 +189,11 @@ allowed under commit-single; more than one ACCEPTED submission is not.
189189
(`experiments.drop_cancelled_task_rows`, with a warning giving the count), the task row included:
190190
the job ended the agent mid-task (T6), so the rows report part of an episode and the token total
191191
prices part of one.
192-
- X9. An arm whose name ends in `-clean` is a re-run of one condition from scratch, launched after
193-
something about the earlier wave was found wrong. It carries the SAME identity, so within an
194-
identity group (`experiment`, `model`, `language`, `device`, `packet`, `harness`) a `task` row
195-
whose arm carries the suffix drops every row of every arm in that group WITHOUT it, at read
196-
(`experiments.drop_superseded_arm_rows`, with a warning giving the count and the number of arms).
197-
The clean tasks supersede the earlier ones rather than pooling with them; the suffix names no
198-
condition, and `arm` and `rep` are deliberately not in the group. The database is not modified (N1).
192+
- X9. An arm whose name ends in `-clean` (`CLEAN=1` waves and every owed rerun) is a re-run of the
193+
arm without the suffix. At read the suffix comes off (`experiments.fold_clean_arms`) and nothing is
194+
dropped: both waves' rows pool, and the latest run per kernel (R3) picks between them (2026-09-18
195+
user rule). An owed rerun of a few kernels therefore replaces only those kernels. The database is
196+
not modified (N1).
199197

200198
## 3. Per-task answer
201199

@@ -444,7 +442,7 @@ family; the table states the pairs it kept.
444442
| X6 | `experiments.drop_foreign_kernel_rows`, called by `experiments.read_observations` | `test_experiments.py`: foreign-kernel rows dropped with a warning, runs without a task row kept |
445443
| X7 | `experiments.drop_pre_relaunch_rows`, called by `experiments.read_observations` | `test_experiments.py`: pre-final judge rows dropped with a warning, a task with no stamp untouched, task start over the kept rows |
446444
| X8 | `experiments.drop_cancelled_task_rows`, called by `experiments.read_observations` | `test_experiments.py`: every row of a cancelled task dropped with a warning, a frame without the column untouched |
447-
| X9 | `experiments.drop_superseded_arm_rows`, called by `experiments.read_observations`; the `-clean` suffix is written by `CLEAN=1` in `experiments/submit-cpf-llr40.sh` | `test_experiments.py`: superseded arms dropped with a warning, another identity group untouched, a frame with no clean arm untouched |
445+
| X9 | `experiments.fold_clean_arms`, called by `experiments.read_observations`; the `-clean` suffix is written by `CLEAN=1` and by `submit-owed-wave.sh` | `test_experiments.py`: clean arm folded with every row kept, a 1-kernel rerun keeps the other kernels and wins its own under `latest_runs`, a blank arm stays blank |
448446
| R1, R2 | `population.graded_episode_rows`, `last_per_episode` | `test_aggregation_population.py`: last submission, non-positive, suspect |
449447
| R3, R4 | `population.latest_runs`, `arm_kernel_answers`, `kernel_tokens` | rerun supersedes; rerun without answer; undated; start-time tie order |
450448
| R5 | `population.arm_kernel_answers`, `kernel_tokens(repeats="median")` | median run and carrier; token median |

‎hpcagent_bench/experiments.py‎

Lines changed: 12 additions & 60 deletions
Original file line numberDiff line numberDiff line change
@@ -308,7 +308,7 @@ def read_observations(path: pathlib.Path) -> "pd.DataFrame":
308308
drop_foreign_kernel_rows,
309309
drop_pre_relaunch_rows,
310310
drop_cancelled_task_rows,
311-
drop_superseded_arm_rows,
311+
fold_clean_arms,
312312
):
313313
frame = rule(frame)
314314
return frame
@@ -436,69 +436,21 @@ def fold_renamed_arms(frame: "pd.DataFrame") -> "pd.DataFrame":
436436
return frame.assign(arm=frame["arm"].astype(str).map(renamed_arm))
437437

438438

439-
#: What a launcher appends to re-run an arm from scratch (``CLEAN=1``). It names no condition: the
440-
#: identity columns are unchanged, and the clean tasks SUPERSEDE the ones before them.
439+
#: What a launcher appends to re-run an arm (``CLEAN=1``, and every owed rerun). It names no
440+
#: condition: the identity columns are unchanged.
441441
CLEAN_SUFFIX: str = "-clean"
442442

443-
#: The identity a clean re-run supersedes within. Not ``arm`` -- the whole point is that the clean arm
444-
#: and the arm it replaces are two names for one condition -- and not ``rep``, since a designed repeat
445-
#: is a task of the same condition and is re-run with it.
446-
CLEAN_GROUP: tuple[str, ...] = ("experiment", "model", "language", "device", "packet", "harness")
447443

448-
449-
def clean_identity(frame: "pd.DataFrame") -> "pd.Series":
450-
"""One identity label per row for :func:`drop_superseded_arm_rows`.
451-
452-
THE RECORDED COLUMNS ARE NOT ENOUGH ON THEIR OWN. An extracted observations table carries
453-
``language``, ``packet`` and ``harness`` and no ``experiment``, ``model`` or ``device``, so a
454-
group built from :data:`CLEAN_GROUP` alone puts every model's C control in one identity, and one
455-
model's clean re-run then drops every other model's arm. That is what it did: six finished
456-
GPT-OSS-120B arms took 18 arms with them, Qwen3.8-27B's and Kimi-K2.7-Code's included.
457-
458-
The arm NAME carries what the columns do not, so the label is the recorded columns plus the arm
459-
with its ``-clean`` suffix removed. Two arms are then one identity exactly when they are the same
460-
name re-run, which is what the suffix means. A name that differs by more than the suffix
461-
supersedes nothing, and the failure direction is to keep both waves rather than to delete one.
462-
"""
463-
columns = [column for column in CLEAN_GROUP if column in frame.columns]
464-
# A missing cell reads as the empty string: a row with no arm or no recorded packet still needs
465-
# one label, and pandas keeps NA through astype(str) and through the string accessors.
466-
condition = frame["arm"].astype(str).fillna("").str.removesuffix(CLEAN_SUFFIX).fillna("")
467-
labelled = frame.assign(clean_condition=condition)
468-
return labelled[[*columns, "clean_condition"]].astype(str).fillna("").agg("\x1f".join, axis=1)
469-
470-
471-
def drop_superseded_arm_rows(frame: "pd.DataFrame") -> "pd.DataFrame":
472-
"""``frame`` without the rows of arms a ``-clean`` re-run superseded (spec X9).
473-
474-
A clean arm re-runs one condition from an empty workspace after something about the earlier wave
475-
was found wrong -- a leaked view, a broken relaunch, a job that requeued onto its own rows. It
476-
carries the SAME identity, so without this rule the two waves pool and the defect the re-run
477-
exists to escape is averaged back in. Within one identity group (:func:`clean_identity`), a task
478-
row whose arm carries the suffix therefore drops every row of every arm in that group without it.
479-
The frame changes, never the database (N1), and the count is warned about.
480-
"""
481-
import warnings
482-
483-
if frame.empty or "arm" not in frame.columns or "record" not in frame.columns:
484-
return frame
485-
groups = clean_identity(frame)
486-
clean = frame["arm"].astype(str).str.endswith(CLEAN_SUFFIX)
487-
superseding = set(groups[clean & (frame["record"] == "task")])
488-
if not superseding:
444+
def fold_clean_arms(frame: "pd.DataFrame") -> "pd.DataFrame":
445+
"""``frame`` with every ``-clean`` arm under the arm it re-ran (spec X9). Nothing is dropped:
446+
the waves pool and the latest run per kernel (``population.latest_runs``) picks between them
447+
(2026-09-18 user rule). Dropping every earlier row on any clean task row cost whole arms: an
448+
owed rerun of 1-7 kernels erased the ~40 kernels of the wave it topped up."""
449+
if frame.empty or "arm" not in frame.columns:
489450
return frame
490-
dropped = groups.isin(superseding) & ~clean
491-
count = int(dropped.sum())
492-
if count:
493-
arms = sorted(set(frame.loc[dropped, "arm"].astype(str)))
494-
message = f"dropped {count} row(s) of {len(arms)} arm(s) superseded by a clean re-run (spec X9)"
495-
warnings.warn(message, stacklevel=2)
496-
kept = frame[~dropped]
497-
# The suffix names a WAVE, not a condition, so it comes off once the wave it replaces is gone.
498-
# Left on, it renames the arm for everything downstream: a pair list, an --arms regex and a
499-
# figure's arm pattern all ask for the condition by name and would find nothing.
500-
renamed = kept["arm"].astype(str).str.removesuffix(CLEAN_SUFFIX)
501-
return kept.assign(arm=renamed.where(renamed.notna(), kept["arm"]))
451+
arms = frame["arm"]
452+
folded = arms.astype(str).str.removesuffix(CLEAN_SUFFIX)
453+
return frame.assign(arm=folded.where(arms.notna(), arms))
502454

503455

504456
def main(argv: list[str] | None = None) -> int:

‎tests/test_experiments.py‎

Lines changed: 28 additions & 75 deletions
Original file line numberDiff line numberDiff line change
@@ -16,6 +16,7 @@
1616
import pytest
1717

1818
from hpcagent_bench import experiments
19+
from hpcagent_bench.stats import population
1920

2021

2122
def test_a_blank_and_filled_arm_reads_as_one_identity() -> None:
@@ -178,37 +179,39 @@ def clean_frame() -> pd.DataFrame:
178179
)
179180

180181

181-
def test_an_arm_superseded_by_a_clean_rerun_is_dropped_with_a_warning() -> None:
182-
"""Spec X9: the clean arm re-ran the condition from scratch because the earlier wave was wrong,
183-
and the two carry one identity -- pooled, the defect the re-run exists to escape is averaged
184-
back in. What survives is reported under the CONDITION's name: the suffix named the wave."""
185-
with pytest.warns(UserWarning, match="dropped 2 row"):
186-
kept = experiments.drop_superseded_arm_rows(clean_frame())
187-
assert set(kept.arm) == {"cpf-llr-focus40-qwen38-c-cpf", "cpf-llr-focus40-qwen38-c-cpfsrc"}
188-
assert len(kept[kept.arm == "cpf-llr-focus40-qwen38-c-cpf"]) == 2
182+
def test_a_campaign_with_no_clean_arm_is_left_alone() -> None:
183+
frame = clean_frame()
184+
frame = frame[~frame.arm.str.endswith("-clean")]
185+
assert experiments.fold_clean_arms(frame).equals(frame)
189186

190187

191-
def test_a_clean_rerun_supersedes_only_its_own_identity_group() -> None:
192-
"""The suffix says which TASKS are live for one condition, not that every other arm of the
193-
campaign was re-run; a cpfsrc arm with no clean wave keeps every row."""
194-
with pytest.warns(UserWarning, match="spec X9"):
195-
kept = experiments.drop_superseded_arm_rows(clean_frame())
196-
assert (kept.packet == "cpfsrc").sum() == 2
188+
def test_a_clean_rerun_folds_into_the_arm_it_re_ran_and_keeps_every_row() -> None:
189+
"""Spec X9 (2026-09-18 user rule): the suffix names a wave, not a condition, so the clean arm is
190+
reported under the arm it re-ran and both waves' rows stay for the latest run to choose from."""
191+
kept = experiments.fold_clean_arms(clean_frame())
192+
assert len(kept) == len(clean_frame())
193+
assert (kept.arm == "cpf-llr-focus40-qwen38-c-cpf").sum() == 4
194+
assert (kept.arm == "cpf-llr-focus40-qwen38-c-cpfsrc").sum() == 2
197195

198196

199-
def test_a_campaign_with_no_clean_arm_is_left_alone() -> None:
200-
"""Every wave so far ran without the suffix, and X9 must be invisible to them."""
201-
frame = clean_frame()
202-
frame = frame[~frame.arm.str.endswith("-clean")]
203-
assert len(experiments.drop_superseded_arm_rows(frame)) == len(frame)
197+
def test_a_one_kernel_owed_rerun_keeps_the_arms_other_kernels_and_wins_its_own() -> None:
198+
"""The bug: an owed rerun is named -clean and covers a few kernels, and X9 used to drop every row
199+
of the wave it topped up -- 40 kernels became the rerun's one."""
200+
common = {"record": "task", "run_root": "r", "language": "c", "packet": "", "harness": "claude"}
201+
rows = [
202+
{**common, "arm": "x-qwen38-c", "job": "1", "run_id": f"a{i}", "benchmark": f"k{i}", "ts_ms": 1}
203+
for i in range(3)
204+
] + [{**common, "arm": "x-qwen38-c-clean", "job": "2", "run_id": "b0", "benchmark": "k0", "ts_ms": 2}]
205+
latest = population.latest_runs(experiments.fold_clean_arms(pd.DataFrame(rows)))
206+
assert sorted(latest.benchmark) == ["k0", "k1", "k2"]
207+
assert latest.set_index("benchmark").loc["k0", "job"] == "2"
208+
assert set(latest.arm) == {"x-qwen38-c"}
204209

205210

206-
def test_a_clean_arm_with_no_task_row_supersedes_nothing() -> None:
207-
"""A judge row alone does not say a re-run happened: the task row is what records that an agent
208-
was launched under the clean arm, the same evidence X6-X8 read."""
209-
frame = clean_frame()
210-
frame = frame[(frame.record != "task") | (~frame.arm.str.endswith("-clean"))]
211-
assert len(experiments.drop_superseded_arm_rows(frame)) == len(frame)
211+
def test_a_blank_arm_stays_blank_through_the_fold() -> None:
212+
frame = pd.DataFrame({"arm": [math.nan, "x-c-clean"], "record": ["call", "task"]})
213+
kept = experiments.fold_clean_arms(frame)
214+
assert experiments.is_blank(kept.arm.iloc[0]) and kept.arm.iloc[1] == "x-c"
212215

213216

214217
@pytest.mark.parametrize(
@@ -287,56 +290,6 @@ def test_nan_is_blank_but_zero_is_not() -> None:
287290
assert not experiments.is_blank("0")
288291

289292

290-
def test_a_clean_rerun_of_one_model_does_not_supersede_another_model(tmp_path: pathlib.Path) -> None:
291-
"""Spec X9 groups by identity, and an extracted table records language, packet and harness but
292-
not the model. Grouping on those alone put every model's C control in one identity, so six
293-
finished GPT-OSS-120B re-runs deleted Qwen3.8-27B's and Kimi-K2.7-Code's arms as well."""
294-
rows = [
295-
{
296-
"arm": "cpf-llr-focus40-oss120b-c-clean",
297-
"record": "task",
298-
"language": "c",
299-
"packet": "",
300-
"harness": "claude",
301-
},
302-
{"arm": "cpf-llr-focus40-oss120b-c", "record": "task", "language": "c", "packet": "", "harness": "claude"},
303-
{"arm": "cpf-llr-focus40-qwen38-c", "record": "task", "language": "c", "packet": "", "harness": "claude"},
304-
{"arm": "cpf-llr-focus40-kimi27sglang-c", "record": "task", "language": "c", "packet": "", "harness": "claude"},
305-
]
306-
with pytest.warns(UserWarning, match="superseded by a clean re-run"):
307-
kept = experiments.drop_superseded_arm_rows(pd.DataFrame(rows))
308-
assert sorted(kept.arm) == [
309-
"cpf-llr-focus40-kimi27sglang-c",
310-
"cpf-llr-focus40-oss120b-c",
311-
"cpf-llr-focus40-qwen38-c",
312-
]
313-
314-
315-
def test_a_clean_rerun_supersedes_the_arm_of_its_own_name(tmp_path: pathlib.Path) -> None:
316-
"""The suffix names no condition, so the clean wave replaces the wave it re-ran and nothing else."""
317-
rows = [
318-
{"arm": "x-qwen38-c-skills-clean", "record": "task", "language": "c", "packet": "lang-skills", "harness": "h"},
319-
{"arm": "x-qwen38-c-skills", "record": "task", "language": "c", "packet": "lang-skills", "harness": "h"},
320-
{"arm": "x-qwen38-c", "record": "task", "language": "c", "packet": "", "harness": "h"},
321-
]
322-
with pytest.warns(UserWarning, match="superseded by a clean re-run"):
323-
kept = experiments.drop_superseded_arm_rows(pd.DataFrame(rows))
324-
assert sorted(kept.arm) == ["x-qwen38-c", "x-qwen38-c-skills"]
325-
326-
327-
def test_a_surviving_clean_arm_is_reported_under_the_condition_it_re_ran() -> None:
328-
"""The suffix names a wave, not a condition. Left on the arm, it renames the condition for every
329-
consumer downstream: a pair list, an --arms regex and a figure's arm pattern all ask by name."""
330-
rows = [
331-
{"arm": "x-qwen38-c-clean", "record": "task", "language": "c", "packet": "", "harness": "h"},
332-
{"arm": "x-qwen38-c", "record": "task", "language": "c", "packet": "", "harness": "h"},
333-
{"arm": "x-oss120b-c", "record": "task", "language": "c", "packet": "", "harness": "h"},
334-
]
335-
with pytest.warns(UserWarning, match="superseded by a clean re-run"):
336-
kept = experiments.drop_superseded_arm_rows(pd.DataFrame(rows))
337-
assert sorted(kept.arm) == ["x-oss120b-c", "x-qwen38-c"]
338-
339-
340293
def test_a_column_no_row_in_the_table_ever_recorded_still_fills_from_the_arm_name() -> None:
341294
"""The bug this guards: a column NOTHING recorded reads back from CSV as all-NaN float64, and
342295
writing an arm's recovered language into that raised ``Invalid value 'c' for dtype 'float64'``

0 commit comments

Comments
 (0)