Fix data pipeline bugs and write GOOD 2.x format - #2
Open
headisbagent wants to merge 5 commits into
Open
headisbagent wants to merge 5 commits into
headisbagent wants to merge 5 commits into
Conversation
Fixes the data problems found in the September 2026 GOOD model review: - Wind profiles had 25 values per day: table_4-39_onshore.csv carries an index column, so iloc-based slicing started each day at "Day Of Month". Hour columns are now selected by name (load, solar and hydro too), which also keeps results unchanged under pandas 3. - Hydro profiles dropped Hour_0 and were never attached to plants (asset key "REGION:hydro:" vs profile key "REGION:hydro"). One profile_key() now names every profile, and hydro uses the monthly capacity factors as a daily energy budget. - New wind and solar capex was $/kW divided by 1e6; it is now $/MW (x1e3) with Table 4-16 fixed O&M and EPA's 9.78% capital charge rate. - Storage gets durations and efficiencies: existing batteries from EIA-860 2021 Schedule 3-4 (added to Raw/, fleet median 2 h otherwise), new batteries from EPA Table 4-35. - Lines carry EPA inter-regional losses and share corridors. Also fixes problems found in the pipeline itself: - Missouri's state code was "MO " (trailing space), so rps_MO matched no assets. - Offshore wind, biomass and landfill gas were not flagged renewable. - Existing wind and solar all took the first resource class's profile; they now use a capacity-weighted average of the region's classes. - Nuclear dispatch cost included fixed O&M and was rescaled to $21.2/MWh; it is now fuel plus variable O&M (mean $6.8/MWh). - Fill-in costs were sampled with an unseeded np.random; a seeded Generator makes builds reproducible. - ffill(inplace=True) on column selections becomes a no-op under pandas 3; replaced with assignment. Outputs use GOOD 2.x units (MW, MWh, $/MWh, $/MW, kg/MWh), classes (wind and solar are curtailable Producers, loads are positive MW) and declarative policy filters. Every assumed value is in PARAMETERS with its source. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build.py rebuilds Data/US/Processed and writes GOOD graphs to Outputs/ (the US, nine interconnect regions and an aggregated California example) in one command, validating each graph against GOOD's input schema. "python build.py --check" rebuilds into a temporary folder and fails if the result differs from the committed files. Removes Raw_Data.ipynb, Graph.ipynb, Aggregation.ipynb and process_data.py, which build.py replaces, and src/graph.py, src/inputs/aggregate.py, src/reload.py and src/progress_bar.py, which good.graph and good.aggregate replace. Outputs/ now also ignores gzipped graphs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Built with "python build.py" (seed 0). Existing capacity is unchanged (1,135,501 MW across the same 18,993 units). All 868 profiles have 8,760 values; wind profiles match those in PR #1. 22 zero-capacity wind cost classes are no longer written as candidate sites. Removes plants.json and combined_assets.json, which nothing produced or read, and adds metadata.json with units and sourced parameters. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tests cover the parsing fixes on synthetic tables (index columns, row order, hour columns, profile keys, state codes, cost units, seeded filling, line losses) and check the committed data: 8,760-value per-unit profiles, resolved profile references, state codes, MW and $/MW ranges, capacity against NEEDS, storage, policies, and GOOD validation of the full US graph. CI runs them and "python build.py --check". The README documents the build, outputs, units, sources and the assumptions that need review. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI's rebuild check failed on three weighted-average wind profiles: the weights were applied with a matrix product, which BLAS computes in a different order on Linux (OpenBLAS) and macOS (Accelerate), and seven values landed on the other side of a 5th-decimal rounding boundary. The average now uses elementwise operations, and "build.py --check" accepts numbers that differ only in their last written digit. Regenerated the three profiles (seven values change by 1e-5). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes the data problems found in the September 2026 GOOD model review, fixes several pipeline bugs, and makes the pipeline write GOOD 2.x format (ucdavis/good_model#3). The notebooks are replaced by one command,
python build.py, which validates every output graph with GOOD. CI checks that the committed data can be rebuilt fromData/US/Raw(numbers must match to their last written digit).Data fixes (from the model review)
table_4-39_onshore.csvhas an index column, solong_wide()'siloc[:, 6:]started each day at "Day Of Month". Hour columns are now selected by name everywhere. The resulting wind profiles are identical to those in Grant 2/20/2026: Fixed onshore wind data #1, which fixed the same problem by editing the CSV; this change works with either version of the file."REGION:hydro:"and profiles"REGION:hydro", andiloc[:, 3:]droppedHour_0. Oneprofile_key()now names every profile, and hydro uses its monthly capacity factors as a daily energy budget.$/kW / 1e6. It is now $/MW, with Table 4-16 fixed O&M and EPA's 9.78% capital charge rate. Range: $929–2,193/kW.Raw/; median 2.0 h, 63 of 102 matched). New batteries use EPA Table 4-35.Pipeline bugs found along the way
'Missouri': 'MO 'had a trailing space. 495 assets now matchrps_MO.np.random.choice. The build now uses a seeded generator.renewable.renewable_transmission_cost,ffill(inplace=True)on a column becomes a no-op, andgroupby().apply()offsets shift. The new code gives the same output on pandas 2.3 and 3.0.Structure
build.pyreplacesRaw_Data.ipynb,Graph.ipynb,Aggregation.ipynbandprocess_data.py. It writesData/US/Processed/and GOOD graphs toOutputs/: the US, nine interconnect regions, and an aggregated California example.good.graphandgood.aggregatereplacesrc/graph.pyand the old aggregator.plants.jsonandcombined_assets.jsonare removed; nothing produced or read them.metadata.jsonrecords units and every assumed parameter with its source.Unchanged
Existing capacity is identical (1,135,501 MW, the same 18,993 units). Coal, gas and oil mean dispatch costs are unchanged. 22 zero-MW wind cost classes are no longer written as candidate sites.
Needs team review
PARAMETERS:capacity_factor.csv(hydro) andrps_fraction.csv(its layout matches NREL ReEDS inputs). Please add them to the README.main. Keeping the cleaner CSV from Grant 2/20/2026: Fixed onshore wind data #1 is fine with this code.requirements.txtpinsgoodto thev2.0-review-fixesbranch of GOOD 2.0: fix review findings, move to linopy and MW units, add docs good_model#3. Switch it to a release tag once that PR merges.Test plan
pytest -q: 21 tests pass on pandas 2.3 and 3.0. They cover the parsing fixes on synthetic tables and checks on the committed data, including GOOD validation of the full US graph.python build.py --check: a fresh build matches the committed files on pandas 2.3 and 3.0.python build.py: all 11 graphs pass GOOD validation.🤖 Generated with Claude Code