Validate a reformatted dataset by producing and inspecting per-variable plots against a reference dataset. Use this after any backfill, after adding a new variable, or when investigating a regression.
The tool is deliberately visual – view the images. Statistics alone cannot catch many of the most common data quality problems (flipped maps, coordinate-axis errors, quantization banding, unit conversion bugs that only bite at certain lead times).
uv run src/scripts/validation/plots.py run-all <DATASET_URL>DATASET_URL is the complete, direct URL to the dataset (bucket-prefix/dataset-id/version). Examples:
s3://dynamical-noaa-hrrr/noaa-hrrr-analysis/v0.1.0.icechunks3://dynamical-dwd-icon-eu/dwd-icon-eu-forecast-5-day/v0.2.0.icechunk
To look up the URLs for a dataset, run uv run main <dataset-id> dataset-urls. It prints the primary and replica URLs; pass --format json for a machine-readable form.
Expect the run to take ~30–60 seconds per variable, mostly bounded by S3 reads (a ~20-variable dataset finishes in ~10–15 minutes). Progress is logged: one line per variable per plot type. Virtual stores are slower and their runtime scales differently — see Virtual and multi-level datasets below.
Long runs and remote sessions: a full virtual run-all takes on the order of an hour or two (see below), and upload of a many-variable run pushes hundreds of images. If you run these on a managed/remote session (e.g. Claude Code on the web), launch run-all as a background task and drive it with a monitor that emits a parsed progress line on a short cadence (≈4 minutes works well — frequent enough that the session is not reclaimed mid-run, and short enough to stay within the prompt-cache TTL so each wake-up stays cheap). The periodic emitted events are what keep the session warm — more reliable than a self-scheduled timer wake-up, which re-reads the whole session context on every fire and is blind to the run's actual state. A monitor that emits only at completion does not keep the session warm: the silent stretch until it finishes looks idle and the session is reclaimed mid-run, taking the ephemeral environment (and the run) with it. Make the monitor signal on both completion and failure — detect completion from the written validation_summary.md and failure from the run process having exited without it — so a crash surfaces immediately instead of looking like "still running". Launch the run detached (e.g. setsid) in its own process group so later cleanup can't cascade into the session's shell, watch the run's actual worker PID rather than the launcher's (and avoid a loose pgrep pattern that also matches the monitor's own command line and so never exits), and re-arm the monitor when it times out, since a single monitor is capped well under a multi-hour run. Also watch memory: a full virtual run-all holds several GB of decoded fields — budget ~6–8 GB RAM (measured peak ~6.5 GB on the 176-variable HRRR virtual store) and don't run it on an undersized host. On a large ensemble archive set MALLOC_ARENA_MAX=2: the scan runs hundreds of probe threads, and glibc's per-thread arenas hold freed memory rather than returning it, which roughly doubled resident size (9 GB against 4.9 GB for the same work on the 268-variable GEFS 35-day store).
On a large ensemble archive also set --probe-workers (see Options): the scan's default probe fan-out is tuned for archives with far fewer source files per region job, and on an ensemble archive it holds several GB more for no gain.
Pass --checkpoint-dir on any whole-archive scan you cannot afford to repeat. The manifest and decode scans are the long pole of a virtual run-all — hours on a large ensemble archive, where one sampled decode job alone takes tens of minutes — and without it an interruption anywhere loses all of it. With it each unit is recorded as it finishes (a manifest slice, a decode region job) and a re-run picks up from there.
When the run completes, stdout prints the path of validation_summary.md (relative to the repo root). Open that file first.
--reference-urlReference dataset for side-by-side comparison. Defaults to the NOAA GEFS analysis icechunk asset URL discovered from the STAC catalog athttps://stac.dynamical.org/catalog.json. Change if your dataset is outside the reference's temporal or spatial coverage. Variables not present in the reference still get validation-only plots and stats.--variable <name>(repeatable, alias-v) Restrict to specific variables. Especially useful while iterating on a single new variable — brings total runtime down to under a minute. To compare related variables side-by-side, pass--variablemultiple times:-v downward_short_wave_radiation_flux_surface -v total_cloud_cover_atmosphere.--start-date <YYYY-MM-DD>/--end-date <YYYY-MM-DD>Restrict the append dimension. Useful for reproducing an issue in a specific window.--init-time/--lead-time(forecast) or--time(analysis) Pin the append-dim position instead of drawing one at random.--init-time(forecast) and--time(analysis) each place both the spatial snapshot and the time series period, so one flag puts the whole report on a chosen date — which is how you keep a report's random draws off an era where part of the archive's variables don't exist yet.--lead-timenarrows the spatial snapshot alone. Also useful to reproduce a spatial anomaly deterministically.--output-dir <path>Write into an existing directory instead of a new one (useful if rerunning a single plot type into an existing run dir).--checkpoint-dir <path>Virtual stores: record the whole-archive scans there as they run and reuse anything already present, so an interrupted scan resumes instead of starting over — worth setting on any archive whose scans run for hours. The manifest scan writes one file per archive slice, the decode scan one per sampled region job. Files are keyed by dataset id and by everything that changes what a unit covers — scan window and variable filter, plus the region sample size and the resolved sampling configuration for the decode scan — so one directory holds both scans for several datasets, and changing a sampling parameter rescans instead of reusing results gathered under the old one. Files are not keyed by store URL, so give each store its own directory (a staging copy shares its dataset id). The scan window's end defaults to the newest committed position, so an operational update between an interruption and the resume moves it, which rescans the manifest scan's last slice and every decode job. A run that must survive updates pins--end-date(inclusive) to a date already fully committed and older than the region job'soperational_update_window, so nothing inside the window changes underneath a recorded result. Checkpoints record results, not an Icechunk snapshot: delete a dataset's files after a backfill or an update that changes references inside the scanned window, and after a change to the scans' own logic.--probe-workers <n>Virtual stores: how many region jobs the whole-archive manifest scan probes at once (default: the tuned value). The scan's in-flight jobs spread across that many append-dim manifest splits at a time, so on an archive with many source files per job the working set outgrows the ref cache and manifests are refetched instead of reused. Lowering it cuts peak memory, and is modestly faster. Measured on the 268-variable GEFS 16 day 0.5 degree virtual store (31 members x 105 leads x 2 file types = 6,510 source files per job), over one whole 360-job checkpoint window each with nothing else running: 8 workers took 20.1 minutes at a 3.6 GB peak, against 23.4 minutes and 7.2 GB at the 48-worker default. Memory is the reason to set it; the time difference is small. An archive with few files per job (the HRRR virtual store is ~100) needs no change.--level <value>For datasets with vertical-level groups, the level value to sample (e.g.500for 500 mb). Selects the nearest level on each variable's own vertical dim; the default is the middle level. Single-level variables ignore it.--point <lat,lon>(repeatable, up to 2) Pin the two spatial points used by the value and time series plots to specific coordinates instead of the random default; the nearest grid cell is used. Pick points that will provide clear validation signal, for example one point on land and another in the ocean for a global dataset, or two spaced-apart but on-land points for a regional dataset. Avoid points very near the poles and avoid points where there are expected geographic coverage gaps. The unpinned random default already draws both points from the middle 50% of each spatial dimension (avoiding the outer-25% margins), but pinning is the deterministic choice for a published report. Also honored by the standalonevalue-timeseriesandcompare-timeseriescommands so a verification re-render can reuse the same points.
You can also run a single plot type with compare-spatial, compare-timeseries, availability, or value-timeseries. The primary entry point run-all runs all of those steps and is the only command that produces the validation_summary.md index.
Two structural differences change how these tools behave; both are detected automatically from the store URL, but it helps to know what's happening.
A virtual dataset stores chunks as references into source GRIB files, decoding on read. One chunk is one GRIB message — a whole spatial field at a single (init, lead, member, level). So the cost model is the opposite of a materialized store: a spatial snapshot is cheap (one decode), while a point time series across the append dimension is catastrophic (one decode per position — millions for a full archive). run-all adapts:
- Availability is manifest-probed, not value-scanned. On a materialized store
availabilityreads.isnull()at two points across all time. On a virtual store that would decode every chunk it touches, and presence is anyway a manifest question (does the reference exist?), not a value question. Sorun-all(and the standaloneavailabilitycommand) runs the manifest scan instead — exhaustive per source file and per variable across the whole archive (seemanifest_scanbelow) — and renders the same availability artifacts either way. value-timeseriesis sampled. Instead of reducing mean ± std over lead × member at every position (which decodes the whole column), it strides the append dimension and decodes one message per sampled position, pinning a representative (lead, member, level) slice so the whole-period series is a single, comparable slice. A chunk is a full spatial field, so both run points are read from the same decode. This still surfaces drift and unit step-changes; per-variable loads run a few ahead of the plotting loop. The summary marks the series "sampled" and records the pinned lead/member/level.- Spatial and time series comparisons are unchanged — both already scope to a single snapshot / single init, the cheap access pattern.
Note that spatial downsampling (_downsample_for_plot) only shrinks the plot, not the read: gribberish decodes the full message regardless.
Virtual runs decode many full-field messages, so scope a run with --variable/-v or --start-date/--end-date when you don't need the whole archive. Each variable's decode count is logged before its read starts.
Runtime scales roughly linearly with variable count: budget ~0.4 minutes per variable end-to-end for a full virtual run-all (measured: 176 variables in ~71 minutes on the HRRR virtual store, including the concurrent whole-archive manifest scan and the sampled decode scan). So a ~150-variable virtual dataset takes ~1–2 hours, depending on archive length (the manifest scan cost grows with the number of expected source files) and S3 latency. Scope with -v/--start-date while iterating to bring this down to well under a minute.
For WeatherNext 2 publication restrictions, use the separate holdback audit and remediation runbook. Completeness checks missing expected refs; it cannot establish the absence of extra restricted refs.
These are the thorough, post-backfill correctness pass for a virtual dataset, distinct from the bounded operational checks run by the final update pod. Both reach the dataset's own region job (for expected source files and manifest chunk-key logic) by resolving the registered dataset from the store's dataset_id attribute, so both take a store URL like every other command. run-all runs both automatically for a virtual store, so one command produces the complete report — the manifest scan runs alongside the decode + plot phases in the same process, but they contend for I/O, so its cost is not fully hidden; the standalone commands are the strict post-backfill gates (exit code).
uv run src/scripts/validation/plots.py availability <DATASET_URL> [--start-date <date>] [--end-date <date>] [--min-fraction 1.0] [--checkpoint-dir <dir>]
uv run src/scripts/validation/plots.py decode-scan <DATASET_URL> [--start-date <date>] [--end-date <date>] [--max-samples 20] [--checkpoint-dir <dir>]
uv run src/scripts/validation/plots.py probe-position <DATASET_URL> <POSITION>- manifest scan (
availabilityon a virtual store) — the strict completeness gate, probing ref existence (no decode) job by job with two measures — the scan streams, folding each region job's probes as they complete, so only the accumulated per-position and per-variable result grows with window length and a whole multi-year archive can be scanned on a modest host. Per source file: every expected file's representative ref; the one-file-one-commit invariant (virtual_datasets.md) extends that probe to every ref the file contributed, so a file that never ingested (any lead / member / level-group file) is exactly the gap this catches. Any incomplete position is listed inunavailable_timestamps.txtto target a backfill. Per variable: each variable's ref at one present source file per position (smallest nonzero lead; any referenced level for group vars), which catches a variable missing from otherwise-ingested files — e.g. a variable the model only started producing partway through the archive shows as a 0→1 availability step. Neither measure establishes completeness of the levels within a variable: a variable is present when any level has a reference. The offline decode scan samples only referenced levels, so it also does not detect missing levels. Expected files respect the store's committedexpected_forecast_lengthcoordinate when the dataset has one, so leads that never existed upstream (e.g. a 36-hour-era init in a 48-hour archive) are not flagged. Exits non-zero if any position is below--min-fractionof expected source files. The scan window ends at the store's committed extent, so not-yet-published positions are never flagged; narrow further with--start-date/--end-date.
For a materialized store, variables with semantic NaNs use the availability of variables co-ingested from the same source files because their own missing values cannot distinguish unavailable data from a valid missing state.
- position probe (
probe-position) — the specific source file URLs a store has no reference for at one append-dim position.unavailable_timestamps.txtgives only present/expected counts, so this is how you get from "110/111 present" to the one file to check upstream when classifying a gap per 3e. - decode scan (
decode-scan) — sampled decode health. Samples evenly spaced append-dim regions (every variable group at each sampled region), decodes a bounded sample of present references — across lead times, members, and levels — and fails if any sampled chunk errors or decodes entirely NaN. Results render as the "Decode health (sampled)" section ofvalidation_summary.md(the standalone command also writesdecode_scan_summary.md). This is a sample, not an exhaustive sweep: a reference that decodes to garbage outside the sample is not caught (a literal every-chunk decode across all leads × members × levels is hours). With--checkpoint-direach sampled region job is recorded as it finishes, so an interrupted scan re-decodes only the jobs it never reached — worth setting whenever the scan's per-job cost (tens of minutes on a large ensemble archive) makes a restart from scratch expensive.
The guarantee split is the thing to keep straight: completeness is exhaustive at file granularity and at (variable, position) granularity, decode health is sampled. "Validated" therefore means every expected source file was ingested, every variable has a ref at every probed position, and a representative sample decodes cleanly — not that every value in the archive was decoded and checked. State this in the published summary so users read it correctly.
A dataset with vertical groups (e.g. pressure_level, model_level) exposes group variables by store path, like pressure_level/temperature, each carrying its level dimension. The plots sample one representative level per variable (the middle level by default), recorded in the summary and the plot title — the same philosophy as sampling one ensemble member and one spatial snapshot. Use --level <value> to inspect a specific level deterministically. Different groups have different level dims and lengths; each variable is sampled on its own dim. The manifest scan's file pass is the exception — it covers every source file (and, through the per-file atomic commit, every level those files contributed); its per-variable pass probes the middle vertical chunk first and expands to the remaining levels only when that reference is absent. A variable is present if any level has a reference. The offline decode scan probes level presence once per variable and source coordinate, then reuses that result to sample evenly among referenced levels, preserving the dataset validator's sampled_levels limit (default 3). Offline presence labels must match the selected variable's group dimension; a mismatch raises a validation error. A decode check fails if no sampled variable has a reference. Operational decode checks retain their configured level sampling and missing-reference checks. Manifest and decode checkpoint versions prevent reuse of results from scans that checked only the middle level.
Group variables are addressed by their store path: -v pressure_level/temperature, not -v temperature.
Each run writes to a fresh directory under data/output/:
data/output/<dataset-id>/<version>_<YYYY-MM-DDTHH-MM>/
├── validation_summary.md # start here
├── availability_heatmap.png # all variables × append dim, one image
├── availability_<var>.png # one per variable
├── value_timeseries_<var>.png # one per variable
├── spatial_<var>.png # one per variable
├── temporal_<var>.png # one per variable
└── unavailable_timestamps.txt # only if any data is missing
The directory path itself is dense enough to identify the run: dataset id + version + minute-precision timestamp. Multiple runs of the same dataset group under a shared <dataset-id>/ parent.
availability_heatmap.png— every variable × the full append dim, colored by fraction of data available (green = available, red = missing, grey = not probed because no source file is present at that position). Grey positions count as incomplete in the availability table because the variable is unavailable there even though no carrier file exists to probe. This is the availability overview: era boundaries (a variable that starts partway through the archive), whole-position gaps shared across variables, and per-variable holes are all visible in one image. Availability is manifest-probed on a virtual store (exhaustive) and value-scanned at the two run points on a materialized store.availability_<var>.png— the same series as one trace, one per variable. Reviewers may skip it for a variable whose summary line reads complete; it is always written so every variable's section renders consistently.value_timeseries_<var>.png— the variable's value over the entire dataset time range at two spatial points: the per-timestep mean (line, left y-axis). On a materialized forecast / ensemble dataset the per-timestep standard deviation across the non-time dims (lead time, ensemble member) is drawn as a second, lighter line on a secondary right y-axis, with amean/std devlegend. It is omitted (mean line only, no second axis or legend) when there is nothing to spread over: analysis datasets (a single value per timestep) and virtual stores (the series is a single pinned lead/member/level slice, so each timestep is one value). On a materialized store this reuses the point data already loaded for the availability scan; a virtual store does its own sampled read. Unliketemporal_<var>.png(a short window vs the reference), this spans all time and is meant to surface slow drifts and sudden discontinuities — e.g. a source that begins emitting a variable in different units shows up as a sharp step change in the mean line; where a std line is present, a distribution change shows up as a step change in it.spatial_<var>.png— 3 panels at one time step: reference map (left), validation map (middle), value distribution histogram (right). If the variable isn't in the reference dataset, the left panel reads "Variable not available" and the histogram plots only the validation distribution. Catches flipped / rotated / mis-projected maps, wrong coordinate extents, unit-scale mismatches (different histograms), and over-quantization (visible as banding).temporal_<var>.png— time series at two spatial points, validation (red) vs reference (blue). The window is drawn at random from the part of the append dim the reference also covers, falling back to the full validation range when the two archives don't overlap at all. If the variable isn't in the reference, only the validation series is plotted. Catches time misalignment, diurnal-cycle phase errors, unit mismatches, trend-level biases, missing/incorrect deaccumulation, and projection errors (uncorrelated / offset timeseries).
Plots are per-variable; availability_heatmap.png is the only all-variable image. Compare variables against each other by opening several *_<var>.png of the same type together.
The entry point for every run. It contains, in order:
- Validation and reference dataset identity (name, id, version, URL), time ranges, and scope.
- Run parameters: spatial points used by the value scans (all ensemble members), and for each of the spatial and time series sections the picked ensemble member plus the chosen init/lead/time (spatial) or timeseries period.
- Availability section: how availability was measured, the heatmap, a table of incomplete variables (complete/total positions, first/last incomplete), and a link to the unavailable-timestamps file (
unavailable_timestamps.txt). - Per-variable details: metadata (units, long/short/standard name, step type, and the variable's
commentattr when it has one), full-period value min/mean/std/max at each point, spatial + temporal min/max/mean for both validation and reference, and an availability line (positions complete, plus null counts at the two points on materialized stores).
Work in this order. AI assistants reviewing a dataset with more than ~25 variables must use the batched process in 3f — a complete-variables dataset's plots do not fit in one agent context — with 3a-3e as the per-batch and aggregation steps.
Open validation_summary.md first. It provides text-based information which can help identify issues that are harder to view in plots.
- Datasets block: confirm the validation dataset id + version match what you intend to review. Note the reference dataset and its time range — if the reference doesn't cover your validation window, temporal + spatial comparisons will be empty (expected, not a bug).
- Run parameters: confirm the ensemble member recorded under each of the spatial and time series subsections (if the dataset has ensembles — the same member is used for both, and the null analysis intentionally runs over all members), the chosen spatial time (init/lead for forecasts), and the timeseries period. The spatial plot is a single snapshot — if you see something weird, you can pass
--init-time/--lead-time/--timeto reproduce deterministically. - Availability: read the method line, the heatmap, and the incomplete-variables table. Any incomplete variable is your first lead. Open
unavailable_timestamps.txt— it lists the positions missing source files, to target a backfill. Use the first/last incomplete columns to spot patterns shared across variables (e.g. a shared first-incomplete date points to source coverage starting later). - Per-variable details: for each variable, compare validation stats to reference stats. Validation min/max that is orders of magnitude off from the reference is a near-certain unit mismatch.
Statistics miss the visual failure modes. Open the images and walk through the checklist in section 4.
Every per-variable PNG must be reviewed: for each variable, open value_timeseries_<var>.png, spatial_<var>.png, and temporal_<var>.png; open availability_<var>.png too unless the variable's summary line reads complete. Do not stop after a representative sample — issues can be variable-specific (a unit bug at one level, a flipped map for one field) and only surface when every plot is checked.
A good working rhythm:
- Open
availability_heatmap.pngfirst for the all-variable availability overview (era boundaries, shared gaps) — the only all-variable image. Per-variable comparison happens in step 2. - Open
value_timeseries_<var>.png,spatial_<var>.png,temporal_<var>.png(plusavailability_<var>.pngfor any incomplete variable) for every variable in the dataset — not a sample. Filenames are consistent, so a pattern like*_<var>.pngopens all of them at once in most viewers; opening a family together surfaces cross-variable patterns (e.g. radiation peaks coinciding with cloud cover minima). - Cross-check against the variable's row in
validation_summary.md(units, long_name, stats). - Apply the checklist below. Note anomalies as
<variable> + <file> + what's wrongso they can be acted on.
For any anomaly, reproduce it deterministically so a fix can be verified:
- If spatial: rerun
run-allwith--variable <name> --init-time <t> --lead-time <h>(forecast) or--time <t>(analysis). - If temporal: the timeseries period is randomized — it's in
validation_summary.md. Narrow with--start-date/--end-date. - If a discontinuity in
value_timeseries_<var>.png: narrow the full-period plot to the transition withvalue-timeseries <DATASET_URL> --variable <name> --start-date / --end-date, then fetch the source files on either side of the step to confirm a unit/scale change at the source. - If availability gaps: use the positions listed in
unavailable_timestamps.txtto backfill just those timestamps.
Before writing the summary, work through every item you flagged during the review and gather enough evidence to either confirm it as an issue, downgrade it to a documented quirk, or resolve it. The goal is that nothing reaches ### For further review without you having looked at it twice. Use the single-step entry points (compare-spatial, compare-timeseries, availability) rather than re-running the full run-all — they're faster and let you target the exact slice that's in question. Each follow-up category below has a required verification step — completing it is what lets you resolve the item or move it from ### For further review to ### Review notes.
If you can name a verification, you must run it. Any time you can write down a specific command, time range, lat/lon, or alternate snapshot that would confirm or rule out an item, run it yourself before stopping. The single-step tools are cheap (under a minute per re-render) and exist for this. Conclusions like "worth re-running on another snapshot", "should re-confirm against the raw GRIB", or "warrants checking at a 3-hourly slot" are unfinished verifications — they must be executed and the bullet updated with what the verification revealed, not left as a to-do for the next reader.
- Rerun a single plot type targeted at the variable, time, and location in question to confirm an anomaly is real and not an artifact of the randomly chosen snapshot or timeseries window. For example,
uv run src/scripts/validation/plots.py compare-spatial <DATASET_URL> --variable <name> --time <t> --output-dir <run-dir>to re-render one spatial plot at a chosen time, orcompare-timeserieswith--start-date/--end-dateand the same lat/lon as the original run to zoom in on a temporal anomaly. Re-rendering into the existing--output-diroverwrites the original PNG so the summary's links stay valid.- Gotcha — a
-v-filtered re-render shrinksavailability_heatmap.pngand desyncs the summary stats. A single-step command rebuildsavailability_heatmap.pngfrom only the variables passed in that invocation, and onlyrun-allwritesvalidation_summary.md— so a subset re-render into the run dir leaves the heatmap and the summary's stats stale. To fix a report whose random snapshot missed part-of-archive variables, re-run the fullrun-allwith the snapshot pinned (--init-time/--lead-timefor forecasts,--timefor analyses, or--start-date/--end-date) to a date where every variable is present. Reserve-v-filtered single-step runs for throwaway investigation in a scratch--output-dir, not for updating the report you intend to publish.
- Gotcha — a
- For unexpected unavailable timesteps, you must fetch a representative sample of the unavailable timestamps from the upstream archive before attributing the gap to an upstream cause — outages happen, but so do ingestion errors, and an LLM or a hurried reviewer will tend to default to "outage" without checking. A single spot-check is not enough: sample across the affected init cycles and lead times so a per-file bug (like a stale sidecar on individual leads) doesn't get mistaken for a clean outage. Verify both the source data file (e.g. the GRIB) and any sidecar the reformatter depends on (e.g. the
.idxbyte index for GRIB-based datasets). A key existing is not proof the data does: check the object's length and that its sidecar describes the position you expect. Upstream archives carry zero-byte placeholder objects (key present,Content-Length: 0, both grib2 and.idx), sidecars orphaned from a 404 object, and objects holding a different run's data entirely (a short object whose.idxnames another init or lead) — all of which a bare existence check reads as present. Useprobe-positionto name the files behind a count first, then check each one's length and sidecar rather than only its status code. Compare what you find against what the reformatted archive has at the same timestamp to determine whether to retry the backfill (the positions listed inunavailable_timestamps.txt) or document the gap as a confirmed source outage with the URL(s) you checked. - For suspected unit, scale, or coordinate bugs, cross-check against a third independent source (the raw GRIB/NetCDF file, a public viewer such as NOAA's nowCOAST or ECMWF's open charts, or another reformatted archive) — the GEFS reference is convenient but it's only one comparison and shares some biases with GFS-derived datasets.
- For ensemble datasets, rerun once more without
--variablefilters so a different ensemble member is selected; an anomaly that only appears for one member is structurally different from one that appears across members.
Track what you found per item so the eventual ### For further review entries cite the evidence (filename, timestamp, source URL) rather than just describing what you saw in the original snapshot.
When your review is complete, insert a ## Summary section into validation_summary.md, placed immediately below the Report generation start time: … line and directly above the ## Datasets section (i.e. right after the report's introductory paragraph and timestamp, not above them). This keeps the report's identifying header — the intro sentence and generation time — first, with your summary as the first content section that follows. This section is the part of the report that downstream readers actually scan — the rest is reference material — so its job is to answer "is this dataset ready to use, and is there anything I should know?" at a glance.
Brevity is part of the job. A summary's value degrades with length: a reader who skims three screens absorbs less than one who reads five tight bullets. Every sentence competes for the same limited attention, so include a detail only if a reader would do something differently for knowing it, and cut what merely proves you looked. No rule here replaces judgement about what a user of this dataset needs — the balance is to be complete on what matters and silent on what doesn't.
Opening sentences (always). Begin with one or two sentences stating where the dataset stands. In an early draft this might just be "Initial validation pass; see ### For further review for open items." In a final draft and in published reports it affirms the dataset has been reviewed and is ready for use, plus at most one brief sentence per ### Review notes item that changes what a user must do. Often there is nothing to add and the whole paragraph is "This dataset has been reviewed and is ready for use." — let readers reach the notes that concern them instead of previewing all of them.
Subsections (up to three, included as needed). Each is a bulleted list. Every bullet leads with a bolded phrase that carries the point on its own — a reader who reads only the bold text should still get it. Following prose is only what the bold phrase can't carry: the actionable specific, the affected window, the evidence link. For many bullets that is nothing, and the bold phrase is the whole bullet.
### For further review— issues an expert dataset creator should look into before the report is published, each with a link to the image(s) where the issue is apparent. Audience is the internal reviewer. Must be removed from final-draft and published reports — any item that ends up worth surfacing to users belongs under### Review notesinstead.### What looks good— brief confirmation of the checks in section 4 that passed. Good data is the expectation, not news; this section shows the review happened rather than cataloguing it. A handful of one-line bullets, and don't restate numbers the per-variable tables already carry.### Review notes— user-facing notes a downstream consumer can act on: true gotchas, not every source data nuance. Typical kinds are unavailable-data patterns, summarised rather than itemised (e.g. "Source data was unavailable for 16 hours on 2022-11-29 → 2022-11-30 across all variables; backfill is not possible from the upstream archive."), variables a user must explicitly mask, and source data changes over time (version-boundary behavior changes, historical low-quality windows, source outages). Group a finding that spans several variables into one bullet, and describe the state of the data rather than the story of how it was checked or a justification for a source model oddity. Items from### For further reviewthat are investigated and turn out to be real but acceptable should be moved here and reworded for an external audience. Omit this section entirely if there is nothing to report. Intrinsic, always-true variable facts (masking sentinels like -999/-50, unit quirks, what the variable is) do not go here; they belong in the variable'scommentattr so they travel with the data. See the comment-vs-review-note rule in AGENTS.md (Metadata conventions).
Cluster related notes under #### group headings. A reader scanning a dozen flat bullets has to work out for themselves which ones are the same kind of thing; a grouped list has that done for them. Group by what a reader would do about a note — the unavailable ranges, the values that are present but wrong — and give each group an #### heading naming the group as a noun phrase (#### Unavailable data, #### Known bad values), not a sentence. A single note that belongs to no group stays a plain top-level bullet; put those first, before the first heading, so the most important one leads the section.
Bullets inside a group follow the same rules as any other: bolded lead phrase, then only what the bold can't carry. Sub-bullets are for the other shape — one note whose items are the eras or windows it covers. There the range is the point, so lead the sub-bullet with it and keep the shared fact in the parent bullet:
- **The record is assembled from three source eras with different native resolutions and time steps.** Any field whose source grid is coarser than 0.25° is resampled onto it bilinearly, in every era.
- 2000-01-01 – 2019-12-31: 0.25°/0.5° retrospective data, 3-hourly at the source.
- 2020-01-01 – 2020-09-23T06:00: 1.0° data, 6-hourly at the source. …Don't group for its own sake — two unrelated notes forced under one heading is worse than two bullets, and a single note needs neither a heading nor sub-bullets.
Verification gate on ### Review notes. Any entry that attributes a gap or sentinel to an upstream cause (source outage, upstream archive gap, source-GRIB precision) must be backed by direct evidence: a fetched source file (per the unavailable-timestamp tactic in 3d) or an inspection of the reference dataset at the same timestamps. If you cannot produce that evidence, the item stays in ### For further review, not ### Review notes — "looks like an outage" without verification is exactly the failure mode this gate prevents. The evidence backs the claim but does not go in the note: the published entry states the fact, not how it was established.
Phrasing gate on ### For further review. Bullets in this section must describe an open question with the evidence already gathered, not a proposed verification that has not been performed. If a bullet contains "worth re-running", "should re-confirm", "warrants checking", "would be nice to verify", or any other phrasing that names a verification you could run, you have not finished §3d — go run it, then either resolve the item or rewrite the bullet around what the verification revealed.
A complete-variables dataset (e.g. noaa-hrrr-forecast-48-hour-virtual: 177 variables, so ~700 per-variable PNGs) cannot be reviewed in a single agent context: at roughly 300-500 tokens per plot image plus each variable's stats block and reasoning, a full pass is on the order of a million tokens before any follow-up work. At ~25 variables a single context works; well above that, split the review across subagents. The lead (orchestrating) agent must not open plot images itself — every image read happens inside a batch or verification subagent whose context is discarded after it returns.
Lead agent process:
- Produce or locate the run directory (
run-all— for a virtual store it runs the manifest scan and the sampled decode scan itself, so availability and decode health are already in the report). - From
validation_summary.mdread only the header: the Datasets block, Run parameters, and the Availability section (everything above## Per-variable details). Availability findings come from that section's table and files, not from plot review. Never read the per-variable details wholesale — at 177 variables that section alone is tens of thousands of tokens. - List the variables (
ls <run-dir>/spatial_*.pngmaps slugs back to names) and partition them into theme-coherent batches of 10-20: keep families together (all radiation, all reflectivity/precipitation, all cloud, each vertical group's variables together) so within-batch cross-variable checks — radiation peaking while cloud cover is high, temperature vs dew point consistency — remain possible. Note which batch got which family so cross-batch questions have an owner. - Spawn one subagent per batch, in parallel, with the prompt contract below.
- Aggregate the structured verdicts. Look specifically for cross-batch patterns the subagents cannot see: the same timestamp flagged in several batches (run-level ingest issue), one member of a physically-linked pair flagged while the other batch reported its partner clean, level-adjacent anomalies in a vertical group split across batches.
- Dispatch each SUSPECT/FAIL finding to a verification subagent per 3d: give it the exact single-step re-render command (scratch
--output-dir), or the upstream GRIB/idx URLs to fetch, and the specific question to answer. It returns confirmed/refuted plus evidence. The 3d rule is unchanged: if you can name a verification, run it — via a subagent. - Write the
## Summaryper 3e from the structured findings, citing the plot filenames and evidence the subagents returned.
Batch subagent prompt contract. Each batch subagent gets: the run directory path; its explicit variable list; the reproduction parameters from the run header it needs (spatial snapshot time, timeseries period, sampled level); any dataset-specific expectations (e.g. hour-0-NaN accumulated variables, projected y/x grid, known source quirks); and instructions to, per variable, read its PNGs and its heading section in validation_summary.md and apply the full section 4 checklist. Be explicit that the three per-variable plots — value_timeseries_, spatial_, and temporal_<slug>.png — are always opened for every variable (including availability-complete ones; the spatial map is the only place the flipped-map / sentinel-bleed / histogram-overlap checks can be made, and stats alone will not catch them), and that only availability_<slug>.png is conditional (opened when the variable's availability line is not complete). Slug = variable name with / → __. Word the contract so a subagent cannot read "conditional" as applying to the spatial/temporal plots — a common misread that silently drops the geometry checks for well-behaved fields. It must return compactly (≤ ~40 lines), e.g.:
VERDICTS
temperature_2m: OK
composite_reflectivity: SUSPECT — histogram mass at -10 dBZ sentinel-like spike
FINDINGS (SUSPECT/FAIL only)
- var: composite_reflectivity; plot: spatial_composite_reflectivity.png; what: ...;
where: init=..., lead=...; suspected-cause: ...; next-verification: <exact command>
BATCH OBSERVATIONS
- radiation family diurnal phase consistent across all 4 vars in this batch
Verdicts + precise pointers, not narration: the lead acts on next-verification lines and never re-reads the images behind an OK.
Sizing. Budget 10-15 variables per batch subagent — a 177-variable dataset is 12-18 parallel batches, then a handful of verification subagents. Keep batches under ~15 variables; a subagent that also does follow-up re-renders needs the headroom.
Look for each of these in every image.
- Map is oriented correctly. North at the top, east on the right. If the reference map shows a continent and the validation map shows its mirror image or upside-down copy, data and coordinates are out of sync (
pcolormeshis getting a flipped latitude or longitude axis relative to the data array). - Longitude axis is increasing with east on the right. If there's a vertical seam in the middle of the map, the data likely uses a 0–360 convention and needs conversion. If features are shifted 180° east/west, the longitude convention is wrong. We expect global datasets to go from −180 to +180 (not 0–360).
- Coastlines/land features land at expected coordinates. Locate a known landmark (e.g. the UK, Italy, the Great Lakes, southern Africa). Verify it sits at the correct lat/lon in the validation map. If it's shifted, the coordinate arrays are offset relative to the data array (off-by-one, wrong origin, or wrong grid spacing).
- Spatial extent matches the declared grid. Domain bounds in the plot should match what the dataset's
TemplateConfigdeclares. An extent truncation or overshoot means the coordinates or slicing is wrong.
- Validation histogram overlaps the reference histogram for variables that are also in the reference dataset. A large horizontal offset is a unit or scale bug. A large width difference is a quantization or smoothing bug.
- Validation min/max is physically plausible. Temperature in °C roughly −50 to +50. Pressure in Pa is ~50000–110000. Precipitation rate mm/s is tiny (10⁻⁵). Wind speed m/s is 0–100. Check
unitsinvalidation_summary.mdvs observed range. - No obvious quantization banding in the spatial map. Large flat patches of identical values or "staircasing" in smooth gradients indicates
keep_mantissa_bitsis too low. - No unexpected source sentinel values show through. A source nodata value may appear as a color extreme or histogram spike. When inspecting stored values, account for
keep_mantissa_bits: rounding can change the stored value and make it collide with legitimate values, so an absolute zero-residual scan may not prove the rewrite. - No invalid set is partly normalized. For each known invalid encoding, either its complete set becomes NaN or none of it does. If it remains unchanged, the variable's
commentmust explain the raw values. Do not infer partial normalization merely from the presence of unrelated NaNs. - Whole plot matches meteorological expectations. Look closely for subtly or obviously wrong new types of problems not enumerated here. Visual plots are a key layer of our defense in depth approach to catching data quality issues. We can't list every possible issue, rather use your meteorological knowledge to first define what you expect to see and compare that to what you actually see in the plots.
- Diurnal cycle is in phase with the reference. Shortwave radiation, 2m temperature, and similar variables should peak at the same local hour as the reference. Phase-shift is a time-coordinate bug (e.g. UTC vs local, or off-by-one lead time).
- Trend magnitudes match. The validation and reference should have similar day-to-day ranges. A consistent offset points to a calibration or unit issue; a consistent scale difference points to a unit conversion.
- No unexpected flatlines or spikes. Flatlines at one value for many steps can be a read error or a stuck sentinel; isolated spikes can be unit bugs at specific lead times (common in accumulated variables).
- Accumulated variables reset as expected. Precipitation and radiation accumulators should typically reset each forecast — check the
step_typeinvalidation_summary.mdand confirm the shape. - No obvious quantization in time series. Time series which are snapped or binned to a limited set of values or "staircasing" in what should be smooth time series indicates
keep_mantissa_bitsis too low. - Whole plot matches meteorological expectations. Look closely for subtly or obviously wrong new types of problems not enumerated here. Visual plots are a key layer of our defense in depth approach to catching data quality issues. We can't list every possible issue, rather use your meteorological knowledge to first define what you expect to see and compare that to what you actually see in the plots.
- No sharp discontinuities in the mean line. A sudden step up or down in the value across the full time range — not explained by a real seasonal/physical transition — is the signature of a source data change such as a unit switch (e.g. a GRIB file that begins emitting Kelvin instead of Celsius, or kg m⁻² instead of mm). These step changes are invisible to the short-window
temporal_<var>.pngplot, which is the reason this plot exists. - No step change in the std line (materialized forecast / ensemble datasets). The secondary-axis std line reflects the spread across lead time / ensemble members at each timestep. An abrupt rise or drop that persists indicates a distribution change in the source (e.g. a change in precision, smoothing, or member generation). Analysis datasets and virtual stores (whose series is a single pinned slice) have no std line, so only the mean-line check applies there.
- Whole plot matches expectations across time. Drifts, ramps, or level shifts that don't correspond to a known seasonal cycle or source transition warrant investigation — cross-check against the source archive at the timestamps where the change appears.
Availability (from availability_heatmap.png, availability_<var>.png, and unavailable_timestamps.txt)
- Every variable is fully available, or the gaps are explained. Any incomplete variable should have a reason: source data unavailable before a date for a specific variable, known source outage, ocean point for a land-only variable. Unexplained gaps are the bug.
- Unavailable pattern is not structural. Gaps concentrated at specific lead times, specific hours of day, or specific forecast cycles suggest a processing or indexing bug, not a random source outage. Use the heatmap and the first/last incomplete columns in the summary table to spot patterns shared across variables (e.g. a consistent first-incomplete date points to source coverage starting later; a variable available only after a date is typically a model-version addition — document it, don't retry it).
- Any availability gap you intend to label as an upstream outage has been verified per the sampling tactic in 3d. If every sampled file (and its sidecar index, if applicable) is present and intact, the gap is an ingestion bug, not an outage — keep it in
### For further reviewand target the positions listed inunavailable_timestamps.txt. - First step of an analysis dataset is NaN for accumulated variables. For analysis datasets,
step_type≠instantvariables (accumulation / average / max / min) are structurally NaN at the very first timestamp — there is no prior window to accumulate / average / extremize over. This is expected and not a bug.
Note on variables without hour-0 values: accumulation / average / max variables (and any variable the dataset marks with hour_0_values_override=False, e.g. instantaneous precipitation-type flags) are structurally NaN at each forecast's first lead time. So is a rate deaccumulated from a running total, even where its source publishes a lead-0 accumulation — differencing has no prior step there. Both availability paths exclude that slice: the value scan skips lead 0 for a variable whose stores_hour_0_values() is false (the store's own question), and the manifest probe, which probes source references, for one whose has_hour_0_values() is false. So "complete" for such a variable means no unexpected gaps.
Note on semantic-NaN materialized variables: a variable with internal_attrs.source_fill_value reads NaN both where data was never ingested and where the source's missing state legitimately applies (e.g. percent frozen precipitation wherever no precipitation is falling), so a value scan cannot measure its availability. The value scan instead reports the mean availability of the variable's source_file_var_groups co-members — "were this position's source files ingested". This is not a run-specific gate for a variable-filtered backfill unless those co-members were rewritten in the same run. Like the virtual manifest scan's per-file probe, this is file-granularity: such a variable individually missing from an otherwise-ingested file is not detectable. So for each of them, cross-check its value time-series plot against its availability line: a variable reporting complete whose time series is empty or drops out over a window means the co-ingested proxy missed a real gap — treat that as an availability gap and investigate.
- Reference availability. If the reference dataset doesn't cover your window,
validation_summary.mdwill showvariable not available in reference dataset— that is not a bug in the validation dataset, but spatial / temporal comparisons lose their signal. The timeseries window is already drawn from the overlap with the reference, so this means the two archives don't overlap at all: pick a different reference. - Ensemble member is plausible. For ensemble datasets
validation_summary.mdrecords the randomly-selected member. Rerun once more to confirm that a different member also looks right.
Once a run is reviewed, render it to a static HTML report, share it as one or more drafts for internal and external review, and finally publish the approved version. Two storage paths exist:
- Draft — every non-final upload goes here, timestamped, kept forever. Reachable only by its direct link and never surfaced to dataset users (the dynamical.org catalog links only the stable path), so uploading a draft does not make the report externally viewable — it's for sharing a run for review.
- Stable — the canonical, published report for a dataset, overwritten by each new publish. Embedded in dynamical.org and therefore seen by external users.
Both paths live in the dataset-validation-reports Cloudflare R2 bucket, served publicly at https://dataset-validation-reports.dynamical.org. Drafts and previously-published reports are archived forever — only the file at the stable path is overwritten.
uv run src/scripts/validation/plots.py render-report <run-dir>Reads <run-dir>/validation_summary.md and writes <run-dir>/validation_report.html next to it. Self-contained HTML (inline CSS/JS, no build step), images referenced at relative paths so the rendered file works locally and after upload.
The HTML mirrors the markdown 1:1 with two viewing affordances:
- Left TOC (slide-out hamburger on mobile) lists the top-level sections plus a Variables group with a checkbox per variable. "All / none" toggles let you compare a subset of variables side-by-side. Each variable section has
id="var-<name>"so URLs can deep-link (e.g.validation_report.html#var-temperature_2m). - Images are full-width on mobile and wrapped in a link to the underlying file so tap/click opens the full-resolution image.
upload re-renders before uploading, so this command is only needed for local-only previews.
uv run src/scripts/validation/plots.py upload <run-dir> # draft
uv run src/scripts/validation/plots.py upload <run-dir> --publish # publish to stable + archiveupload re-renders the HTML, then uploads the entire run directory (validation_summary.md, all *.png, unavailable_timestamps.txt if present, validation_report.html) to R2.
Without --publish:
<dataset-id>/drafts/<version>_<YYYY-MM-DDTHH-MM>/
Use this to share a single run for review without committing to it. Drafts are kept forever; iterate by re-running run-all (a new timestamped run dir) and re-uploading. Re-upload a fresh draft after every change to the validation summary so the shared link always reflects the current review state. Drafts go to timestamped paths that are never overwritten, so uploading a draft is non-destructive and does not require confirmation.
--publish is the opposite: it overwrites the stable, website-linked report, so never run it without explicit direction from the user — a draft is always the right default.
With --publish:
<dataset-id>/latest/ # stable, overwritten each publish
<dataset-id>/published/<version>_<YYYY-MM-DDTHH-MM>/ # archive copy, kept forever
The stable path is what the dynamical-stac catalog links to. The archive copy preserves what latest/ was before this publish; previously published reports are never lost.
Prints the public URL of validation_report.html on completion (e.g. https://dataset-validation-reports.dynamical.org/<dataset-id>/latest/validation_report.html).
The report moves through three audiences before it can be published. Each phase is just another upload (no --publish until the final step), but the ## Summary block is rewritten between phases for the next audience.
Phase 1 — Internal-review drafts. End each internal-review pass with upload <run-dir> (no --publish). Audience: internal data reviewers (you and other repo contributors). The summary's ### For further review section drives the conversation; internal jargon ("P1/P2", filenames, repo paths, ticket numbers) is fine here because every reader has the repo context. Iterate — investigate each item per 3d, update the summary, rerun run-all if new plots are needed, re-upload — until ### For further review is empty (every item is either resolved or has been moved to ### Review notes).
Phase 2 — External-audience draft. When ### For further review is empty, rewrite the ## Summary block for an external audience and upload one more draft (still no --publish). Audience: external dataset users who have never seen the run directory or our review process. Rewrite involves:
- Update the opening sentences to affirm the dataset is reviewed and ready for use, and to call out anything from
### Review notesthat needs special care. - Drop the (now empty)
### For further reviewsubsection. - Reword every remaining bullet for a public dataset consumer: spell out variable names, expand acronyms, avoid
P1/P2(use the explicit lat/lon or describe the point), avoid ticket numbers, internal codenames, file paths, and process shorthand. Each bullet must make sense in isolation to someone with no repo context.
Phase 3 — Publish. Only after a human reviewer approves the Phase 2 draft, run upload <run-dir> --publish to write to the stable path. Do not run --publish while ### For further review is non-empty or while the summary still reads as internal-audience prose — share another draft instead.
When the dataset already has a published report and you are only adding a variable (not creating the dataset), do not rebuild the summary from scratch:
- Run the full
run-all(all variables, not-v-filtered).--publishoverwrites the dataset's single website-linked report, so a variable-filtered run would drop every other variable from the catalog page. - Carry the already-approved commentary forward: read the current published
validation_summary.md(https://dataset-validation-reports.dynamical.org/<dataset-id>/latest/validation_summary.md) and reuse its approved## Summarytext for every finding that isn't new. Update or add only the pieces about the new variable, then run the draft → publish cycle from that merged summary. This keeps human-approved wording that the review cycle already vetted rather than regenerating (and re-litigating) it. - New-variable notes still follow the 3e split — time-windowed archive facts go in
### Review notes, intrinsic variable facts go in the variable'scomment. - If no published report exists yet (
latest/404s), there is nothing to carry forward — write a fresh full summary.
Once published, the report is picked up automatically on the next deploy of the dynamical.org site — the site's build pulls the latest published report for each dataset and incorporates it into the catalog page for that data product. No per-dataset wiring is required; just redeploy dynamical.org to refresh the catalog with the new report.
To skip the manual redeploy, set PAGES_DEPLOY_HOOK_URL (see Configuration) to a dynamical.org Cloudflare Pages deploy hook: a --publish upload then POSTs it after the files land, triggering one rebuild per publish. Draft uploads never trigger a deploy.
upload (both for uploading drafts and for publishing) reads R2 credentials from these environment variables:
R2_VALIDATION_REPORTS_ENDPOINT_URLR2_VALIDATION_REPORTS_ACCESS_KEY_IDR2_VALIDATION_REPORTS_SECRET_ACCESS_KEY
Scoped to the dataset-validation-reports bucket. Set them in the environment before running upload.
Optionally set PAGES_DEPLOY_HOOK_URL to a dynamical.org Cloudflare Pages deploy hook URL; when set, a --publish upload POSTs it after uploading to trigger a site rebuild (unset → publishing uploads only, and the site refreshes on its next deploy).