atlas-outcomes: a place-decision pack built from real outcomes (USGS flow, NWS flood warnings) - #62
Merged
Merged
Conversation
A stdlib dataset factory (apps/atlas-outcomes) that turns USGS daily flow and IEM VTEC flood-warning polygons into jev-decision-schema-1.0.0 records about a gauge at the end of day t (flow percentiles vs its own history, upstream gauge, Atlas construct stack and active layers, recent warnings), labelled by recorded futures: flow on t+1 above p90, a FL/FA/FF warning polygon over the gauge within 24 h, and the next-day change level. - run.py harvest | features | curate | baseline | train-command - 60 gauges, 25 NWS offices; committed gzipped snapshots under sources/, offline byte-identical rebuild (checked in CI) - time-split holdout (latest 365 days, one-day embargo); easy negatives downsampled in train only, with sample weights in provenance.jsonl - stdlib logistic baseline vs persistence and climatology (BASELINE.json) - provenance tier outcome-real added to the factory; datasets registered; ARCHITECTURE.md places the pack in the learning plane Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WH1hPpf3u1xWecEhS6AHD2
The block named 07a2146, a claude/decision-plane commit that the squash merge left off main, so check_handoff_fresh.py --check failed lint-and-test on main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WH1hPpf3u1xWecEhS6AHD2 (cherry picked from commit d9aaa9b)
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


This adds
apps/atlas-outcomes, a strictjev-decision-schema-1.0.0pack. Every label is something that actually happened later, not a label an LLM wrote. Its provenance tier is the newoutcome-real, from the amended data policy: it may train candidates and is evaluated on a time-split holdout. Nothing is ever promoted automatically.Data
sources/total 3.2 MB.run.py curaterebuilds the pack offline, byte for byte. A new CI job checks this.Questions in each record
high_next(yes/no): flow on day t+1 is above the gauge's historical 90th percentile.warn_next(yes/no): a flood warning covering the gauge is issued within 24 hours.flow_change(score 0–4): how much tomorrow's flow changes.Features use only data available at day t. Percentiles are computed from strictly earlier days, and tests check that there is no leakage.
Split
One day between the two is left out, so no training label looks into the holdout.
In training only, easy negatives are downsampled to 10%. Their weights are recorded in
provenance.jsonl.Baselines on the holdout (CPU, run here)
high_nextwarn_nextflow_changerun.py train-commandprints the System One training and serving commands for the GPU host.run.py baseline --score preds.jsonlscores a checkpoint with the same metrics.Other changes
factory/config.pyanddocs/FACTORY.md: the newoutcome-realtier.factory/datasets.json: the registered snapshots and pack.docs/ARCHITECTURE.md: where place decisions fit in the learning plane.ci.yml: a newatlas-outcomesjob running ruff, 16 unit tests and the byte-identical rebuild check.Caveats
🤖 Generated with Claude Code
https://claude.ai/code/session_01WH1hPpf3u1xWecEhS6AHD2
Generated by Claude Code