Skip to content

atlas-outcomes: a place-decision pack built from real outcomes (USGS flow, NWS flood warnings) - #62

Merged
jcdavis131 merged 2 commits into
mainfrom
claude/atlas-outcomes
Sep 24, 2026
Merged

jcdavis131 merged 2 commits into
mainfrom
claude/atlas-outcomes

Conversation

@jcdavis131

Copy link
Copy Markdown
Owner

This adds apps/atlas-outcomes, a strict jev-decision-schema-1.0.0 pack. Every label is something that actually happened later, not a label an LLM wrote. Its provenance tier is the new outcome-real, from the amended data policy: it may train candidates and is evaluated on a time-split holdout. Nothing is ever promoted automatically.

Data

  • 60 USGS gauges across 19 states and 25 NWS offices. 25 of them have a paired upstream gauge on the same river.
  • 626,718 gauge-days of daily flow, from 1995 to 2026-09-22.
  • 37,383 NWS flood warning polygons from the IEM VTEC archive (river flood, areal flood and flash flood), 2015–2026.
  • A warning "covers" a gauge when the gauge's point is inside the warning polygon.
  • The committed gzipped snapshots in sources/ total 3.2 MB. run.py curate rebuilds the pack offline, byte for byte. A new CI job checks this.

Questions in each record

  • high_next (yes/no): flow on day t+1 is above the gauge's historical 90th percentile.
  • warn_next (yes/no): a flood warning covering the gauge is issued within 24 hours.
  • flow_change (score 0–4): how much tomorrow's flow changes.

Features use only data available at day t. Percentiles are computed from strictly earlier days, and tests check that there is no leakage.

Split

Rows Dates
Train 79,491 2015 to 2025-09-20
Holdout 20,921 2025-09-22 to 2026-09-21

One day between the two is left out, so no training label looks into the holdout.

In training only, easy negatives are downsampled to 10%. Their weights are recorded in provenance.jsonl.

Baselines on the holdout (CPU, run here)

Label Measure Climatology Persistence Logistic regression
high_next AUC / Brier 0.730 / 0.0498 0.874 / 0.0228 0.982 / 0.0157
warn_next AUC / Brier 0.697 / 0.0089 0.533 / 0.0167 0.831 / 0.0089
flow_change accuracy / MAE 0.435 / 0.819 0.500 / 0.783 0.465 / 0.743

run.py train-command prints the System One training and serving commands for the GPU host. run.py baseline --score preds.jsonl scores a checkpoint with the same metrics.

Other changes

Caveats

  • USGS values are fetched as they stand today. Later revisions to the data count as a small leak.
  • The jev-v0 trainer ignores per-example weights, so a model trained on this pack learns the downsampled positive rates (29% and 2.7%), not the natural ones.
  • Warning updates and extensions aren't harvested.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WH1hPpf3u1xWecEhS6AHD2


Generated by Claude Code

A stdlib dataset factory (apps/atlas-outcomes) that turns USGS daily flow
and IEM VTEC flood-warning polygons into jev-decision-schema-1.0.0 records
about a gauge at the end of day t (flow percentiles vs its own history,
upstream gauge, Atlas construct stack and active layers, recent warnings),
labelled by recorded futures: flow on t+1 above p90, a FL/FA/FF warning
polygon over the gauge within 24 h, and the next-day change level.

- run.py harvest | features | curate | baseline | train-command
- 60 gauges, 25 NWS offices; committed gzipped snapshots under sources/,
  offline byte-identical rebuild (checked in CI)
- time-split holdout (latest 365 days, one-day embargo); easy negatives
  downsampled in train only, with sample weights in provenance.jsonl
- stdlib logistic baseline vs persistence and climatology (BASELINE.json)
- provenance tier outcome-real added to the factory; datasets registered;
  ARCHITECTURE.md places the pack in the learning plane

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WH1hPpf3u1xWecEhS6AHD2
The block named 07a2146, a claude/decision-plane commit that the squash merge
left off main, so check_handoff_fresh.py --check failed lint-and-test on main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WH1hPpf3u1xWecEhS6AHD2
(cherry picked from commit d9aaa9b)
@vercel

vercel Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
arxiviq Ready Ready Preview Sep 23, 2026 11:27pm UTC
dottie-os-console Ready Ready Preview Sep 23, 2026 11:27pm UTC

Request Review

@jcdavis131
jcdavis131 marked this pull request as ready for review September 24, 2026 00:45
@jcdavis131
jcdavis131 merged commit 1409f1a into main Sep 24, 2026
11 checks passed

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved. Cursor Bugbot was not present after the first check poll, so that signal was skipped; remaining checks, review comments, and approval policies did not require human review. No reviewers were assigned.

Open in Web View Automation 

Sent by Cursor Approval Agent: Pull Request Router and Approver

This branch was successfully deployed

2 active deployments
Preview – dottie-os-console — 787a60a3 Deployed Sep 23, 2026 by vercel[bot]
Preview – arxiviq — 787a60a3 Deployed Sep 23, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants