Skip to content

Add model integration tests and advisory regression reports - #23

Merged
i-am-sijia merged 4 commits into
CTPSSTAFF:mainfrom
wsp-sag:model-ci
Sep 30, 2026
Merged

i-am-sijia merged 4 commits into
CTPSSTAFF:mainfrom
wsp-sag:model-ci

Conversation

@jpn--

@jpn-- jpn-- commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

Changes

Add automated model validation against the released dependencies in uv.lock. Every pull request runs input/contract tests and the full model sequence for a fixed sample of 2,000 complete households, retaining the full zone system and skims. Structural failures fail CI; changes in modeled distributions remain advisory.

  • Check input schemas, household IDs, skim dimensions and zone mappings, population preservation, output relationships, and trip structure.
  • Add weekly/manual 10,000-household and single-process runs.
  • Publish diagnostic logs, advisory baseline comparisons, runtime, memory, versions, and input/configuration hashes as Actions artifacts and summaries.
  • Document local execution, fixture provenance, baseline review, and how to extend the tests.

No behavioral model configurations, input population files, or dependency versions are changed. Constraint/telework component-specific tests can be added when those extensions land on main.

Validation

All jobs passed on the exact branch head in GitHub Actions on the driftlesslabs fork:

  • 17 input and contract tests.
  • 2,000-household multiprocess run: 139 seconds, 1.72 GiB peak summed process RSS.
  • 10,000-household multiprocess run: 166 seconds, 1.99 GiB peak summed process RSS.
  • 2,000-household single-process run.

Local model runs, Ruff checks, YAML lint, and whitespace checks also passed. Summed RSS may count shared pages more than once. Scheduling fallback counts remain advisory (7 trips in the smaller run and 52 in the larger run).

Weekly scheduling activates on the default branch after merge. Branch protection can then require the contracts and model checks.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Tour-sequence validation is incorrect for normal round trips, and extension loading and summary metadata requirements are incomplete.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity · 1 Medium severity

Open (2)
What changed in this PR

Adds automated model input/contract tests, regression baselines, diagnostics, and CI workflows.

Changes:

  • Adds fixed-sample and scheduled model validation.
  • Adds advisory baselines and runtime/memory diagnostics.
  • Documents local testing and CI workflows.
File Description
tests/​test_model_inputs.py Validates inputs, IDs, skims, and zones.
tests/​test_model_ci.py Tests output-contract validation.
tests/​model/​settings.yaml Configures test-scale model runs.
tests/​baselines/​2000.json Stores small-run baseline metrics.
tests/​baselines/​10000.json Stores large-run baseline metrics.
scripts/​model_ci.py Runs and validates model integrations; tour sequencing, extension loading, and summary metadata require changes.
README.md Documents automated model tests.
docs/​testing.md Documents testing, fixtures, diagnostics, and baselines.
.github/​workflows/​model-tests.yml Defines CI and scheduled validation jobs.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/model_ci.py Outdated
Comment thread scripts/model_ci.py Outdated
jpn-- and others added 2 commits September 23, 2026 16:21
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
ActivitySim 1.5.1 initialize_from_tours assigns trip_num and trip_count per direction. Preserve that contract and the independently scheduled leg times, while requiring both directions, matching tour endpoints, and a connected full-tour path. Add multistop round-trip coverage and corrupt-output regression cases.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Address the two moderate validation gaps before approval.

Review effort: Lite
Findings: None

Resolved since last review (2)

@i-am-sijia
i-am-sijia merged commit ef8e166 into CTPSSTAFF:main Sep 30, 2026
3 checks passed
@i-am-sijia
i-am-sijia deleted the model-ci branch October 6, 2026 16:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants