Skip to content

datasets: validators for rag and incident; point crossframework at the crosswalk - #92

Merged
emmanuelgjr merged 1 commit into
mainfrom
datasets/rag-incident-validators
Oct 9, 2026
Merged

emmanuelgjr merged 1 commit into
mainfrom
datasets/rag-incident-validators

Conversation

@emmanuelgjr

Copy link
Copy Markdown
Contributor

Summary

Three datasets showed NO VALIDATOR in run_all_checks.py. This PR closes that gap for two of them and retires the third.

rag_dataset and incident_dataset: new validate.py

Both follow the pattern of the existing dataset validators. Each one checks every entries/*.json:

An empty entries/ passes ("no entries yet"), so incident_dataset stays green until its first contribution. A schema violation, bad date, unknown DSGAI ID, ID/file-name mismatch or duplicate ID each fail (tested on a scratch copy).

RAG-0001.json (contributed in #58) moves to rag_dataset/entries/ and gains a $schema line, like entries in the other datasets. Its content is unchanged.

The two READMEs replace their "TODO: Define schema" placeholders with the schema link, the entries/ layout and the validate.py command. Incident entries are JSON only now; the README used to also allow Markdown, which can't be validated.

crossframework_mapping_dataset: superseded

This folder never held data. Framework mappings live in the GenAI Crosswalk (26 frameworks, 3,771 mappings), and new mapping proposals go there. Its README (which still pointed at a /mappings directory that no longer exists) and the datasets table now say so. The folder and its schema stay, so existing links still resolve.

Verification

Merge order

Merge this before #89: #89's README and docstring examples will be updated to the new entries/RAG-0001.json path.

🤖 Generated with Claude Code

…e crosswalk

- rag_dataset/validate.py and incident_dataset/validate.py check every
  entries/*.json against data_validation/schemas/{rag,incident}.schema.json
  (draft-07, with date formats checked), DSGAI IDs against the shared
  taxonomy, ID-matches-filename and duplicate IDs. An empty entries/ passes,
  so incident_dataset is green until its first entry.
- RAG-0001.json moves to rag_dataset/entries/ and gains a $schema line, like
  the other datasets' entries; its content is unchanged.
- rag and incident READMEs: the "TODO: define schema" placeholders now point
  at the schema, the entries/ layout and validate.py; incident entries are
  JSON only, since Markdown cannot be validated.
- crossframework_mapping_dataset never held data and framework mappings now
  live in the GenAI Crosswalk repo, so its README and the datasets table
  point there instead of inviting contributions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
emmanuelgjr added a commit that referenced this pull request Oct 9, 2026
…path)

rag_dataset and incident_dataset get their own validate.py in #92,
crossframework_mapping_dataset is superseded by the crosswalk repo, and
RAG-0001.json moves to rag_dataset/entries/. Update the README dataset
table, the example paths in README, SETUP and the validator docstrings,
and the suggested first contribution.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@emmanuelgjr
emmanuelgjr merged commit 9cd3711 into main Oct 9, 2026
6 checks passed
emmanuelgjr added a commit that referenced this pull request Oct 9, 2026
… data (#89)

* data_validation: implement the validators and QC tools, fix reference data

The validators/ and qc_tools/ scripts were stubs that printed "contributions
welcome", the tests passed without testing anything, and the MITRE ATLAS
reference table had the wrong name for 21 of its 28 IDs.

- run_all_checks.py also finds validate.py in dataset sub-collections (the
  Turkish pairs validator now runs in CI), gives each a timeout, and then
  runs the shared checks on every data file. Errors fail the run; heuristic
  hits are warnings (--strict makes them fail). #87's output format and
  exit codes are kept.
- validators/: schema (finds each file's schema; covers datasets with no
  validate.py, folding in completeness_check), DSGAI IDs, CVE/GHSA/CWE/ATLAS
  IDs, anonymization patterns, duplicates. Shared helpers in _common.py.
- qc_tools/: coverage report, anomaly list, cross-dataset reference check.
  CI posts the report and anomalies to the job summary.
- reference_data/: ATLAS (2026.09, with retired IDs marked) and CWE (4.20)
  tables regenerated from MITRE by update_reference_data.py; sources in
  SOURCES.md. Header-only framework control CSVs removed.
- datasets/_shared/dsgai_taxonomy.json: DSGAI07 and DSGAI12 use the paper's
  full titles, matching the README and dsgai_entries.json.
- Real tests for every check (82); CI runs the whole suite.
- README and SETUP describe what exists; Python 3.10+; unused pandas and
  regex dependencies dropped; venv/ gitignored as SETUP already claimed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* data_validation: docs follow #92 (rag/incident validators, RAG entry path)

rag_dataset and incident_dataset get their own validate.py in #92,
crossframework_mapping_dataset is superseded by the crosswalk repo, and
RAG-0001.json moves to rag_dataset/entries/. Update the README dataset
table, the example paths in README, SETUP and the validator docstrings,
and the suggested first contribution.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@emmanuelgjr
emmanuelgjr deleted the datasets/rag-incident-validators branch October 9, 2026 20:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant