File tree Expand file tree Collapse file tree
livemathematicianbench_id_split Expand file tree Collapse file tree Original file line number Diff line number Diff line change 1+ {
2+ "fork_synced_at" : " 2026-07-26T09:38:34.869801+00:00" ,
3+ "commits_behind_before_sync" : 210 ,
4+ "action_taken" : " synced"
5+ }
Original file line number Diff line number Diff line change @@ -29,7 +29,7 @@ Each `items.json` contains only stable IDs or source-path hints.
2929| Manifest directory | Benchmark | Counts | Coverage | Raw data source | ` split_dir ` |
3030| ---| ---| ---:| ---| ---| ---|
3131| ` searchqa_id_split/ ` | SearchQA | 400 / 200 / 1400 | Official HF dataset IDs | [ lucadiliello/searchqa] ( https://huggingface.co/datasets/lucadiliello/searchqa ) | ` data/searchqa_split ` |
32- | ` livemathematicianbench_id_split/ ` | LiveMathematicianBench | 35 / 18 / 124 | Four official monthly files | [ LiveMathematicianBench/LiveMathematicianBench] ( https://huggingface.co/datasets/LiveMathematicianBench/LiveMathematicianBench ) | ` data/livemathematicianbench_split ` |
32+ | ` livemathematicianbench_id_split/ ` | LiveMathematicianBench | 35 / 17 / 125 | Four official monthly files | [ LiveMathematicianBench/LiveMathematicianBench] ( https://huggingface.co/datasets/LiveMathematicianBench/LiveMathematicianBench ) | ` data/livemathematicianbench_split ` |
3333| ` docvqa_id_split/ ` | DocVQA | 107 / 53 / 374 | 10% subset of validation | [ lmms-lab/DocVQA] ( https://huggingface.co/datasets/lmms-lab/DocVQA ) | ` data/docvqa/splits ` |
3434| ` officeqa_id_split/ ` | OfficeQA | 50 / 24 / 172 | OfficeQA Full | [ databricks/officeqa] ( https://huggingface.co/datasets/databricks/officeqa ) | ` data/officeqa_split ` |
3535| ` spreadsheetbench_id_split/ ` | SpreadsheetBench | 80 / 40 / 280 | SpreadsheetBench Verified 400 | [ KAKA22/SpreadsheetBench] ( https://huggingface.co/datasets/KAKA22/SpreadsheetBench ) | ` data/spreadsheetbench_split ` |
Original file line number Diff line number Diff line change 1616 "split_seed" : 42 ,
1717 "counts" : {
1818 "train" : 35 ,
19- "val" : 18 ,
20- "test" : 124
19+ "val" : 17 ,
20+ "test" : 125
2121 },
2222 "item_fields" : [
2323 " id" ,
Original file line number Diff line number Diff line change 1+ import json
2+ from pathlib import Path
3+
4+ from skillopt .datasets .base import SPLIT_NAMES
5+
6+ DATA_DIR = Path (__file__ ).resolve ().parents [1 ] / "data"
7+
8+
9+ def test_split_manifest_counts_match_item_files () -> None :
10+ manifest_paths = sorted (DATA_DIR .glob ("*/split_manifest.json" ))
11+ assert manifest_paths , f"No split manifests found under { DATA_DIR } "
12+
13+ for manifest_path in manifest_paths :
14+ manifest = json .loads (manifest_path .read_text (encoding = "utf-8" ))
15+
16+ for split_name in SPLIT_NAMES :
17+ declared_count = manifest ["counts" ][split_name ]
18+ items_path = manifest_path .parent / split_name / "items.json"
19+ actual_count = len (json .loads (items_path .read_text (encoding = "utf-8" )))
20+
21+ assert declared_count == actual_count , (
22+ f"{ manifest_path } split { split_name !r} : declared count { declared_count } , actual count { actual_count } "
23+ )
You can’t perform that action at this time.
0 commit comments