Skip to content

Commit 780168a

Browse files
anandgupta42claude
andcommitted
test: [#1406] learn benchmark harness, task project and lesson pools
Moved out of #1405 unchanged (from `feat/rsi-workspace-learning`), so the product PR stays reviewable. Produces the numbers in `research/rsi-workspace-learning-2026-09-30/learn-v1-results.md`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
1 parent 3ab191c commit 780168a

131 files changed

Lines changed: 12285 additions & 0 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.
Lines changed: 130 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,130 @@
1+
# acme-shop RSI demo: does the agent learn team conventions over time?
2+
3+
For humans running the experiment. **The agent never sees `demo/`**; it only gets a workdir made by
4+
`prepare_workdir.py` (a copy of `project/`) and the one-line ticket prompt.
5+
6+
```
7+
demo/
8+
project/ dbt-duckdb project "acme_shop" (the agent's repo; README does not document the conventions)
9+
verifier/
10+
check.py hidden CI: check.py <workdir> <task_id> -> JSON
11+
tasks/*.json 14 tasks: id, split, prompt, source, table, target_model, check_type, setup, verify
12+
gold/ hand-written gold solutions (train-refunds, heldout-invoices, 2 controls)
13+
gold_playbook.md upper-bound arm: the conventions as a SKILL.md body
14+
selftest.py gold must pass; naive / over-applied / garbage must fail
15+
prepare_workdir.py <task_id> <dest>: copy project, git init (+fake origin), commit, dbt seed. Idempotent.
16+
run_task.sh prepare | check | selftest | list wrapper
17+
```
18+
19+
## Commands
20+
21+
```bash
22+
export DBT_BIN=/path/to/dbt # default: /private/tmp/claude-501/-Users-anandgupta-codebase-altimate-code/5e228db8-69ac-4824-86f1-4a9ad4ff2e5c/scratchpad/dbtenv/bin/dbt
23+
# (dbt-core 1.12 + dbt-duckdb; DBT_PYTHON defaults to the python next to it, needs duckdb)
24+
./run_task.sh list
25+
./run_task.sh prepare train-refunds /tmp/wd # fresh workdir; prints the ticket prompt. Re-running rebuilds it.
26+
# ... agent works in /tmp/wd with the prompt ...
27+
./run_task.sh check train-refunds /tmp/wd # JSON; exit 0 iff all checks ok (~6s, never modifies /tmp/wd)
28+
./run_task.sh selftest # ~55s; --all-gold solves all 12 staging tasks programmatically; --all-naive
29+
```
30+
31+
A generic harness reads each `tasks/<id>.json`: `setup` and `verify` are argv templates (run from `demo/`, `{workdir}`
32+
substituted). Output: `{"task_id","pass","score","checks":[{"name","kind","ok","message"}]}`; `score` = fraction of checks ok,
33+
`kind` is `lint` (message may name the rule) or `data` (symptom only). Per-check `ok` gives per-convention scores.
34+
35+
The verifier copies the workdir to a temp dir, restores pristine `seeds/`, `macros/`, `dbt_project.yml`, `profiles.yml`,
36+
reseeds a fresh duckdb and builds there, so editing seeds/macros or leaving a stale db cannot change the verdict.
37+
38+
## Hidden conventions (the ground truth)
39+
40+
| id | convention | check kind |
41+
|----|-----------|-----------|
42+
| C1 | `models/staging/<source>/stg_<source>__<entity>.sql` (entity = plural table minus `raw_`) | lint |
43+
| C2 | PK `id` -> `<singular>_id`; model declared in `_<source>__models.yml` with `unique` + `not_null` on it | lint |
44+
| C3 | `*_cents` columns via `{{ cents_to_dollars() }}`, renamed without `_cents`; no `*_cents` in output (values verified) | lint |
45+
| C4 | every timestamp column via `{{ to_utc() }}`, output named `<stem>_at` (`refunded_ts`->`refunded_at`, `created`->`created_at`); `date` columns exempt | data |
46+
| C5 | sources with `_is_deleted`: `where not _is_deleted` and column not exposed (verified by row count in duckdb) | data |
47+
| C6 | `dbt build --select <model>` passes (model + tests) | lint |
48+
49+
Controls (conventions must NOT be over-applied): `control-customers-vip` (add boolean `is_vip` = >=3 orders to
50+
`stg_shop__customers`; K1 build+tests, K2 values, K3 existing columns/rows unchanged, K4 `stg_shop__orders` untouched) and
51+
`control-payments-by-month` (ad-hoc `analyses/payments_by_month.sql`, totals must stay in cents; K1 exists, K2 compiles,
52+
K3 columns, K4 numbers match). Applying `cents_to_dollars` or `_is_deleted` filtering blindly fails them.
53+
54+
## Splits
55+
56+
| split | task | source | C3 money | C4 timestamps | C5 soft delete |
57+
|-------|------|--------|:--:|:--:|:--:|
58+
| train | train-refunds | shop | x | x | x |
59+
| train | train-payments | billing | x | x | |
60+
| train | train-shipments | shop | x | x | x |
61+
| train | train-coupons | shop | x | | x |
62+
| val | val-subscriptions | billing | x | x | x |
63+
| val | val-products | shop | x | x | x |
64+
| val | val-plans | billing | x | x | |
65+
| val | val-csat-surveys | support | | x | x |
66+
| heldout | heldout-invoices | billing | x | x | x |
67+
| heldout | heldout-support-tickets | support | | x | x |
68+
| heldout | heldout-disputes | billing | x | x | x |
69+
| heldout | heldout-ledger-entries | billing | x | x | |
70+
| control | control-customers-vip | shop | existing model; no convention applies | | |
71+
| control | control-payments-by-month | billing | analysis; keep cents, no convention applies | | |
72+
73+
Every learning split exercises C3, C4 and C5. No non-staging "heldout" variant was added (it could not be made fair:
74+
the conventions are staging-specific); the two control tasks play that role.
75+
76+
Note on `heldout-support-tickets`: the first runs showed every arm (including the gold-playbook arm) naming the model
77+
`stg_support__tickets` instead of the expected `stg_support__support_tickets`, which failed all of C1-C6 and scored 0.00
78+
everywhere. The prompt now names the entity ("...; the entity is support_tickets") to remove that ambiguity. C1 remains
79+
strict on the name; when the expected file is missing and exactly one new/changed `stg_support__*.sql` exists (vs the
80+
pristine `project/`), C2-C6 are evaluated on that model so the task still measures the conventions. Zero or several
81+
candidates keep the all-fail behaviour.
82+
83+
## Fair-evaluation notes
84+
85+
- **Inferable from existing code** (`stg_shop__customers/orders`, `_shop__models.yml`, `macros/`): C1 file naming/location,
86+
the CTE pattern, C2 (`customer_id`/`order_id` PK rename, YAML location, `unique`+`not_null`), and C3 (orders uses
87+
`cents_to_dollars` and drops `_cents`). Also `macros/to_utc.sql` exists, so the macro is discoverable, but no model uses it.
88+
- **Not inferable from the repo**: C4 (use of `to_utc`, the `_at` suffix) and C5 (soft-delete filtering; no existing source
89+
has `_is_deleted`). These can only be learned from verifier feedback (or the playbook). C4/C5 feedback is deliberately
90+
symptom-only (e.g. "returned 40 rows; reconciliation expects 37"); C1-C3 feedback names the rule, like a linter.
91+
- First attempts should therefore pass C1/C2/C3/C6 often and fail C4/C5; gains over time should show on C4/C5 in val/heldout.
92+
- C3/C4 text checks are regexes on comment-stripped SQL (`cents_to_dollars(col)`, `to_utc(col)`); equivalent hand-written SQL fails them.
93+
C4 expects `<col minus _ts>_at`; any other `_at` spelling is reported as violating the suffix convention.
94+
- Timestamp columns are detected from the seed types (`timestamp`); `date` columns (`order_date`, `expires_on`) are exempt.
95+
- `train-coupons` has no timestamp column and `*-plans`, `ledger-entries`, `payments` have no `_is_deleted`; the checks report "not applicable" (ok).
96+
- Over-eager agents that blanket-apply `where not _is_deleted` to customers fail the control build; those converting to dollars in the analysis fail K4.
97+
- Self-test results below were produced with dbt-core 1.12 / dbt-duckdb; each verifier call takes about 6s.
98+
99+
## Self-test results (`./run_task.sh selftest`)
100+
101+
```
102+
gold:train-refunds pass=True score=1.0 (C1..C6 all ok; 22 rows == 25 minus 3 deleted)
103+
gold:heldout-invoices pass=True score=1.0 (37 rows == 40 minus 3 deleted)
104+
gold:control-customers-vip pass=True score=1.0 (5 VIPs)
105+
gold:control-payments-by-month pass=True score=1.0 (35 month/method rows match)
106+
naive:train-refunds pass=False score=0.333 (C1, C6 ok; C2..C5 FAIL)
107+
naive:heldout-invoices pass=False score=0.333
108+
overapplied:control-payments-by-month pass=False score=0.75
109+
overapplied:control-customers-vip pass=False score=0.25
110+
empty/garbage workdir, unknown task pass=False (no crash; every check reports a message)
111+
./run_task.sh selftest --all-gold: all 12 staging tasks solved by a generator pass (score 1.0)
112+
SELFTEST OK
113+
```
114+
115+
Naive solution (`select * from {{ source('shop', 'refunds') }}`, no YAML) on `train-refunds`:
116+
117+
```
118+
[ok] C1_location_naming (lint): found models/staging/shop/stg_shop__refunds.sql
119+
[FAIL] C2_primary_key_and_tests (lint): stg_shop__refunds: model is not declared in models/staging/shop/_shop__models.yml. Team rule: each source folder has one _shop__models.yml with its models; column `refund_id` is missing unique and not_null test(s) in models/staging/shop/_shop__models.yml; primary key `id` is not renamed to `refund_id` (Team rule: <singular_entity>_id).
120+
[FAIL] C3_money_cents_to_dollars (lint): stg_shop__refunds: column `amount_cents` exposed raw. Team rule: money columns must be converted with {{ cents_to_dollars() }} and renamed without the _cents suffix (amount); column(s) `amount_cents` still carry the _cents suffix. Team rule: no *_cents columns in staging output.
121+
[FAIL] C4_timestamps_utc_at (data): stg_shop__refunds: column `refunded_ts` is not timezone-normalized; the downstream join against the finance calendar (UTC) fails on it; timestamp column(s) `refunded_ts` pass through un-normalized and un-renamed.
122+
[FAIL] C5_soft_deletes (data): stg_shop__refunds: returned 25 rows; the reconciliation against the source system expects 22. 3 row(s) should not reach analytics; exposes internal column `_is_deleted`, which must not reach analytics.
123+
[ok] C6_dbt_build (lint): `dbt build --select stg_shop__refunds` succeeded (model and its tests)
124+
```
125+
126+
Over-applied control (`cents_to_dollars` in the analysis):
127+
128+
```
129+
[FAIL] K4_numbers_match_finance_export (data): 70 month/method row(s) differ from the finance export, e.g. ('2025-01-01', 'bank_transfer'): got payments/total_cents ('1', '94') vs expected ('1', '9410')
130+
```
Lines changed: 64 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,64 @@
1+
#!/usr/bin/env python3
2+
"""prepare_workdir.py <task_id> <dest>
3+
4+
Copy project/ into a fresh dir, init a repo with a fixed fake remote, commit,
5+
and run `dbt seed`. Idempotent: an existing dest previously created by this
6+
script (marker file) is rebuilt from scratch; any other non-empty dest is refused.
7+
Env: DBT_BIN (default documented in README.md). Prints the task prompt.
8+
"""
9+
import json
10+
import os
11+
import shutil
12+
import subprocess
13+
import sys
14+
15+
HERE = os.path.dirname(os.path.abspath(__file__))
16+
PROJECT = os.path.join(HERE, "project")
17+
MARKER = ".prepared"
18+
DBT_BIN = os.environ.get(
19+
"DBT_BIN",
20+
"/private/tmp/claude-501/-Users-anandgupta-codebase-altimate-code/"
21+
"5e228db8-69ac-4824-86f1-4a9ad4ff2e5c/scratchpad/dbtenv/bin/dbt",
22+
)
23+
GIT_ENV = {
24+
"GIT_AUTHOR_NAME": "Acme Dev", "GIT_AUTHOR_EMAIL": "dev@acme.example",
25+
"GIT_COMMITTER_NAME": "Acme Dev", "GIT_COMMITTER_EMAIL": "dev@acme.example",
26+
"GIT_AUTHOR_DATE": "2026-01-01T00:00:00Z", "GIT_COMMITTER_DATE": "2026-01-01T00:00:00Z",
27+
}
28+
29+
30+
def sh(cmd, cwd, env=None):
31+
p = subprocess.run(cmd, cwd=cwd, env=dict(os.environ, **(env or {})), capture_output=True, text=True)
32+
if p.returncode != 0:
33+
sys.exit(f"command failed: {' '.join(cmd)}\n{p.stdout}{p.stderr}")
34+
35+
36+
def main():
37+
if len(sys.argv) != 3:
38+
sys.exit(__doc__)
39+
task_id, dest = sys.argv[1], os.path.abspath(sys.argv[2])
40+
tpath = os.path.join(HERE, "verifier", "tasks", task_id + ".json")
41+
if not os.path.isfile(tpath):
42+
sys.exit(f"unknown task {task_id}")
43+
task = json.load(open(tpath))
44+
if os.path.exists(dest):
45+
if os.path.isfile(os.path.join(dest, MARKER)) or (os.path.isdir(dest) and not os.listdir(dest)):
46+
shutil.rmtree(dest)
47+
else:
48+
sys.exit(f"refusing to overwrite {dest}: not created by prepare_workdir.py")
49+
shutil.copytree(PROJECT, dest, ignore=shutil.ignore_patterns("target", "logs", "*.duckdb", "*.duckdb.wal", ".user.yml"))
50+
open(os.path.join(dest, MARKER), "w").write(task_id + "\n")
51+
with open(os.path.join(dest, ".gitignore"), "a") as f:
52+
f.write(MARKER + "\n")
53+
sh(["git", "init", "-q", "-b", "main"], dest)
54+
sh(["git", "remote", "add", "origin", "git@github.com:acme/acme-shop.git"], dest)
55+
sh(["git", "add", "-A"], dest)
56+
sh(["git", "commit", "-q", "-m", "chore: initial acme-shop dbt project"], dest, GIT_ENV)
57+
sh([DBT_BIN, "seed", "--profiles-dir", dest, "--project-dir", dest], dest,
58+
{"DBT_SEND_ANONYMOUS_USAGE_STATS": "false", "DO_NOT_TRACK": "1"})
59+
print(f"workdir ready: {dest}")
60+
print(f"task {task_id} ({task['split']}): {task['prompt']}")
61+
62+
63+
if __name__ == "__main__":
64+
main()
Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,4 @@
1+
target/
2+
logs/
3+
*.duckdb
4+
*.duckdb.wal
Lines changed: 23 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,23 @@
1+
# acme-shop
2+
3+
dbt project for Acme's e-commerce analytics (duckdb locally).
4+
5+
## Layout
6+
7+
- `seeds/` - raw extracts loaded into the `raw` schema (stand-ins for the warehouse landing tables)
8+
- `models/staging/<source>/` - one folder per source system (`shop`, `billing`, `support`)
9+
- `macros/` - shared SQL helpers
10+
11+
## Running
12+
13+
```bash
14+
dbt seed --profiles-dir . --project-dir .
15+
dbt build --profiles-dir . --project-dir .
16+
```
17+
18+
Run from the project root; the duckdb file `acme.duckdb` is created here.
19+
`profiles.yml` lives in this directory, so pass `--profiles-dir .`.
20+
21+
## Contributing
22+
23+
Open a PR against `main`. CI runs `dbt build` and the team's review checks.
Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,19 @@
1+
name: acme_shop
2+
version: "1.0.0"
3+
config-version: 2
4+
profile: acme_shop
5+
6+
model-paths: ["models"]
7+
seed-paths: ["seeds"]
8+
macro-paths: ["macros"]
9+
10+
clean-targets: ["target", "logs"]
11+
12+
seeds:
13+
acme_shop:
14+
+schema: raw
15+
16+
models:
17+
acme_shop:
18+
staging:
19+
+materialized: view
Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
1+
{% macro cents_to_dollars(col) %}round({{ col }} / 100.0, 2){% endmacro %}
Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,3 @@
1+
{% macro generate_schema_name(custom_schema_name, node) -%}
2+
{%- if custom_schema_name is none -%}{{ target.schema }}{%- else -%}{{ custom_schema_name | trim }}{%- endif -%}
3+
{%- endmacro %}
Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
1+
{% macro to_utc(col) %}cast({{ col }} as timestamp){% endmacro %}
Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,18 @@
1+
version: 2
2+
3+
sources:
4+
- name: billing
5+
schema: raw
6+
tables:
7+
- name: payments
8+
identifier: raw_payments
9+
- name: subscriptions
10+
identifier: raw_subscriptions
11+
- name: invoices
12+
identifier: raw_invoices
13+
- name: plans
14+
identifier: raw_plans
15+
- name: disputes
16+
identifier: raw_disputes
17+
- name: ledger_entries
18+
identifier: raw_ledger_entries
Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,18 @@
1+
version: 2
2+
3+
models:
4+
- name: stg_shop__customers
5+
description: Customers, one row per customer.
6+
columns:
7+
- name: customer_id
8+
tests:
9+
- unique
10+
- not_null
11+
12+
- name: stg_shop__orders
13+
description: Orders, one row per order.
14+
columns:
15+
- name: order_id
16+
tests:
17+
- unique
18+
- not_null

0 commit comments

Comments
 (0)