Skip to content

Commit 7c41bf8

Browse files
anandgupta42claude
andcommitted
docs(learn): record the verifier weaknesses kept unchanged for comparability
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
1 parent 3f22fea commit 7c41bf8

1 file changed

Lines changed: 1 addition & 0 deletions

File tree

‎research/rsi-workspace-learning-2026-09-30/learn-v1-results.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -145,3 +145,4 @@ Most control failures are runs that did not complete (usually an exhausted turn
145145
- **Over-application from keyword retrieval:** a lesson scoped to staging models can be retrieved for an analysis request that mentions the same word ("cents") and is sometimes applied there (1 of 6 control runs in 3 retrieval arms). Showing the lesson's scope reduced but did not remove it. Candidate follow-up: weaker framing for retrieved lessons than "Team rules for this request", or require path compatibility before showing a scoped lesson.
146146
- **Retrieval vs loading everything:** on Gemini, loading all 1,000 lessons scored as well as retrieval; retrieval's value is cost (−26% per run at 1,000) and context headroom, not quality. Not measured on Claude (blocked).
147147
- **One project, one model, 9 held-out runs per arm:** differences of one run are noise.
148+
- **Verifier and fixtures are kept as run, with known weaknesses:** C3/C4 recognise the conventions by the macro call in SQL rather than by comparing every output value; some synthetic seed rows are inconsistent (for example refunds dated before their orders); and `to_utc` does not perform a real timezone conversion. Tightening any of these changes what passes, so it needs a rerun, not a rescore.

0 commit comments

Comments
 (0)