You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 84bbde9
Browse filesBrowse the repository at this point in the historyBrowse files
feat(sleep): baseline-relative adversarial probes with rollout aggregation
- blocking is baseline-relative: identical source/probe pairs scored under
baseline and candidate docs; a row is brittle only when the candidate gap
worsens beyond the margin in a strict majority of rollout indices
- evidence retains all four aggregated scores plus per-rollout samples
- dream_adversarial_rollouts (cap 8); blocking requires >= 2
- _strip_polite_frame restricted to politeness-marked requests with negative
tests for ability/permission/desire forms
- baseline documents are required arguments so no caller can silently
compare against an unintended baseline
`evolve_skill`. Recalled archive tasks are restricted to that same skill hint;
314
315
shared memory is read-only in fan-out runs. Setting `evolve_skill` to `false`
315
316
therefore disables per-skill proposals as well as the managed skill proposal.
@@ -341,20 +342,32 @@ unless explicit adversarial blocking adds a second rejection condition.
341
342
|`recall_k`|`0`| Associative recall — pull the K most-similar past tasks (from a persisted archive) into tonight's dream. |
342
343
|`dream_factor`|`0`| Add N lightweight synthetic variants of each task. |
343
344
|`dream_adversarial`|`0`| Score up to N harmless request-frame variants per real training task against each gate-eligible candidate. The factor is capped at 3 per task and 256 probes per candidate. |
344
-
|`dream_adversarial_blocking`|`false`| When `true`, reject a candidate if any probe drops beyond the configured margin. When `false`, surface the same evidence without changing the gate decision. |
345
-
|`dream_adversarial_margin`|`0.0`| Tolerated source-to-probe score drop in `[0, 1]` before a row is marked brittle. |
345
+
|`dream_adversarial_blocking`|`false`| When `true`, reject a candidate whose brittleness is candidate-introduced under the baseline-relative rule below. When `false`, surface the same evidence without changing the gate decision. Requires `dream_adversarial_rollouts >= 2`. |
346
+
|`dream_adversarial_margin`|`0.0`| Tolerated worsening of the candidate gap relative to the baseline gap, in `[0, 1]`, before a row is marked brittle. Calibrate it on your own task mix before enabling blocking. |
347
+
|`dream_adversarial_rollouts`|`1`| Repeated samples per task and arm (capped at 8). Blocking requires at least 2 so one stochastic sample can never reject a candidate. |
346
348
347
349
Adversarial probes preserve the source reference and judge but change only the
348
-
request frame (for example, removing polite boilerplate or adding explicit
349
-
request delimiters). They are generated from real, underived training tasks
350
-
only. Recalled, already-synthetic, validation, and test tasks are excluded.
351
-
The report records every source/probe score and the exact perturbation that
352
-
failed. Probes are advisory first because any fixed robustness suite is an
353
-
incomplete proxy; enable blocking only after reviewing its behavior on your
354
-
task mix. Blocking mode fails closed if no eligible probe can be generated.
355
-
Each candidate adds one source rollout per eligible task plus N probe rollouts,
356
-
so token and latency cost grow with the number of real training tasks and the
357
-
selected factor.
350
+
request frame (for example, removing explicitly politeness-marked boilerplate
351
+
or adding request delimiters). They are generated from real, underived
352
+
training tasks only. Recalled, already-synthetic, validation, and test tasks
353
+
are excluded.
354
+
355
+
The decision is **baseline-relative** so pre-existing frame sensitivity never
356
+
flags a candidate: every source/probe pair is scored under both the current
357
+
(baseline) documents and the candidate documents, each score is the mean of
358
+
`dream_adversarial_rollouts` repeated samples, and a row is brittle only when
359
+
the candidate's probe-minus-source gap worsens beyond the margin relative to
360
+
the baseline gap AND the worsening holds in a strict majority of rollout
361
+
indices. All four aggregated scores and the per-rollout samples are retained
362
+
in the evidence so the decision is auditable. Any non-finite score fails
363
+
closed.
364
+
365
+
Probes are advisory first because any fixed robustness suite is an incomplete
366
+
proxy; enable blocking only after reviewing the advisory evidence and
367
+
calibrating the margin on your task mix. Blocking mode fails closed if no
368
+
eligible probe can be generated. The replay cost per gate-eligible candidate
369
+
is `rollouts * 2 * (sources + probes)`, so token and latency cost grow with
370
+
the number of real training tasks, the factor, and the rollout count.
0 commit comments