diff --git a/apps/dottie/README.md b/apps/dottie/README.md index 7a6d32cb..858297cd 100644 --- a/apps/dottie/README.md +++ b/apps/dottie/README.md @@ -22,7 +22,7 @@ See [`DOTTIE_PRIME_SOTA.md`](./DOTTIE_PRIME_SOTA.md) for the full prime → Dott ## Honest capability statement (read first) -- **Ollama is the working brain today.** Only backend that does useful work is `ollama` (default `qwen3:32b`) served locally. +- **Ollama is the working brain today.** Only backend that does useful work is `ollama` (default `qwen3:8b`) served locally. - **Ava is the trainee.** `ava` backend decodes from real smoke-scale checkpoint (~14M nano). Zero task capability today, emits noise — honestly. Exists so flywheel has a trainee and serving path is built for day a capable ckpt exists. - **Echo is plumbing.** Deterministic CI harness (`plumbing_only=True`). - **Anti-fabrication everywhere.** Unreachable Ollama, missing ckpt, missing torch → Dottie refuses with true reason (`DottiePolicyUnavailable` / 503). Every metric computed from real inputs; `r_task` for free-form is `null` (no verifier), never invented. Verified tasks (`compute`, `extract`, `tool_chain`, `file_ops`, `constraint`) have deterministic verifier from same values rendered into prompt — automated no-leakage check enforces it. @@ -41,7 +41,7 @@ See [`DOTTIE_PRIME_SOTA.md`](./DOTTIE_PRIME_SOTA.md) for the full prime → Dott rlm("...") ──▶│ + RLM Runner (rlm.py) ──▶ Harness v2 (harness_continual.py) │ │ │ ContinualHarness refined via evidence, snapshots, rollback │ │ │ Sessions daemon Registry + inbox messaging + goals + heartbeat │ - │ ├─ OllamaPolicy (qwen3:32b) ─ brain │ + │ ├─ OllamaPolicy (qwen3:8b) ─ brain │ │ ├─ AvaPolicy (TorchModelPolicy+ckpt) ─ trainee │ │ └─ EchoPolicy (deterministic CI) │ │ │ │ @@ -62,7 +62,7 @@ See [`DOTTIE_PRIME_SOTA.md`](./DOTTIE_PRIME_SOTA.md) for the full prime → Dott ```bash # 1. Ollama brain ollama serve & -ollama pull qwen3:32b +ollama pull qwen3:8b # 2. Install Dottie SOTA pip install -e apps/dottie # provides dottie CLI diff --git a/apps/dottie/RUNBOOK_4080.md b/apps/dottie/RUNBOOK_4080.md index 4ef3f085..3160dae4 100644 --- a/apps/dottie/RUNBOOK_4080.md +++ b/apps/dottie/RUNBOOK_4080.md @@ -10,7 +10,7 @@ with Ollama as the working brain and your fresh mini checkpoint as the trainee. | backend | what a climb iteration will measure | |---|---| -| `ollama` | Real task capability of your local model (e.g. `qwen3:32b`). Verified-task success rates here are the first REAL capability numbers in the ecosystem. | +| `ollama` | Real task capability of your local model (e.g. `qwen3:8b`). Verified-task success rates here are the first REAL capability numbers in the ecosystem. | | `ava` | Your homegrown checkpoint. A smoke/mini-scale checkpoint has **zero task capability** — expect `success_rate 0.0`. That number is the honest baseline the flywheel exists to move. | | `echo` | Deterministic plumbing (`plumbing_only`), never a capability measurement. | @@ -31,7 +31,7 @@ pip install torch --index-url https://download.pytorch.org/whl/cu128 # 4080 CU # Ollama brain ollama serve & # if not already running -ollama pull qwen3:32b # or the model you prefer; set DOTTIE_OLLAMA_MODEL to match +ollama pull qwen3:8b # or the model you prefer; set DOTTIE_OLLAMA_MODEL to match ``` ## 1. Point the trainee at your fresh mini checkpoint diff --git a/apps/dottie/RUNBOOK_RESEARCH_LOOP.md b/apps/dottie/RUNBOOK_RESEARCH_LOOP.md index 2952bd2b..f470055a 100644 --- a/apps/dottie/RUNBOOK_RESEARCH_LOOP.md +++ b/apps/dottie/RUNBOOK_RESEARCH_LOOP.md @@ -54,9 +54,9 @@ pip install torch --index-url https://download.pytorch.org/whl/cu128 # 4080 CU # Ollama is the brain for ideation + implementation ollama serve & # if not already running -ollama pull qwen3:32b # or your preferred model +ollama pull qwen3:8b # or your preferred model export DOTTIE_OLLAMA_URL=http://localhost:11434 -export DOTTIE_OLLAMA_MODEL=qwen3:32b +export DOTTIE_OLLAMA_MODEL=qwen3:8b ``` ## 1. One-time — seed the baseline @@ -180,6 +180,14 @@ Everything is under `apps/dottie/data/research/`: broken fix verbatim (observed: identical wrong output shape four attempts running). Prefer a qwen3-class model (`qwen3:14b` fits a 12 GB card); the policy strips `` blocks, and `DOTTIE_OLLAMA_THINK=false` is available when latency matters more than fix quality. +- **A completion stops mid-JSON**: every call is capped at 2048 generated tokens + (`DOTTIE_OLLAMA_NUM_PREDICT`; thinking tokens count against it). On the 4080 box's ledger + (177 experiments, measured 2026-09-27) the longest stored implementation `code` is 4,369 + characters, well inside the cap under `THINK=false`. + Raise it if a model genuinely needs more; `0` removes the cap, which brings back the + uncapped runaway (708 s / 81,920 tokens on one call, 2026-09-27). + `DOTTIE_OLLAMA_TEMPERATURE` and `DOTTIE_OLLAMA_NUM_CTX` are also read; the research stages + still set their own temperature per call. - **Everything sits in `pending`**: `implement` hasn't run (or keeps hitting `failed_validation`); check `logs/implement.log` for the failing validator level. - **`train` says "no experiments ready for training"**: nothing has passed validation yet — run diff --git a/apps/dottie/docker-compose.dottie.yml b/apps/dottie/docker-compose.dottie.yml index 7b2fce90..4db7e720 100644 --- a/apps/dottie/docker-compose.dottie.yml +++ b/apps/dottie/docker-compose.dottie.yml @@ -6,7 +6,7 @@ # curl http://localhost:8100/status # # extra_hosts maps host.docker.internal to the host gateway so the Ollama server running on -# the user's box (qwen3:32b at :11434) is reachable from inside the container — the same +# the user's box (qwen3:8b at :11434) is reachable from inside the container — the same # convention apps/ava-factory/docker-compose.yml uses for its ava-train service. services: @@ -19,7 +19,7 @@ services: - "127.0.0.1:8100:8100" environment: DOTTIE_OLLAMA_URL: ${DOTTIE_OLLAMA_URL:-http://host.docker.internal:11434} - DOTTIE_OLLAMA_MODEL: ${DOTTIE_OLLAMA_MODEL:-qwen3:32b} + DOTTIE_OLLAMA_MODEL: ${DOTTIE_OLLAMA_MODEL:-qwen3:8b} DOTTIE_DATA_DIR: /data DOTTIE_WORKERS: ${DOTTIE_WORKERS:-2} DOTTIE_QUEUE_MAX: ${DOTTIE_QUEUE_MAX:-32} diff --git a/apps/dottie/dottie/policy.py b/apps/dottie/dottie/policy.py index 0bf33f20..1d4a8b61 100644 --- a/apps/dottie/dottie/policy.py +++ b/apps/dottie/dottie/policy.py @@ -5,7 +5,7 @@ ``apps/ava-factory/ava/rl/codeact_loop.py``: ``transcript: str -> next assistant turn: str``. Three backends: - * :class:`OllamaPolicy` — real HTTP calls to an Ollama server (the user's local qwen3:32b by + * :class:`OllamaPolicy` — real HTTP calls to an Ollama server (the user's local qwen3:8b by default). This is the only backend with real task capability today. * :class:`AvaPolicy` — the trainee: wraps the factory's real ``TorchModelPolicy`` over a smoke-scale ava checkpoint. Zero capability today; exists for the training flywheel. @@ -21,14 +21,29 @@ import os import re from pathlib import Path -from typing import Any +from typing import TYPE_CHECKING, Any import httpx from dottie import resolve +if TYPE_CHECKING: + from collections.abc import Callable + DEFAULT_OLLAMA_URL = "http://host.docker.internal:11434" -DEFAULT_OLLAMA_MODEL = "qwen3:32b" +# The model the 4080 box actually serves. qwen3:32b (the old default) is not pulled there, so +# a bare run refused while the research loop, which pins qwen3:8b through its env, worked. +DEFAULT_OLLAMA_MODEL = "qwen3:8b" +# Per-call generation cap, sent as options.num_predict. With no cap, a degenerate generation +# is bounded only by the server, and llama-server runs with --context-shift, so it slides the +# window and keeps going. Measured 2026-09-27 (local-LLM A/B, granite4.2:3b at T0.2, -c 8192): +# one call ran 708 s to 81,920 tokens, done_reason=length. On the research loop's CPU setting +# the same runaway would hold the call until the 1800 s read timeout. Thinking tokens count +# against the cap, so with DOTTIE_OLLAMA_THINK unset a long qwen3 thought can use it up; the +# research loop runs think=false. DOTTIE_OLLAMA_NUM_PREDICT overrides; 0 or negative sends +# no cap. +DEFAULT_OLLAMA_NUM_PREDICT = 2048 +DEFAULT_OLLAMA_TEMPERATURE = 0.2 # Transcript markers — must match ava/rl/codeact_loop.py (and the factory datagen) exactly. USER = "<|user|>" @@ -89,6 +104,19 @@ def transcript_to_messages(transcript: str) -> list[dict[str, str]]: return messages +def _env_number(name: str, cast: Callable[[str], Any], kind: str) -> Any: + """Parse env var ``name``; unset or blank means None. A value that does not parse + raises ValueError naming the variable, so a typo refuses instead of quietly running + with the default (for the generation cap, that would mean running uncapped).""" + raw = os.environ.get(name) + if raw is None or raw.strip() == "": + return None + try: + return cast(raw.strip()) + except ValueError as e: + raise ValueError(f"{name} must be {kind}, got {raw!r}") from e + + def strip_think(text: str) -> str: """Remove closed ``...`` blocks (qwen3-style reasoning preamble). @@ -102,9 +130,21 @@ class OllamaPolicy(PolicyProvider): """Real next-turn generation over HTTP against an Ollama server (``/api/chat``). Base URL from ``DOTTIE_OLLAMA_URL`` (default ``http://host.docker.internal:11434``), model - from ``DOTTIE_OLLAMA_MODEL`` (default ``qwen3:32b``). Non-streaming; sensible timeouts - (connect fast-fails, generation may take minutes on a 32b local model). Unreachable server - or HTTP error => :class:`DottiePolicyUnavailable` with the true cause.""" + from ``DOTTIE_OLLAMA_MODEL`` (default ``qwen3:8b``). Non-streaming; sensible timeouts + (connect fast-fails, generation may take minutes on a CPU-pinned local model). Unreachable + server or HTTP error => :class:`DottiePolicyUnavailable` with the true cause. + + Request options, each an explicit argument > env > default: + + * ``num_predict`` / ``DOTTIE_OLLAMA_NUM_PREDICT`` — generation cap, default 2048 on + every call; 0 or negative sends none. + * ``temperature`` / ``DOTTIE_OLLAMA_TEMPERATURE`` — default 0.2. A per-call + ``complete(temperature=...)`` beats all three. + * ``num_ctx`` / ``DOTTIE_OLLAMA_NUM_CTX`` — context window; unset, 0 or negative sends + none, so the model's own default applies. + + ``DOTTIE_OLLAMA_NUM_GPU`` / ``_KEEP_ALIVE`` / ``_THINK`` / ``_READ_TIMEOUT_S`` are + documented where they are read.""" name = "ollama" @@ -115,7 +155,9 @@ def __init__( *, connect_timeout_s: float = 5.0, read_timeout_s: float | None = None, - temperature: float = 0.2, + temperature: float | None = None, + num_predict: int | None = None, + num_ctx: int | None = None, ) -> None: self.base_url = ( base_url or os.environ.get("DOTTIE_OLLAMA_URL") or DEFAULT_OLLAMA_URL @@ -143,7 +185,19 @@ def __init__( write=30.0, pool=connect_timeout_s, ) - self.temperature = float(temperature) + if temperature is None: + temperature = _env_number("DOTTIE_OLLAMA_TEMPERATURE", float, "a number") + self.temperature = float( + DEFAULT_OLLAMA_TEMPERATURE if temperature is None else temperature + ) + if num_predict is None: + num_predict = _env_number("DOTTIE_OLLAMA_NUM_PREDICT", int, "an integer") + if num_predict is None: + num_predict = DEFAULT_OLLAMA_NUM_PREDICT + self.num_predict: int | None = num_predict if num_predict > 0 else None + if num_ctx is None: + num_ctx = _env_number("DOTTIE_OLLAMA_NUM_CTX", int, "an integer") + self.num_ctx: int | None = num_ctx if num_ctx and num_ctx > 0 else None def __call__(self, transcript: str) -> str: messages = [{"role": "system", "content": CODEACT_SYSTEM_PROMPT}] @@ -174,6 +228,10 @@ def _chat(self, messages: list, *, temperature: float | None = None) -> str: if temperature is None else float(temperature) } + if self.num_predict is not None: + options["num_predict"] = self.num_predict + if self.num_ctx is not None: + options["num_ctx"] = self.num_ctx # DOTTIE_OLLAMA_NUM_GPU=0 pins inference to CPU (doctrine: the GPU belongs to # model TRAINING; LLM inference and everything else run on CPU). Any integer is # passed through as the number of offloaded layers. diff --git a/apps/dottie/dottie/status.py b/apps/dottie/dottie/status.py index 50e00b7e..ba3b47f8 100644 --- a/apps/dottie/dottie/status.py +++ b/apps/dottie/dottie/status.py @@ -18,7 +18,7 @@ from dottie.engine import DottieEngine CAPABILITY_NOTE = ( - "Ollama (external local model, e.g. qwen3:32b) is the only backend with real task " + "Ollama (external local model, e.g. qwen3:8b) is the only backend with real task " "capability today. The ava backend decodes from a smoke-scale checkpoint (~90 base + ~25 " "agentic optimizer steps, capability_claim=none) and exists to close the training " "flywheel, not to assist. The echo backend is a deterministic plumbing test." diff --git a/apps/dottie/dottie/tasks.py b/apps/dottie/dottie/tasks.py index 51676be1..ce94a456 100644 --- a/apps/dottie/dottie/tasks.py +++ b/apps/dottie/dottie/tasks.py @@ -17,6 +17,9 @@ observations' recorded tool_calls). * ``file_ops`` — write a derived file under the sandbox scratch dir (its cwd), read it back, and report a content digest; verify by re-deriving the exact expected bytes. + Both line-ending forms of those bytes are accepted (LF, and the CRLF a + Windows text-mode write produces), so the score does not depend on the + OS the sandbox runs on. HONEST LIMIT: the sandbox scratch dir is ephemeral and inaccessible to the parent after the run, so the verifier proves the derived CONTENT (digest), not the write syscall itself. @@ -115,6 +118,10 @@ class VerifiedTask: tool_names: tuple[str, ...] = () # display signatures, e.g. "part_lookup(part_id)" tool_sources: dict[str, str] = field(default_factory=dict) expected: str = "" # canonical answer token (provider-side truth) + # Other tokens the verifier also accepts: the SAME computed truth in another encoding + # (file_ops: the digest of the expected bytes with CRLF line endings). The no-leakage + # guard checks these too, so none of them can be scored off the prompt either. + alt_expected: tuple[str, ...] = () grading: str = "binary" # "binary" | "graded" verifier_note: str = "" @@ -127,18 +134,24 @@ def verifier_detail(self) -> dict[str, Any]: "family": self.family_id, "seed": self.seed, "expected": self.expected, + "alt_expected": list(self.alt_expected), "grading": self.grading, "note": self.verifier_note, } def _binary_token_verify( - expected: str, *, ignore_case: bool = False + expected: str, *, ignore_case: bool = False, alternatives: Sequence[str] = () ) -> Callable[[str, Sequence[Any]], float]: + tokens = (expected, *alternatives) + def verify(final_text: str, observations: Sequence[Any]) -> float: return ( 1.0 - if answer_token_present(expected, final_text, ignore_case=ignore_case) + if any( + answer_token_present(t, final_text, ignore_case=ignore_case) + for t in tokens + ) else 0.0 ) @@ -284,6 +297,19 @@ def _build_file_ops(rng: random.Random, seed: int) -> VerifiedTask: # Exact expected file bytes, re-derived provider-side (the digest verifies the content). content = "\n".join(line.upper() for line in lines) + "\n" digest12 = hashlib.sha256(content.encode("utf-8")).hexdigest()[:12] + # The same content as a Windows text-mode write stores it. open(..., "w") there turns + # every "\n" into "\r\n", so a model that truthfully hashes the bytes it wrote reports + # this digest, and an LF-only verifier scored the family on the host OS instead of on + # the content. Measured 2026-09-27 (local-LLM A/B on this Windows box): 0 of 50 + # file_ops finals across five arms carried the LF digest; accepting this one moved + # qwen3:8b from 0/10 to 6/10. It is reachable only from the exact expected lines, so it + # credits no wrong content. Accepting it here, rather than forcing newline= inside the + # shared factory sandbox, keeps the prompt byte-identical and holds however the model + # opens the file; a sandbox shim would also change what "w+"/"r+" reads return, for + # every consumer of that sandbox, not just this family. + digest12_crlf = hashlib.sha256( + content.replace("\n", "\r\n").encode("utf-8") + ).hexdigest()[:12] doc = "\n".join(lines) prompt = ( "In the sandbox, take these lines:\n" @@ -299,12 +325,17 @@ def _build_file_ops(rng: random.Random, seed: int) -> VerifiedTask: seed=seed, prompt=prompt, expected=digest12, - verify_fn=_binary_token_verify(digest12, ignore_case=True), + alt_expected=(digest12_crlf,), + verify_fn=_binary_token_verify( + digest12, ignore_case=True, alternatives=(digest12_crlf,) + ), verifier_note=( "binary: sha256[:12] of the re-derived expected file bytes must appear in " - "the FINAL. Limit (honest): proves the derived content, not the write " - "syscall — the sandbox scratch dir is destroyed before the parent could " - "inspect it." + "the FINAL; the digest of the same bytes with CRLF line endings (what a " + "Windows text-mode write stores) is accepted too, so the score does not " + "depend on the sandbox host's OS. Limit (honest): proves the derived " + "content, not the write syscall — the sandbox scratch dir is destroyed " + "before the parent could inspect it." ), ) @@ -377,7 +408,10 @@ def build(self, family: str, seed: int) -> VerifiedTask: # str seeds hash via sha512 in random.Random — stable across runs and processes. rng = random.Random(f"dottie-task:{family}:{seed}:{attempt}") task = builder(rng, seed) - if not answer_token_present(task.expected, task.prompt, ignore_case=True): + if not any( + answer_token_present(token, task.prompt, ignore_case=True) + for token in (task.expected, *task.alt_expected) + ): return task raise TaskBuildError( # pragma: no cover - needs 8 consecutive leak coincidences f"could not build a leak-free {family!r} task for seed {seed} in " diff --git a/apps/dottie/research_orchestration/crontab b/apps/dottie/research_orchestration/crontab index ab65a4eb..68a2d6c2 100644 --- a/apps/dottie/research_orchestration/crontab +++ b/apps/dottie/research_orchestration/crontab @@ -10,7 +10,7 @@ DOTTIE_ROOT=/home/user/workspace/dottie DOTTIE_OLLAMA_URL=http://localhost:11434 -DOTTIE_OLLAMA_MODEL=qwen3:32b +DOTTIE_OLLAMA_MODEL=qwen3:8b SHELL=/bin/bash # 1. Ideation — daily at midnight. Grounds N hypotheses in the current real baseline + dead ends. diff --git a/apps/dottie/tests/test_policy.py b/apps/dottie/tests/test_policy.py index c8ed76ae..15a3d7ce 100644 --- a/apps/dottie/tests/test_policy.py +++ b/apps/dottie/tests/test_policy.py @@ -3,10 +3,16 @@ from __future__ import annotations +from pathlib import Path + +import httpx import pytest +import dottie.policy as policy_mod from dottie import resolve from dottie.policy import ( + DEFAULT_OLLAMA_MODEL, + DEFAULT_OLLAMA_NUM_PREDICT, AvaPolicy, DottiePolicyUnavailable, EchoPolicy, @@ -87,6 +93,135 @@ def test_ollama_env_config(monkeypatch): assert p.model == "some-model:7b" +# -- OllamaPolicy request options (fake transport, no network) ------------------------ + +_OPTION_ENV = ( + "DOTTIE_OLLAMA_MODEL", + "DOTTIE_OLLAMA_NUM_PREDICT", + "DOTTIE_OLLAMA_TEMPERATURE", + "DOTTIE_OLLAMA_NUM_CTX", + "DOTTIE_OLLAMA_NUM_GPU", + "DOTTIE_OLLAMA_KEEP_ALIVE", + "DOTTIE_OLLAMA_THINK", +) + + +class _FakeOllama: + """Stands in for ``httpx.post``: records every /api/chat payload, answers 200.""" + + def __init__(self) -> None: + self.payloads: list[dict] = [] + + def __call__(self, url, json=None, timeout=None): + self.payloads.append(json) + return httpx.Response( + 200, + json={"message": {"role": "assistant", "content": "FINAL: ok"}}, + request=httpx.Request("POST", url), + ) + + @property + def options(self) -> dict: + return self.payloads[-1]["options"] + + +@pytest.fixture() +def fake_ollama(monkeypatch): + for var in _OPTION_ENV: + monkeypatch.delenv(var, raising=False) + fake = _FakeOllama() + monkeypatch.setattr(policy_mod.httpx, "post", fake) + return fake + + +def test_ollama_default_model_is_the_pulled_qwen3_8b(fake_ollama): + # qwen3:32b is not pulled on the 4080 box (`ollama list`, 2026-09-27: qwen3:8b only); the + # research loop already pins qwen3:8b through its env, so only the bare default refused. + assert DEFAULT_OLLAMA_MODEL == "qwen3:8b" + p = OllamaPolicy(base_url="http://x") + assert p.model == "qwen3:8b" + p("<|user|>\nhello") + assert fake_ollama.payloads[-1]["model"] == "qwen3:8b" + + +def test_compose_default_model_matches_the_policy_default(): + app_root = Path(policy_mod.__file__).resolve().parent.parent + compose = (app_root / "docker-compose.dottie.yml").read_text(encoding="utf-8") + line = f"DOTTIE_OLLAMA_MODEL: ${{DOTTIE_OLLAMA_MODEL:-{DEFAULT_OLLAMA_MODEL}}}" + assert line in compose + + +def test_ollama_caps_every_generation_by_default(fake_ollama): + """No cap meant a degenerate generation ran until the server stopped it, and + llama-server's --context-shift keeps sliding the window instead of stopping. Measured + 2026-09-27: one call ran 708 s to 81,920 tokens (done_reason=length).""" + p = OllamaPolicy(base_url="http://x", model="m") + assert p("<|user|>\nhello") == "FINAL: ok" # the CodeAct path + assert fake_ollama.options["num_predict"] == DEFAULT_OLLAMA_NUM_PREDICT == 2048 + p.complete("hi", temperature=0.9) # the research path, per-call temperature + assert fake_ollama.options["num_predict"] == 2048 + # Knobs left unset keep today's request: the 0.2 temperature, no num_ctx. + p("<|user|>\nhello") + assert fake_ollama.options["temperature"] == 0.2 + assert "num_ctx" not in fake_ollama.options + + +def test_ollama_num_predict_env_and_arg(fake_ollama, monkeypatch): + monkeypatch.setenv("DOTTIE_OLLAMA_NUM_PREDICT", "512") + OllamaPolicy(base_url="http://x", model="m").complete("hi") + assert fake_ollama.options["num_predict"] == 512 + # An explicit constructor value beats the env. + OllamaPolicy(base_url="http://x", model="m", num_predict=64).complete("hi") + assert fake_ollama.options["num_predict"] == 64 + # 0 or negative = no cap: the key is omitted (Ollama's own default applies). + for off in ("0", "-1"): + monkeypatch.setenv("DOTTIE_OLLAMA_NUM_PREDICT", off) + p = OllamaPolicy(base_url="http://x", model="m") + assert p.num_predict is None + p.complete("hi") + assert "num_predict" not in fake_ollama.options + # Blank is unset, so the default cap applies. + monkeypatch.setenv("DOTTIE_OLLAMA_NUM_PREDICT", " ") + OllamaPolicy(base_url="http://x", model="m").complete("hi") + assert fake_ollama.options["num_predict"] == 2048 + # A typo refuses loudly instead of silently running uncapped. + monkeypatch.setenv("DOTTIE_OLLAMA_NUM_PREDICT", "lots") + with pytest.raises(ValueError, match="DOTTIE_OLLAMA_NUM_PREDICT"): + OllamaPolicy(base_url="http://x", model="m") + + +def test_ollama_temperature_env_precedence(fake_ollama, monkeypatch): + monkeypatch.setenv("DOTTIE_OLLAMA_TEMPERATURE", "0.7") + p = OllamaPolicy(base_url="http://x", model="m") + p("<|user|>\nhello") + assert fake_ollama.options["temperature"] == 0.7 + p.complete("hi") + assert fake_ollama.options["temperature"] == 0.7 + # Per-call temperature (the research stages set it) still wins over the env... + p.complete("hi", temperature=1.0) + assert fake_ollama.options["temperature"] == 1.0 + # ...and so does an explicit constructor value. + OllamaPolicy(base_url="http://x", model="m", temperature=0.1).complete("hi") + assert fake_ollama.options["temperature"] == 0.1 + monkeypatch.setenv("DOTTIE_OLLAMA_TEMPERATURE", "warm") + with pytest.raises(ValueError, match="DOTTIE_OLLAMA_TEMPERATURE"): + OllamaPolicy(base_url="http://x", model="m") + + +def test_ollama_num_ctx_env(fake_ollama, monkeypatch): + monkeypatch.setenv("DOTTIE_OLLAMA_NUM_CTX", "8192") + OllamaPolicy(base_url="http://x", model="m").complete("hi") + assert fake_ollama.options["num_ctx"] == 8192 + OllamaPolicy(base_url="http://x", model="m", num_ctx=4096).complete("hi") + assert fake_ollama.options["num_ctx"] == 4096 + monkeypatch.setenv("DOTTIE_OLLAMA_NUM_CTX", "0") + OllamaPolicy(base_url="http://x", model="m").complete("hi") + assert "num_ctx" not in fake_ollama.options + monkeypatch.setenv("DOTTIE_OLLAMA_NUM_CTX", "8k") + with pytest.raises(ValueError, match="DOTTIE_OLLAMA_NUM_CTX"): + OllamaPolicy(base_url="http://x", model="m") + + # -- AvaPolicy ---------------------------------------------------------------------- diff --git a/apps/dottie/tests/test_tasks.py b/apps/dottie/tests/test_tasks.py index cd5a5f5d..d99307be 100644 --- a/apps/dottie/tests/test_tasks.py +++ b/apps/dottie/tests/test_tasks.py @@ -34,9 +34,11 @@ def test_no_answer_leakage_self_check(family): """The scoring token must never appear in the prompt — so echoing the prompt can't score.""" for seed in LEAK_SEEDS: task = provider.build(family, seed) - assert not answer_token_present(task.expected, task.prompt, ignore_case=True), ( - f"{family} seed {seed} leaked expected {task.expected!r} into its prompt" - ) + # Every token the verifier accepts is guarded, not just the canonical one. + for token in (task.expected, *task.alt_expected): + assert not answer_token_present(token, task.prompt, ignore_case=True), ( + f"{family} seed {seed} leaked accepted token {token!r} into its prompt" + ) # The exact guarantee the echo e2e test relies on: grading the prompt itself scores 0. assert task.verify(task.prompt, []) == 0.0 @@ -96,6 +98,39 @@ def test_file_ops_expected_digest_rederivable_from_prompt(): assert t.verify(f"Digest prefix: {corrupted}", []) == 0.0 +def _sha12(text: str) -> str: + return hashlib.sha256(text.encode("utf-8")).hexdigest()[:12] + + +def _file_ops_content(t) -> str: + """The expected file content, re-derived from the prompt per its stated spec.""" + body = t.prompt.split("take these lines:\n", 1)[1].split("\nWrite them", 1)[0] + return "\n".join(line.upper() for line in body.splitlines()) + "\n" + + +@pytest.mark.parametrize("seed", LEAK_SEEDS) +def test_file_ops_scores_windows_text_mode_crlf_bytes(seed): + """A Windows text-mode write (``open('report.txt', 'w')``) stores every '\\n' as + '\\r\\n', so a model that hashes the bytes it really wrote reports the digest of the + CRLF form. The 2026-09-27 local-LLM A/B (Windows host) scored file_ops 0/10 on all five + arms because of this; no final in 50 carried the LF digest. The verifier must credit + the correct content whichever line ending the sandbox's OS gave it.""" + t = provider.build("file_ops", seed) + content = _file_ops_content(t) + lf, crlf = _sha12(content), _sha12(content.replace("\n", "\r\n")) + assert lf == t.expected and crlf != lf + assert t.verify(f"Digest prefix: {crlf}", []) == 1.0 + assert t.verify(f"Digest prefix: {crlf.upper()}", []) == 1.0 # hex case-insensitive + # The LF (POSIX) form still scores. + assert t.verify(f"Digest prefix: {lf}", []) == 1.0 + # CRLF bytes of the WRONG content still score 0: no trailing newline, not uppercased. + for wrong in (content.rstrip("\n"), content.lower()): + wrong_crlf = _sha12(wrong.replace("\n", "\r\n")) + assert t.verify(f"Digest prefix: {wrong_crlf}", []) == 0.0 + corrupted = ("0" if crlf[0] != "0" else "1") + crlf[1:] + assert t.verify(f"Digest prefix: {corrupted}", []) == 0.0 + + def test_constraint_verify_graded_and_token_gated(): t = provider.build("constraint", 9) m = re.search(r"between (\d+) and (\d+) words", t.prompt) diff --git a/apps/dottie/tests/test_verified_engine.py b/apps/dottie/tests/test_verified_engine.py index b68600ea..22d20e8e 100644 --- a/apps/dottie/tests/test_verified_engine.py +++ b/apps/dottie/tests/test_verified_engine.py @@ -12,6 +12,7 @@ from __future__ import annotations +import os import re import pytest @@ -127,6 +128,24 @@ def code_for(self, transcript: str) -> str: ) +class _FileOpsTextModeSolver(_ScriptedSolver): + """The write a model naturally emits: text mode, no ``newline=`` argument. On Windows + the file bytes come out CRLF, on Linux LF; both are the requested content.""" + + def code_for(self, transcript: str) -> str: + body = transcript.split("take these lines:\n", 1)[1].split("\nWrite them", 1)[0] + return ( + f"lines = {body.splitlines()!r}\n" + "content = '\\n'.join(l.upper() for l in lines) + '\\n'\n" + "with open('report.txt', 'w') as f:\n" + " f.write(content)\n" + "import hashlib\n" + "with open('report.txt', 'rb') as f:\n" + " data = f.read()\n" + "hashlib.sha256(data).hexdigest()[:12]" + ) + + def _run_scripted(engine, monkeypatch, family: str, seed: int, solver_cls): monkeypatch.setattr(engine_mod, "get_policy", lambda backend, **kw: solver_cls()) return engine.run_task(family=family, seed=seed, backend="scripted") @@ -160,6 +179,22 @@ def test_scripted_file_ops_solver_writes_in_real_scratch(engine, monkeypatch): assert rec["reward_components"]["r_task"] == 1.0 +def test_scripted_file_ops_text_mode_write_scores(engine, monkeypatch): + """The same correct content written in text mode must score the same on Windows and + Linux. Before the fix this was 0.0 on Windows: the sandbox wrote CRLF, the solver + truthfully hashed those bytes, and the verifier only knew the LF digest.""" + rec = _run_scripted(engine, monkeypatch, "file_ops", 2, _FileOpsTextModeSolver) + assert rec["steps"][0]["ok"] is True, rec["steps"][0]["error"] + assert rec["reward_components"]["r_task"] == 1.0 + detail = rec["verified_task"] + if os.linesep == "\r\n": + # The CRLF path really ran here: the LF digest is not in the FINAL. + assert detail["expected"] not in rec["final"] + assert detail["alt_expected"][0] in rec["final"] + else: + assert detail["expected"] in rec["final"] + + def test_scripted_solver_fails_honestly_on_wrong_computation(engine, monkeypatch): """Same scripted machinery with a wrong program -> the verifier scores 0.0 (it really discriminates; passing is not an artifact of the harness)."""