Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions apps/dottie/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ See [`DOTTIE_PRIME_SOTA.md`](./DOTTIE_PRIME_SOTA.md) for the full prime → Dott

## Honest capability statement (read first)

- **Ollama is the working brain today.** Only backend that does useful work is `ollama` (default `qwen3:32b`) served locally.
- **Ollama is the working brain today.** Only backend that does useful work is `ollama` (default `qwen3:8b`) served locally.
- **Ava is the trainee.** `ava` backend decodes from real smoke-scale checkpoint (~14M nano). Zero task capability today, emits noise — honestly. Exists so flywheel has a trainee and serving path is built for day a capable ckpt exists.
- **Echo is plumbing.** Deterministic CI harness (`plumbing_only=True`).
- **Anti-fabrication everywhere.** Unreachable Ollama, missing ckpt, missing torch → Dottie refuses with true reason (`DottiePolicyUnavailable` / 503). Every metric computed from real inputs; `r_task` for free-form is `null` (no verifier), never invented. Verified tasks (`compute`, `extract`, `tool_chain`, `file_ops`, `constraint`) have deterministic verifier from same values rendered into prompt — automated no-leakage check enforces it.
Expand All @@ -41,7 +41,7 @@ See [`DOTTIE_PRIME_SOTA.md`](./DOTTIE_PRIME_SOTA.md) for the full prime → Dott
rlm("...") ──▶│ + RLM Runner (rlm.py) ──▶ Harness v2 (harness_continual.py) │
│ │ ContinualHarness refined via evidence, snapshots, rollback │
│ │ Sessions daemon Registry + inbox messaging + goals + heartbeat │
│ ├─ OllamaPolicy (qwen3:32b) ─ brain │
│ ├─ OllamaPolicy (qwen3:8b) ─ brain │
│ ├─ AvaPolicy (TorchModelPolicy+ckpt) ─ trainee │
│ └─ EchoPolicy (deterministic CI) │
│ │ │
Expand All @@ -62,7 +62,7 @@ See [`DOTTIE_PRIME_SOTA.md`](./DOTTIE_PRIME_SOTA.md) for the full prime → Dott
```bash
# 1. Ollama brain
ollama serve &
ollama pull qwen3:32b
ollama pull qwen3:8b

# 2. Install Dottie SOTA
pip install -e apps/dottie # provides dottie CLI
Expand Down
4 changes: 2 additions & 2 deletions apps/dottie/RUNBOOK_4080.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ with Ollama as the working brain and your fresh mini checkpoint as the trainee.

| backend | what a climb iteration will measure |
|---|---|
| `ollama` | Real task capability of your local model (e.g. `qwen3:32b`). Verified-task success rates here are the first REAL capability numbers in the ecosystem. |
| `ollama` | Real task capability of your local model (e.g. `qwen3:8b`). Verified-task success rates here are the first REAL capability numbers in the ecosystem. |
| `ava` | Your homegrown checkpoint. A smoke/mini-scale checkpoint has **zero task capability** — expect `success_rate 0.0`. That number is the honest baseline the flywheel exists to move. |
| `echo` | Deterministic plumbing (`plumbing_only`), never a capability measurement. |

Expand All @@ -31,7 +31,7 @@ pip install torch --index-url https://download.pytorch.org/whl/cu128 # 4080 CU

# Ollama brain
ollama serve & # if not already running
ollama pull qwen3:32b # or the model you prefer; set DOTTIE_OLLAMA_MODEL to match
ollama pull qwen3:8b # or the model you prefer; set DOTTIE_OLLAMA_MODEL to match
```

## 1. Point the trainee at your fresh mini checkpoint
Expand Down
12 changes: 10 additions & 2 deletions apps/dottie/RUNBOOK_RESEARCH_LOOP.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,9 +54,9 @@ pip install torch --index-url https://download.pytorch.org/whl/cu128 # 4080 CU

# Ollama is the brain for ideation + implementation
ollama serve & # if not already running
ollama pull qwen3:32b # or your preferred model
ollama pull qwen3:8b # or your preferred model
export DOTTIE_OLLAMA_URL=http://localhost:11434
export DOTTIE_OLLAMA_MODEL=qwen3:32b
export DOTTIE_OLLAMA_MODEL=qwen3:8b
```

## 1. One-time — seed the baseline
Expand Down Expand Up @@ -180,6 +180,14 @@ Everything is under `apps/dottie/data/research/`:
broken fix verbatim (observed: identical wrong output shape four attempts running). Prefer a
qwen3-class model (`qwen3:14b` fits a 12 GB card); the policy strips `<think>` blocks, and
`DOTTIE_OLLAMA_THINK=false` is available when latency matters more than fix quality.
- **A completion stops mid-JSON**: every call is capped at 2048 generated tokens
(`DOTTIE_OLLAMA_NUM_PREDICT`; thinking tokens count against it). On the 4080 box's ledger
(177 experiments, measured 2026-09-27) the longest stored implementation `code` is 4,369
characters, well inside the cap under `THINK=false`.
Raise it if a model genuinely needs more; `0` removes the cap, which brings back the
uncapped runaway (708 s / 81,920 tokens on one call, 2026-09-27).
`DOTTIE_OLLAMA_TEMPERATURE` and `DOTTIE_OLLAMA_NUM_CTX` are also read; the research stages
still set their own temperature per call.
- **Everything sits in `pending`**: `implement` hasn't run (or keeps hitting `failed_validation`);
check `logs/implement.log` for the failing validator level.
- **`train` says "no experiments ready for training"**: nothing has passed validation yet — run
Expand Down
4 changes: 2 additions & 2 deletions apps/dottie/docker-compose.dottie.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
# curl http://localhost:8100/status
#
# extra_hosts maps host.docker.internal to the host gateway so the Ollama server running on
# the user's box (qwen3:32b at :11434) is reachable from inside the container — the same
# the user's box (qwen3:8b at :11434) is reachable from inside the container — the same
# convention apps/ava-factory/docker-compose.yml uses for its ava-train service.

services:
Expand All @@ -19,7 +19,7 @@ services:
- "127.0.0.1:8100:8100"
environment:
DOTTIE_OLLAMA_URL: ${DOTTIE_OLLAMA_URL:-http://host.docker.internal:11434}
DOTTIE_OLLAMA_MODEL: ${DOTTIE_OLLAMA_MODEL:-qwen3:32b}
DOTTIE_OLLAMA_MODEL: ${DOTTIE_OLLAMA_MODEL:-qwen3:8b}
DOTTIE_DATA_DIR: /data
DOTTIE_WORKERS: ${DOTTIE_WORKERS:-2}
DOTTIE_QUEUE_MAX: ${DOTTIE_QUEUE_MAX:-32}
Expand Down
74 changes: 66 additions & 8 deletions apps/dottie/dottie/policy.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
``apps/ava-factory/ava/rl/codeact_loop.py``: ``transcript: str -> next assistant turn: str``.
Three backends:

* :class:`OllamaPolicy` — real HTTP calls to an Ollama server (the user's local qwen3:32b by
* :class:`OllamaPolicy` — real HTTP calls to an Ollama server (the user's local qwen3:8b by
default). This is the only backend with real task capability today.
* :class:`AvaPolicy` — the trainee: wraps the factory's real ``TorchModelPolicy`` over a
smoke-scale ava checkpoint. Zero capability today; exists for the training flywheel.
Expand All @@ -21,14 +21,29 @@
import os
import re
from pathlib import Path
from typing import Any
from typing import TYPE_CHECKING, Any

import httpx

from dottie import resolve

if TYPE_CHECKING:
from collections.abc import Callable

DEFAULT_OLLAMA_URL = "http://host.docker.internal:11434"
DEFAULT_OLLAMA_MODEL = "qwen3:32b"
# The model the 4080 box actually serves. qwen3:32b (the old default) is not pulled there, so
# a bare run refused while the research loop, which pins qwen3:8b through its env, worked.
DEFAULT_OLLAMA_MODEL = "qwen3:8b"
# Per-call generation cap, sent as options.num_predict. With no cap, a degenerate generation
# is bounded only by the server, and llama-server runs with --context-shift, so it slides the
# window and keeps going. Measured 2026-09-27 (local-LLM A/B, granite4.2:3b at T0.2, -c 8192):
# one call ran 708 s to 81,920 tokens, done_reason=length. On the research loop's CPU setting
# the same runaway would hold the call until the 1800 s read timeout. Thinking tokens count
# against the cap, so with DOTTIE_OLLAMA_THINK unset a long qwen3 thought can use it up; the
# research loop runs think=false. DOTTIE_OLLAMA_NUM_PREDICT overrides; 0 or negative sends
# no cap.
DEFAULT_OLLAMA_NUM_PREDICT = 2048
DEFAULT_OLLAMA_TEMPERATURE = 0.2

# Transcript markers — must match ava/rl/codeact_loop.py (and the factory datagen) exactly.
USER = "<|user|>"
Expand Down Expand Up @@ -89,6 +104,19 @@ def transcript_to_messages(transcript: str) -> list[dict[str, str]]:
return messages


def _env_number(name: str, cast: Callable[[str], Any], kind: str) -> Any:
"""Parse env var ``name``; unset or blank means None. A value that does not parse
raises ValueError naming the variable, so a typo refuses instead of quietly running
with the default (for the generation cap, that would mean running uncapped)."""
raw = os.environ.get(name)
if raw is None or raw.strip() == "":
return None
try:
return cast(raw.strip())
except ValueError as e:
raise ValueError(f"{name} must be {kind}, got {raw!r}") from e


def strip_think(text: str) -> str:
"""Remove closed ``<think>...</think>`` blocks (qwen3-style reasoning preamble).

Expand All @@ -102,9 +130,21 @@ class OllamaPolicy(PolicyProvider):
"""Real next-turn generation over HTTP against an Ollama server (``/api/chat``).

Base URL from ``DOTTIE_OLLAMA_URL`` (default ``http://host.docker.internal:11434``), model
from ``DOTTIE_OLLAMA_MODEL`` (default ``qwen3:32b``). Non-streaming; sensible timeouts
(connect fast-fails, generation may take minutes on a 32b local model). Unreachable server
or HTTP error => :class:`DottiePolicyUnavailable` with the true cause."""
from ``DOTTIE_OLLAMA_MODEL`` (default ``qwen3:8b``). Non-streaming; sensible timeouts
(connect fast-fails, generation may take minutes on a CPU-pinned local model). Unreachable
server or HTTP error => :class:`DottiePolicyUnavailable` with the true cause.

Request options, each an explicit argument > env > default:

* ``num_predict`` / ``DOTTIE_OLLAMA_NUM_PREDICT`` — generation cap, default 2048 on
every call; 0 or negative sends none.
* ``temperature`` / ``DOTTIE_OLLAMA_TEMPERATURE`` — default 0.2. A per-call
``complete(temperature=...)`` beats all three.
* ``num_ctx`` / ``DOTTIE_OLLAMA_NUM_CTX`` — context window; unset, 0 or negative sends
none, so the model's own default applies.

``DOTTIE_OLLAMA_NUM_GPU`` / ``_KEEP_ALIVE`` / ``_THINK`` / ``_READ_TIMEOUT_S`` are
documented where they are read."""

name = "ollama"

Expand All @@ -115,7 +155,9 @@ def __init__(
*,
connect_timeout_s: float = 5.0,
read_timeout_s: float | None = None,
temperature: float = 0.2,
temperature: float | None = None,
num_predict: int | None = None,
num_ctx: int | None = None,
) -> None:
self.base_url = (
base_url or os.environ.get("DOTTIE_OLLAMA_URL") or DEFAULT_OLLAMA_URL
Expand Down Expand Up @@ -143,7 +185,19 @@ def __init__(
write=30.0,
pool=connect_timeout_s,
)
self.temperature = float(temperature)
if temperature is None:
temperature = _env_number("DOTTIE_OLLAMA_TEMPERATURE", float, "a number")
self.temperature = float(
DEFAULT_OLLAMA_TEMPERATURE if temperature is None else temperature
)
if num_predict is None:
num_predict = _env_number("DOTTIE_OLLAMA_NUM_PREDICT", int, "an integer")
if num_predict is None:
num_predict = DEFAULT_OLLAMA_NUM_PREDICT
self.num_predict: int | None = num_predict if num_predict > 0 else None
if num_ctx is None:
num_ctx = _env_number("DOTTIE_OLLAMA_NUM_CTX", int, "an integer")
self.num_ctx: int | None = num_ctx if num_ctx and num_ctx > 0 else None

def __call__(self, transcript: str) -> str:
messages = [{"role": "system", "content": CODEACT_SYSTEM_PROMPT}]
Expand Down Expand Up @@ -174,6 +228,10 @@ def _chat(self, messages: list, *, temperature: float | None = None) -> str:
if temperature is None
else float(temperature)
}
if self.num_predict is not None:
options["num_predict"] = self.num_predict
if self.num_ctx is not None:
options["num_ctx"] = self.num_ctx
# DOTTIE_OLLAMA_NUM_GPU=0 pins inference to CPU (doctrine: the GPU belongs to
# model TRAINING; LLM inference and everything else run on CPU). Any integer is
# passed through as the number of offloaded layers.
Expand Down
2 changes: 1 addition & 1 deletion apps/dottie/dottie/status.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@
from dottie.engine import DottieEngine

CAPABILITY_NOTE = (
"Ollama (external local model, e.g. qwen3:32b) is the only backend with real task "
"Ollama (external local model, e.g. qwen3:8b) is the only backend with real task "
"capability today. The ava backend decodes from a smoke-scale checkpoint (~90 base + ~25 "
"agentic optimizer steps, capability_claim=none) and exists to close the training "
"flywheel, not to assist. The echo backend is a deterministic plumbing test."
Expand Down
48 changes: 41 additions & 7 deletions apps/dottie/dottie/tasks.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,9 @@
observations' recorded tool_calls).
* ``file_ops`` — write a derived file under the sandbox scratch dir (its cwd), read it back,
and report a content digest; verify by re-deriving the exact expected bytes.
Both line-ending forms of those bytes are accepted (LF, and the CRLF a
Windows text-mode write produces), so the score does not depend on the
OS the sandbox runs on.
HONEST LIMIT: the sandbox scratch dir is ephemeral and inaccessible to the
parent after the run, so the verifier proves the derived CONTENT (digest),
not the write syscall itself.
Expand Down Expand Up @@ -115,6 +118,10 @@ class VerifiedTask:
tool_names: tuple[str, ...] = () # display signatures, e.g. "part_lookup(part_id)"
tool_sources: dict[str, str] = field(default_factory=dict)
expected: str = "" # canonical answer token (provider-side truth)
# Other tokens the verifier also accepts: the SAME computed truth in another encoding
# (file_ops: the digest of the expected bytes with CRLF line endings). The no-leakage
# guard checks these too, so none of them can be scored off the prompt either.
alt_expected: tuple[str, ...] = ()
grading: str = "binary" # "binary" | "graded"
verifier_note: str = ""

Expand All @@ -127,18 +134,24 @@ def verifier_detail(self) -> dict[str, Any]:
"family": self.family_id,
"seed": self.seed,
"expected": self.expected,
"alt_expected": list(self.alt_expected),
"grading": self.grading,
"note": self.verifier_note,
}


def _binary_token_verify(
expected: str, *, ignore_case: bool = False
expected: str, *, ignore_case: bool = False, alternatives: Sequence[str] = ()
) -> Callable[[str, Sequence[Any]], float]:
tokens = (expected, *alternatives)

def verify(final_text: str, observations: Sequence[Any]) -> float:
return (
1.0
if answer_token_present(expected, final_text, ignore_case=ignore_case)
if any(
answer_token_present(t, final_text, ignore_case=ignore_case)
for t in tokens
)
else 0.0
)

Expand Down Expand Up @@ -284,6 +297,19 @@ def _build_file_ops(rng: random.Random, seed: int) -> VerifiedTask:
# Exact expected file bytes, re-derived provider-side (the digest verifies the content).
content = "\n".join(line.upper() for line in lines) + "\n"
digest12 = hashlib.sha256(content.encode("utf-8")).hexdigest()[:12]
# The same content as a Windows text-mode write stores it. open(..., "w") there turns
# every "\n" into "\r\n", so a model that truthfully hashes the bytes it wrote reports
# this digest, and an LF-only verifier scored the family on the host OS instead of on
# the content. Measured 2026-09-27 (local-LLM A/B on this Windows box): 0 of 50
# file_ops finals across five arms carried the LF digest; accepting this one moved
# qwen3:8b from 0/10 to 6/10. It is reachable only from the exact expected lines, so it
# credits no wrong content. Accepting it here, rather than forcing newline= inside the
# shared factory sandbox, keeps the prompt byte-identical and holds however the model
# opens the file; a sandbox shim would also change what "w+"/"r+" reads return, for
# every consumer of that sandbox, not just this family.
digest12_crlf = hashlib.sha256(
content.replace("\n", "\r\n").encode("utf-8")
).hexdigest()[:12]
doc = "\n".join(lines)
prompt = (
"In the sandbox, take these lines:\n"
Expand All @@ -299,12 +325,17 @@ def _build_file_ops(rng: random.Random, seed: int) -> VerifiedTask:
seed=seed,
prompt=prompt,
expected=digest12,
verify_fn=_binary_token_verify(digest12, ignore_case=True),
alt_expected=(digest12_crlf,),
verify_fn=_binary_token_verify(
digest12, ignore_case=True, alternatives=(digest12_crlf,)
),
verifier_note=(
"binary: sha256[:12] of the re-derived expected file bytes must appear in "
"the FINAL. Limit (honest): proves the derived content, not the write "
"syscall — the sandbox scratch dir is destroyed before the parent could "
"inspect it."
"the FINAL; the digest of the same bytes with CRLF line endings (what a "
"Windows text-mode write stores) is accepted too, so the score does not "
"depend on the sandbox host's OS. Limit (honest): proves the derived "
"content, not the write syscall — the sandbox scratch dir is destroyed "
"before the parent could inspect it."
),
)

Expand Down Expand Up @@ -377,7 +408,10 @@ def build(self, family: str, seed: int) -> VerifiedTask:
# str seeds hash via sha512 in random.Random — stable across runs and processes.
rng = random.Random(f"dottie-task:{family}:{seed}:{attempt}")
task = builder(rng, seed)
if not answer_token_present(task.expected, task.prompt, ignore_case=True):
if not any(
answer_token_present(token, task.prompt, ignore_case=True)
for token in (task.expected, *task.alt_expected)
):
return task
raise TaskBuildError( # pragma: no cover - needs 8 consecutive leak coincidences
f"could not build a leak-free {family!r} task for seed {seed} in "
Expand Down
2 changes: 1 addition & 1 deletion apps/dottie/research_orchestration/crontab
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@

DOTTIE_ROOT=/home/user/workspace/dottie
DOTTIE_OLLAMA_URL=http://localhost:11434
DOTTIE_OLLAMA_MODEL=qwen3:32b
DOTTIE_OLLAMA_MODEL=qwen3:8b
SHELL=/bin/bash

# 1. Ideation — daily at midnight. Grounds N hypotheses in the current real baseline + dead ends.
Expand Down
Loading
Loading