Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .github/workflows/release-readiness.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,10 @@ jobs:
test -z "$(git status --porcelain)"
- name: Verify Claude guardrails
run: bash tests/test-hooks.sh
- name: Verify pi extension guardrails
run: |
npm install -g @earendil-works/pi-coding-agent
bash tests/test-pi.sh
- name: Verify orchestration and clean installs
run: bash tests/test-orchestrator.sh
- name: Verify regression scoring
Expand Down
35 changes: 35 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,25 @@ cd autoresearch

Invoke via the `$autoresearch` mention syntax: `$autoresearch <subcommand> [flags]`.

### pi (extension)

```bash
pi install git:github.com/uditgoenka/autoresearch
```

Then enable the `pi-extension` package as an extension in `~/.pi/agent/settings.json`:

```json
{ "extensions": ["@uditgoenka/autoresearch-pi"] }
```

Or try it without installing: `pi -e ./pi-extension`. Invoke commands as pi
prompt templates: `/autoresearch`, `/autoresearch_debug`, `/autoresearch_fix`, …
The safety guardrails (scout/privacy/dangerous-cmd block, iteration context,
simplify gate, stop-notify) are ported from the Claude hooks to pi's extension
events. See [`pi-extension/README.md`](pi-extension/README.md) for the full
hook→event mapping and per-hook disable flags.

### Manual (any agent)

Copy the skill files into your agent's skill directory:
Expand All @@ -45,6 +64,9 @@ cp autoresearch/claude-plugin/commands/autoresearch.md .claude/commands/autorese

# Codex
cp -r autoresearch/plugins/autoresearch ~/.agents/plugins/autoresearch

# pi (extension)
cp -r autoresearch/pi-extension ~/.pi/agent/extensions/autoresearch-pi
```

---
Expand Down Expand Up @@ -298,6 +320,15 @@ iteration commit metric delta status description
- Plugin files: `plugins/autoresearch/` with `skills/`
- Command contracts live in each command file under `plugins/autoresearch/skills/autoresearch/`

### pi

- Commands are pi **prompt templates**: `/autoresearch`, `/autoresearch_debug`, `/autoresearch_fix`, … (underscore names)
- The dispatcher is a pi **skill**: `/skill:autoresearch`
- Interactive setup uses `ctx.ui` (confirm/input/select) via the extension when context is missing
- Extension files: `pi-extension/src/` (guardrails) + `pi-extension/skills/autoresearch/` + `pi-extension/prompts/`
- Safety guardrails are ported from the Claude hooks to pi extension events (`tool_call`, `before_agent_start`, `input`, `session_start`, `session_shutdown`)
- Per-hook disable flags: `AR_DISABLE_*` env vars (see `pi-extension/README.md`)

### Other Agents (OpenCode, Gemini CLI, etc.)

- Read this file for the command surface and configuration contract
Expand All @@ -319,6 +350,10 @@ autoresearch/
├── claude-plugin/ ← Claude Code distribution package
│ ├── skills/autoresearch/SKILL.md ← Main skill + references/
│ └── commands/autoresearch/ ← Subcommand registrations
├── pi-extension/ ← pi (coding agent) extension package
│ ├── src/ ← Guardrails ported from Claude hooks → pi events
│ ├── skills/autoresearch/SKILL.md ← Skill dispatcher + references/
│ └── prompts/ ← 14 command prompt templates
└── plugins/autoresearch/ ← Codex distribution package
└── skills/autoresearch/SKILL.md ← Codex skill router + references/
```
Expand Down
51 changes: 45 additions & 6 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Whether you're fixing a typo, adding examples, creating a new sub-command, or im

## Quick Start

Autoresearch is Markdown files that Claude Code, OpenCode, and Codex discover from `skills/` and `commands/` directories. No build step, no compilation — edit a `.md` file, invoke the skill, see your changes.
Autoresearch is Markdown files that Claude Code, OpenCode, Codex, and pi discover from `skills/` and `commands/` directories. No build step, no compilation — edit a `.md` file, invoke the skill, see your changes.

```bash
# 1. Clone the repo
Expand All @@ -15,6 +15,7 @@ cd autoresearch
./scripts/install.sh --claude --global # Claude Code
./scripts/install.sh --opencode --global # OpenCode
./scripts/install.sh --codex --global # Codex
./scripts/install.sh --pi --global # pi (coding agent)

# 3. Or symlink for live editing (recommended for development)
ln -s $(pwd)/.claude/skills/autoresearch ~/.claude/skills/autoresearch
Expand All @@ -27,9 +28,10 @@ ln -s $(pwd)/.claude/commands/autoresearch.md ~/.claude/commands/autoresearch.md
The canonical source is `.claude/`. After making changes, run the transform to sync all platforms:

```bash
./scripts/transform.sh # sync to OpenCode + Codex
./scripts/transform.sh # sync to OpenCode + Codex + pi
./scripts/transform.sh --opencode # OpenCode only
./scripts/transform.sh --codex # Codex only
./scripts/transform.sh --pi # pi only
```

## Repository Structure (v2.2.2)
Expand All @@ -45,10 +47,14 @@ autoresearch/
│ └── autoresearch/ ← 13 subcommand files (self-contained)
├── .opencode/ ← OpenCode port (generated via transform.sh)
├── .agents/ + plugins/ ← Codex port (generated via transform.sh)
├── pi-extension/ ← pi (coding agent) extension (generated via transform.sh)
│ ├── src/ ← Guardrails ported from Claude hooks → pi events
│ ├── skills/autoresearch/ ← Skill + references + scripts
│ └── prompts/ ← 14 command prompt templates
├── claude-plugin/ ← Distribution package (Claude Code plugin install)
├── scripts/
│ ├── install.sh ← Guided installer (3 platforms)
│ ├── transform.sh ← .claude/ → .opencode/ + .agents/ sync
│ ├── install.sh ← Guided installer (4 platforms)
│ ├── transform.sh ← .claude/ → .opencode/ + .agents/ + pi-extension/ sync
│ ├── release.sh ← Release automation
│ └── release.md ← Release checklist
├── guide/ ← Guides — one per command + advanced patterns
Expand All @@ -67,8 +73,9 @@ autoresearch/
| `references/security-checklist.md` | STRIDE + OWASP checklist (loaded by security command) | Adding security checks |
| `references/predict-personas.md` | 5 expert personas (loaded by predict command) | Adding/modifying personas |
| `references/reason-judge-protocol.md` | Adversarial refinement protocol (loaded by reason command) | Changing judge/critic behavior |
| `scripts/transform.sh` | Canonical transform (.claude/ → .opencode/ + .agents/ + claude-plugin/) | Adding new commands, reference files, or generated helper updates |
| `scripts/transform.sh` | Canonical transform (.claude/ → .opencode/ + .agents/ + claude-plugin/ + pi-extension/) | Adding new commands, reference files, or generated helper updates |
| `claude-plugin/` | Distribution package — synced from .claude/ during release | Don't edit directly — edit .claude/ |
| `pi-extension/` | pi extension — skills + prompts synced from .claude/ via transform.sh; guardrails in `src/` ported from Claude hooks | Edit `src/` for guardrail logic; skills/prompts are generated — edit .claude/ |

## What to Contribute

Expand Down Expand Up @@ -114,7 +121,7 @@ Only create a reference in `references/` if shared by multiple commands. Single-
### 4. Run transform + update docs

```bash
./scripts/transform.sh # sync to OpenCode + Codex
./scripts/transform.sh # sync to OpenCode + Codex + pi
```

Update: README.md (commands table), guide/ (new guide file), COMPARISON.md (subcommand count).
Expand Down Expand Up @@ -154,6 +161,7 @@ For maintainer workflows, the canonical checks are:
- `bash scripts/transform.sh` — regenerate platform distributions and bundled runtime helpers
- `bash tests/test-maintenance.sh` — transform idempotence and release-prep guards
- `bash tests/test-hooks.sh` — Claude hook contracts and fail-open behavior
- `bash tests/test-pi.sh` — pi extension guardrail contracts (ported hooks) and TypeScript type-check

## Release Process

Expand Down Expand Up @@ -218,3 +226,34 @@ echo "Exit code: $?"
# Full test suite
bash tests/test-hooks.sh
```

## pi Hook Development

The pi extension ports the Claude hooks to pi's extension event system. Guardrail logic lives in `pi-extension/src/hooks/*.ts` as pure functions (no stdin/stdout — pi passes events in-process); `pi-extension/src/guardrails.ts` wires them to pi events (`tool_call`, `before_agent_start`, `input`, `session_start`, `session_shutdown`). Shared helpers are in `pi-extension/src/lib/`.

### Adding a New pi Hook

1. Create `pi-extension/src/hooks/{name}.ts` exporting a pure function that returns a `{ block?, reason?, warning?, needsConfirm?, text? }` result.
2. Wire it to the matching pi event in `pi-extension/src/guardrails.ts`.
3. Guard with `isHookEnabled("{name}")` (honors `AR_DISABLE_{NAME}`).
4. Fail open — wrap logic in try/catch, never throw to the event handler.
5. Add cases to `tests/test-pi.sh`.
6. Run `bash tests/test-pi.sh` (and `tsc --strict` is covered by the suite's type-check step).

### pi Hook Rules

- **Fail-open:** wrap in try/catch; a guardrail malfunction never blocks work
- **No external deps:** Node.js builtins only (vendored `lib/ignore.ts`)
- **No stdin/stdout:** pi passes events in-process; return an object, don't `process.exit`
- **State:** use `loadSessionState()` / `saveSessionState()` (OS temp dir, keyed by cwd + session id)
- **Logs:** `~/.pi/agent/autoresearch/.logs/<projectHash>/hook-log.jsonl` (bounded metadata only — no paths, commands, or secrets). Do NOT write under `~/.pi/agent/hooks/` — pi renamed hooks→extensions and warns on a legacy `hooks/` dir.

### Testing pi Hooks

```bash
# Full test suite (53 cases: guardrail logic + tsc --strict type-check)
bash tests/test-pi.sh

# Type-check only (needs pi-coding-agent types resolvable)
cd pi-extension && npx -y -p typescript@5.6 tsc --noEmit --strict --moduleResolution Bundler --module ESNext --target ES2022 --skipLibCheck --esModuleInterop --lib ES2022 --types "[]" src/index.ts
```
29 changes: 26 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,13 +2,14 @@

# Autoresearch

**Turn [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), or [OpenAI Codex](https://developers.openai.com/codex) into a relentless improvement engine.**
**Turn [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [OpenAI Codex](https://developers.openai.com/codex), or [pi](https://github.com/earendil-works/pi-coding-agent) into a relentless improvement engine.**

Based on [Karpathy's autoresearch](https://github.com/karpathy/autoresearch) — constraint + mechanical metric + autonomous iteration = compounding gains.

[![Claude Code Skill](https://img.shields.io/badge/Claude_Code-Skill-blue?logo=anthropic&logoColor=white)](https://docs.anthropic.com/en/docs/claude-code)
[![OpenCode](https://img.shields.io/badge/OpenCode-Skill-purple)](https://opencode.ai)
[![Codex](https://img.shields.io/badge/Codex-Skill-green?logo=openai&logoColor=white)](https://developers.openai.com/codex)
[![pi](https://img.shields.io/badge/pi-Extension-teal)](https://github.com/earendil-works/pi-coding-agent)
[![Version](https://img.shields.io/badge/version-2.2.2-blue.svg)](https://github.com/uditgoenka/autoresearch/releases)
[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)

Expand All @@ -22,7 +23,7 @@ Based on [Karpathy's autoresearch](https://github.com/karpathy/autoresearch) —

*You don't need AGI. You need a goal, a metric, and a loop that never quits.*

**Supports Claude Code, OpenCode, and OpenAI Codex for the core skill, bundled runtime, installation, and verification surface. Hook guardrails are Claude Code-only.**
**Supports Claude Code, OpenCode, OpenAI Codex, and pi for the core skill, bundled runtime, installation, and verification surface. Hook guardrails ship for Claude Code (native hooks) and pi (extension events).**

> **v2.2.0 — Autonomous Orchestrator:** Type a plain-language goal to `/autoresearch` and it classifies your goal, derives a Success predicate, confirms it once, then loops across subcommands until done. No manual chaining required. `Metric:`/`Verify:` invocations run the classic loop unchanged. See [guide/autoresearch-orchestrator.md](guide/autoresearch-orchestrator.md).

Expand Down Expand Up @@ -116,7 +117,7 @@ Before looping, Claude performs a one-time setup:

## Hooks & Safety

Hooks are defense-in-depth guardrails, not a security sandbox. Claude Code ships the hook surface; OpenCode and Codex ship the core skill/runtime/install surface without hook parity.
Hooks are defense-in-depth guardrails, not a security sandbox. Claude Code ships the hook surface natively; pi ships the same guardrails as extension events (`tool_call`, `before_agent_start`, `input`, `session_start`, `session_shutdown`). OpenCode and Codex ship the core skill/runtime/install surface without hook parity.

### What's Protected

Expand Down Expand Up @@ -314,6 +315,28 @@ cp -r autoresearch/.agents/skills/autoresearch ~/.codex/skills/autoresearch
> Invoke via `$autoresearch` mention syntax. Subcommands are keywords: `$autoresearch plan`, `$autoresearch debug`, `$autoresearch evals`, etc.
> The installed Codex package includes the bundled orchestrator and regression helpers under `plugins/autoresearch/skills/autoresearch/` and `.agents/skills/autoresearch/`.

### pi Quick Start

**Option A — Guided installer (recommended):**
```bash
git clone https://github.com/uditgoenka/autoresearch.git
cd autoresearch
./scripts/install.sh --pi --global
```

**Option B — pi package install:**
```bash
pi install git:github.com/uditgoenka/autoresearch
```

**Option C — Try without installing:**
```bash
pi -e ./pi-extension
```

> Invoke commands as pi prompt templates: `/autoresearch`, `/autoresearch_debug`, `/autoresearch_fix`, … (underscore names). The dispatcher is a skill: `/skill:autoresearch`.
> The pi extension ports the Claude hook guardrails (scout/privacy/dangerous-cmd block, iteration context, simplify gate, stop-notify) to pi extension events. See [`pi-extension/README.md`](pi-extension/README.md) for the full hook→event mapping and `AR_DISABLE_*` flags.

### Run It

```
Expand Down
Loading