Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 21 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,28 @@
# Claude Configuration
# =============================================================================
# Claude Code manages its own API configuration.
# Configure with: claude config (inside container)
# Configure with: claude config

# Optional: Claude model to use
# Options: claude-sonnet-4-5-20250929, claude-opus-4-20250514
LLM_MODEL=claude-sonnet-4-5-20250929

# =============================================================================
# Modernized legacy PentestGPT (pentestgpt-legacy) — provider API keys
# =============================================================================
# Set keys only for the providers you intend to use. Run `pentestgpt-legacy
# --list-models` to see which providers are configured, and `--smoke-test` to
# verify each model actually responds.

OPENAI_API_KEY=
ANTHROPIC_API_KEY=
GEMINI_API_KEY= # or GOOGLE_API_KEY
DEEPSEEK_API_KEY=
GROK_API_KEY= # xAI (or XAI_API_KEY)
QWEN_API_KEY= # Alibaba DashScope (or DASHSCOPE_API_KEY)
KIMI_API_KEY= # Moonshot (or MOONSHOT_API_KEY)

# Optional base-URL overrides (defaults are built in per provider)
# OLLAMA_BASE_URL=http://localhost:11434/v1
# OPENAI_BASE_URL=
# DEEPSEEK_BASE_URL=https://api.deepseek.com
3 changes: 0 additions & 3 deletions .gitmodules

This file was deleted.

118 changes: 43 additions & 75 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,16 +4,16 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co

## Project Overview

PentestGPT is an AI-powered autonomous penetration testing agent with a terminal user interface (TUI). It uses an agentic pipeline to solve CTF challenges, Hack The Box machines, and authorized security assessments.
PentestGPT is an AI-powered autonomous penetration testing agent. It uses an agentic pipeline to solve CTF challenges, Hack The Box machines, and authorized security assessments. Output is raw streaming to stdout.

**Published at USENIX Security 2024**: [Paper](https://www.usenix.org/conference/usenixsecurity24/presentation/deng)

**Stack:** Python 3.12+, uv, Docker (Ubuntu 24.04), Textual (TUI), Rich (CLI), Agent SDK
**Stack:** Python 3.12+, uv, Agent SDK

## Common Commands

```bash
# Development
# Setup
uv sync # Install dependencies
uv run pentestgpt --target X # Run locally

Expand All @@ -27,41 +27,32 @@ make lint # Run ruff linter
make format # Format code with ruff
make typecheck # Run mypy type checking
make check # All checks (lint + typecheck)

# Docker Workflow
make install # Build Docker image
make connect # Connect to container (main usage)
make stop # Stop container
make clean-docker # Remove everything including config
```

## Architecture

### Entry Point
- `pentestgpt/interface/main.py` - CLI entry, argument parsing, mode selection
- Command: `pentestgpt --target <IP/URL> [--instruction "hint"] [--non-interactive] [--raw] [--debug]`
- `pentestgpt/interface/main.py` - CLI entry, argument parsing, raw streaming output
- Command: `pentestgpt --target <IP/URL> [--max-iterations N] [--instruction "hint"] [--debug]`

### Core Layer (`pentestgpt/core/`)
- **agent.py** - `PentestAgent`: Wraps the LLM agent, handles flag detection, logs to `/workspace/pentestgpt-debug.log`
- **pipeline.py** - `PipelineOrchestrator`: Runs an iteration loop, each iteration with a fresh backend + controller. Data classes: `IterationResult`, `LoopResult`. The agent writes a context file (`pentestgpt_context.md`) as it works; the orchestrator reads it after each iteration and feeds it into the next. Loop terminates on flag capture, error, or max iterations.
- **backend.py** - `AgentBackend` interface + `ClaudeCodeBackend` implementation (framework-agnostic design)
- **controller.py** - `AgentController`: 5-state lifecycle (IDLE->RUNNING->PAUSED->COMPLETED->ERROR), pause/resume at message boundaries
- **events.py** - `EventBus`: Singleton pub/sub for TUI-agent decoupling (STATE_CHANGED, MESSAGE, TOOL, FLAG_FOUND events)
- **controller.py** - `AgentController`: 5-state lifecycle (IDLE->RUNNING->PAUSED->COMPLETED->ERROR), pause/resume at message boundaries. Used per-iteration by the pipeline orchestrator
- **events.py** - `EventBus`: Singleton pub/sub for agent-output decoupling (STATE_CHANGED, MESSAGE, TOOL, FLAG_FOUND events)
- **session.py** - `SessionStore`: File-based persistence in `~/.pentestgpt/sessions/`, supports session resumption
- **config.py** - Pydantic settings with `.env` file support

### Interface Layer (`pentestgpt/interface/`)
- **tui.py** - Textual TUI app with real-time activity feed, F1 help, Ctrl+P pause, Ctrl+Q quit
- **components/** - ActivityFeed, SplashScreen, tool-specific Renderers
- **config.py** - Pydantic settings with `.env` file support; includes `max_iterations` (default 10) and `context_file` (default `pentestgpt_context.md`)

### System Prompts (`pentestgpt/prompts/`)
- **pentesting.py** - `CTF_SYSTEM_PROMPT`: CTF methodology, flag formats, persistence directives
- **system_prompt.py** - Unified prompt builders: `build_system_prompt`, `build_first_task_prompt`, `build_continuation_task_prompt`. Shared fragments (`_IDENTITY`, `_TOOLS`, `_FLAG_PATTERNS`, `_PERSISTENCE`, `_FALLBACK_STRATEGIES`, `_CTF_CATEGORIES`, `_METHODOLOGY`, `_CONTEXT_PERSISTENCE`)

## Key Patterns

- **Event-Driven**: TUI subscribes to EventBus; agent emits events for state changes, messages, flags
- **Singletons**: `EventBus.get()`, `get_global_tracer()` for global access
- **Iteration Loop**: `PipelineOrchestrator` runs iterations in a loop. Each iteration gets a fresh `ClaudeCodeBackend` + `AgentController` (system prompt is set at connect time, so a new backend is needed per iteration). The agent maintains a context file; the orchestrator reads it after each iteration and injects it into the next iteration's task prompt. Falls back to truncated prior output if the context file is missing.
- **Event-Driven**: Raw mode subscribes to EventBus; agent emits events for state changes, messages, flags
- **Singletons**: `EventBus.get()` for global access
- **Abstract Backend**: `AgentBackend` interface allows swapping LLM backends
- **Flag Detection**: Regex patterns in agent.py match `flag{}`, `HTB{}`, `CTF{}`, 32-char hex
- **Flag Detection**: Regex patterns in controller.py match `flag{}`, `HTB{}`, `CTF{}`, 32-char hex

## Testing

Expand All @@ -73,67 +64,44 @@ uv run pytest tests/test_controller.py -v # Single file
uv run pytest tests/test_controller.py::test_name # Single test
```

## Docker Notes

- Non-root user: `pentester` with sudo
- Workdir: `/workspace` (mounted from `./workspace`)
- LLM config persisted in `claude-config` volume
- Pre-installed: nmap, netcat, curl, wget, git, ripgrep, tmux

## Legacy Version

The previous multi-LLM version (v0.15) is archived in `legacy/`. It supports:
- OpenAI (GPT-4o, o3, o4-mini)
- Google Gemini
- Deepseek
- Ollama (local LLMs)
- GPT4All

To develop on the legacy version:
```bash
cd legacy
pip install -e .
```

## Benchmark System

Use the standalone benchmark runner at `benchmark/standalone-xbow-benchmark-runner/`:

```bash
cd benchmark/standalone-xbow-benchmark-runner

python3 run_benchmarks.py --range 1-10 --pattern-flag # Run benchmarks 1-10
python3 run_benchmarks.py --all --pattern-flag # Run all 104 benchmarks
python3 run_benchmarks.py --retry-failed # Retry failed benchmarks
python3 run_benchmarks.py --dry-run --range 1-5 # Preview without executing
```

See `benchmark/standalone-xbow-benchmark-runner/README.md` for full documentation.

## Repository Structure

```
.
├── pentestgpt/ # Main package (agentic version)
│ ├── core/ # Agent, controller, events, session
│ ├── interface/ # TUI and CLI
│ ├── prompts/ # System prompts
│ ├── benchmark/ # Benchmark runner module
│ └── tools/ # Tool framework
├── benchmark/ # Benchmark suites
│ ├── xbow-validation-benchmarks/ # 104 XBOW benchmarks
│ └── standalone-xbow-benchmark-runner/ # Benchmark runner
├── tests/ # Test suite
├── workspace/ # Runtime workspace (Docker mount)
├── legacy/ # Archived v0.15 (multi-LLM)
├── Dockerfile # Ubuntu 24.04 container
├── docker-compose.yml # Container orchestration
├── pentestgpt/ # Autonomous agent (Claude-only, claude CLI backend)
│ ├── core/ # Pipeline, controller, events, session, backend
│ ├── interface/ # CLI entry point (raw streaming)
│ └── prompts/ # System prompts (system_prompt.py)
├── pentestgpt_legacy/ # Modernized legacy: interactive 3-session + PTT, multi-LLM
│ ├── llm/ # Native per-provider LLM layer (registry, factory, client, providers)
│ ├── utils/ # Orchestrator (pentest_gpt.py) + REPL helpers
│ └── prompts/ # Classic PTT/session prompts
├── tests/ # Test suite (tests/legacy/ covers pentestgpt_legacy)
└── Makefile # Development commands
```

### Modernized Legacy (`pentestgpt_legacy/`)

The classic USENIX-2024 human-in-the-loop tool, rebuilt on a native multi-provider LLM
layer. CLI: `pentestgpt-legacy` (`--list-models`, `--smoke-test`, `--reasoning-model`,
`--parsing-model`, `--base-url`).

- **llm/registry.py** — single source of truth for supported models (`ModelSpec`/`PROVIDERS`),
web-verified IDs. `--list-models` and the README table render from it.
- **llm/factory.py** — `get_client(model_name)` -> `LLMClient`; resolves provider, builds it.
- **llm/client.py** — `LLMClient` bridges async providers to the core's synchronous
`send_new_message`/`send_message` (drop-in for the old `LLMAPI`); holds per-conversation history.
- **llm/providers/** — `OpenAICompatibleProvider` (OpenAI + DeepSeek/Ollama/xAI/Qwen/Moonshot via
base_url; Responses-API fallback for `*-pro`/`*-codex`), `AnthropicProvider`, `GeminiProvider`.
- **llm/config.py** — pydantic-settings credentials (per-provider keys + base-url overrides).
- **smoke_test.py** — `--smoke-test` makes a real round-trip per configured model (acceptance gate).

Note: `make typecheck` is scoped to `pentestgpt/`; the new package is covered by ruff
(`make lint`) and `tests/legacy/`. Run `uv run mypy pentestgpt_legacy/llm/` for its typed core.

## Modification Requirements

When modifying code, ensure:
- Adherence to existing architecture and patterns
- Comprehensive tests for new features
- Ensure to run tests after changes, and do further updates to ensure code quality. Always keep the documentation up to date with any architectural changes. Also ensure all tests pass after modifications.
- Ensure to run tests after changes, and do further updates to ensure code quality. Always keep the documentation up to date with any architectural changes. Also ensure all tests pass after modifications.
23 changes: 10 additions & 13 deletions Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -45,21 +45,18 @@ RUN apt-get update && \
&& apt-get autoclean \
&& rm -rf /var/lib/apt/lists/*

# Install Node.js v20 (required for Claude Code CLI)
# Install Node.js v20 (required for Claude Code Router)
RUN curl -fsSL https://deb.nodesource.com/setup_20.x | bash - && \
apt-get install -y nodejs && \
rm -rf /var/lib/apt/lists/*

# Remove EXTERNALLY-MANAGED marker to allow pip/poetry in Docker
# Also remove system Python packages that conflict with Poetry dependencies
# Remove EXTERNALLY-MANAGED marker to allow pip installs in Docker
# Also remove system Python packages that conflict with dependencies
RUN rm -f /usr/lib/python3.*/EXTERNALLY-MANAGED && \
apt-get remove -y python3-cryptography && \
apt-get autoremove -y

# Install Claude Code CLI globally
RUN npm install -g @anthropic-ai/claude-code

# Install Claude Code Router globally (for OpenRouter support)
# Install Claude Code Router globally (for OpenRouter/local LLM support)
RUN npm install -g @musistudio/claude-code-router

# Create non-root user
Expand All @@ -75,9 +72,11 @@ RUN mkdir -p /workspace /app /home/pentester/.claude /home/pentester/.claude-cod
USER pentester
WORKDIR /app

# Install Poetry for Python dependency management
RUN curl -sSL https://install.python-poetry.org | python3 - && \
echo 'export PATH="/home/pentester/.local/bin:$PATH"' >> /home/pentester/.bashrc
# Install Claude Code CLI (native installer — no npm needed)
RUN curl -fsSL https://claude.ai/install.sh | bash

# Install uv for Python dependency management
RUN curl -LsSf https://astral.sh/uv/install.sh | sh

ENV PATH="/home/pentester/.local/bin:$PATH"

Expand All @@ -88,11 +87,9 @@ COPY --chown=pentester:pentester scripts/entrypoint.sh /home/pentester/entrypoin
COPY --chown=pentester:pentester scripts/ccr-config-template.json /app/scripts/ccr-config-template.json

# Install Python dependencies as root to system Python
# Allow pip to override system packages in Docker
ENV PIP_BREAK_SYSTEM_PACKAGES=1
USER root
RUN poetry config virtualenvs.create false && \
poetry install --only main && \
RUN /home/pentester/.local/bin/uv pip install --system /app && \
chmod +x /home/pentester/entrypoint.sh

# Switch back to pentester user for runtime
Expand Down
Loading