This document describes how Skill3 is structured and how data flows through it. For what it does, see SPEC.md; for the build order, see PLAN.md.
Skill3 is a single-process Java CLI. A learn invocation runs a linear pipeline;
each stage enriches or filters a list of Source objects until a SKILL.md is
synthesized, optionally verified, and vetted — on both its input corpus and its
output skill.
The guiding invariant: a generated skill is a post-cutoff delta, not a primer. The target model already knows the topic up to its knowledge cutoff; the pipeline gathers only what changed after that cutoff and compiles it into a skill. Discovery is topic-agnostic — the model itself plans the searches, so nothing is hardcoded per topic.
graph TD
CLI[Skill3App / picocli] --> Learn[LearnCommand]
Learn --> Cutoff[CutoffResolver]
Learn --> Chat[ChatModel:<br/>LocalLlmClient or AnthropicChatModel]
Learn --> Pipe[LearnPipeline + Options]
Pipe --> Plan[QueryPlanner → facet queries]
Plan --> Retrieve[RetrievalService]
Retrieve --> Brave[BraveSearchClient<br/>one search per query<br/>or FileCorpus --input-file]
Retrieve --> Fetch[HttpPageFetcher<br/>parallel, virtual threads<br/>or FileCorpus --input-file]
Retrieve --> Dates[DateExtractor]
Retrieve --> Authority[AuthorityScorer]
Retrieve --> Sources[(merged, de-duped<br/>List<Source>)]
Sources --> Ingest[IngestionPipeline]
Ingest --> Consensus[ConsensusValidator]
Ingest --> Fresh[FreshnessFilter<br/>cutoff ≤ published ≤ today]
Ingest --> Ranked[(ranked List<Source>)]
Ranked --> InVet[InputVetter<br/>redact secrets · scan · quarantine<br/>Phase 3a, pre-synthesis]
InVet --> Spector[SkillSpectorRunner<br/>--no-llm static scan]
InVet --> Synth[Synthesizer + ChatModel<br/>delta prompt]
Synth --> Verify[Verifier + ChatModel<br/>optional, --verify]
Verify --> Post[SkillMdPostProcessor<br/>format guarantees + footer]
Post --> Loop[SelfCorrectionLoop]
Loop --> Spector
Loop --> Out[SKILL.md]
Out --> Web[WebPreviewGenerator → index.html]
Out --> Gate{high-severity<br/>findings remain?}
Gate -->|yes, gate on| Fail[exit 3]
Every model-driven stage (QueryPlanner, Synthesizer, Verifier, and the
self-correction reviser) talks to the same injected ChatModel, so the choice of
provider (local Ollama, an OpenAI-compatible gateway, or Claude) applies uniformly.
| Package | Responsibility |
|---|---|
se.deversity.skill3 |
Skill3App entry point and LearnPipeline orchestration (with LearnPipeline.Options). |
se.deversity.skill3.cli |
SetupCommand, LearnCommand (picocli), Venv. |
se.deversity.skill3.model |
Source (mutable carrier), ContextBundle (immutable record), Cutoff — plain data. |
se.deversity.skill3.pipeline |
Discovery + ingestion: QueryPlanner, RetrievalService, BraveSearchClient, FileCorpus (offline --input-file, implements both SearchClient and PageFetcher), HttpPageFetcher/PageFetcher, DateExtractor, AuthorityScorer, ConsensusValidator, FreshnessFilter, IngestionPipeline, CutoffResolver, SearchClient. |
se.deversity.skill3.llm |
Synthesis: ChatModel (SAM); LocalLlmClient (OpenAI-compatible — local or hosted gateway via Bearer key); AnthropicChatModel (native anthropic-java SDK); Synthesizer, Verifier, SkillMdPostProcessor, NameSanitizer. |
se.deversity.skill3.skillspector |
SkillSpectorRunner (ProcessBuilder), SkillSpectorReport, Finding, InputVetter (pre-synthesis input scan + secret redaction), Reviser, SelfCorrectionLoop, SkillSpectorUnavailableException. |
se.deversity.skill3.console |
ProgressBar — neutral leaf for console progress output (drives the input-vetting progress bar). |
se.deversity.skill3.web |
WebPreviewGenerator. |
Every package carries a @NullMarked package-info.java (JSpecify). Layering is
enforced by an ArchUnit test: model depends on nothing internal, only the
Skill3App root touches cli, and the sub-packages stay acyclic.
The mermaid graph above is drawn by hand: it shows the pipeline as intended, including
stages that are conceptual rather than a single class. The SVGs below are the other half —
parsed straight out of src/main/java by code-karta,
so they show the code as written and cannot quietly drift from it.
Regenerate them with ./gradlew diagrams and commit the result with the change that moved
them. The task pins code-karta-cli from Maven Central, so it needs no local checkout, and
the renderer is deterministic — a re-run with unchanged source produces byte-identical SVGs,
which is what keeps them out of review diffs unless the structure really moved.
| Diagram | What it answers |
|---|---|
Data carriers ( model) |
What flows between stages: Source and its scoring fields, the immutable ContextBundle, Cutoff, and the RunManifest written as run.json. |
The provider seam ( llm) |
Why --llm-provider works: LocalLlmClient and AnthropicChatModel both implement the one ChatModel interface every model-driven stage takes. |
Discovery and ingestion ( pipeline) |
The collaborators behind discovery, and the SearchClient / PageFetcher seams that FileCorpus implements both of for offline runs. |
| Call graph ( pipeline, stitched across files) |
The real call sequence through discovery and ingestion, resolved across files by symbol solving. Large — open it directly rather than inline. |
There is deliberately no JPMS module diagram: Skill3 has no module-info.java, so it would
have nothing to show.
The Synthesizer system prompt instructs the model to include only what changed
after the cutoff and to treat fundamentals as already-known. The cutoff (and today's
date) are passed in the user prompt. This is the core premise — re-explaining what the
model already knows wastes the mechanism.
QueryPlanner asks the ChatModel to expand a topic into up to six distinct
post-cutoff facet queries. There is no per-topic logic: the same planner serves a
protocol, a person, or a release. It falls back to the bare topic if planning fails.
RetrievalService runs every query through BraveSearchClient, merges and
de-duplicates the result URLs, and fetches the pages concurrently on a
virtual-thread-per-task executor (blocking I/O, independent work); results are merged
on the caller thread, so no Source is shared between workers.
FileCorpus is a no-network alternative to Brave: the user pastes the relevant
documents into one curated file and passes --input-file. It implements both
discovery seams — SearchClient (returns every URL in the file) and PageFetcher
(replays each document's body as synthesized HTML) — so LearnCommand injects one
instance into both slots and the rest of the pipeline is identical to a live run.
Bodies may be plain text, Markdown, or HTML; the file is the curated result set, so
the whole corpus is used regardless of the model's planned queries. No Brave key or
network egress is needed for discovery in this mode.
CutoffResolver maps --target-model to a knowledge-cutoff month (overridable with
--cutoff-time); that date is also the lower bound of the Brave freshness window.
FreshnessFilter scores each source authority × recency (post-cutoff publication
earns full recency) and sorts best-first, so a post-cutoff authoritative source
(1.0 × 1.0) outranks pre-cutoff content (≤ 1.0 × 0.5). --strict-cutoff promotes
the soft floor to a hard filter. An upper bound (today) drops sources dated after
the run — Brave's freshness window is only a hint, and a future-dated page is a leak or
mis-dated, never valid evidence.
AuthorityScorer scores by host: configured authoritative hosts → 1.0, known
low-authority/blog hosts → 0.2, standard repos → 0.7, default 0.5. Authoritative
hosts are supplied per run via --authoritative (kept topic-agnostic — the operator
names the primary sources), so the official spec/site can outrank content farms.
- Target model (
--target-model) — only a cutoff lookup; never called. - Synthesis provider (
--llm-provider):local/openaiuseLocalLlmClient(OpenAI-compatible chat-completions;openaiadds a Bearer key);anthropicusesAnthropicChatModelover the native Anthropic SDK (no OpenAI shim). Opus 4.8 rejectstemperature/budget_tokens, so neither is sent on that path.--rich-contextwidens the per-source prompt budget for big-context models.
The LLM drafts; SkillMdPostProcessor + NameSanitizer guarantee format
compliance regardless of the model: exactly one YAML frontmatter block; a sanitized
name ([a-z0-9-], ≤64, reserved words anthropic/claude stripped); a non-empty,
tag-free, ≤1024-char description (degenerate model descriptions are rejected and
derived from the body); leaked scaffolding and secondary frontmatter removed; and a
Created with skill3 provenance footer stamped idempotently. It also recovers a skill
the model wrapped in a code fence after a chatty preamble.
The pipeline otherwise checks only format (post-processor) and safety
(SkillSpector) — never truth. The Verifier (--verify) re-grounds the draft
against the very sources it was built from in one model call: it removes claims the
sources do not support and demotes future-dated releases from "shipped" to "announced".
Its output is re-run through the post-processor. The gate is only worthwhile with a
capable model — a weak local model can rewrite rather than ground, so it is opt-in.
Both the input corpus and the output skill are scanned by the same
SkillSpectorRunner (--no-llm, static analysis only — no API key needed).
- Input (Phase 3a, before synthesis):
InputVetterfirst deterministically redacts secret-shaped tokens from every source in place — so a leaked key never reaches the model (or a hosted provider), even when the scanner is unavailable — then scans the assembled corpus (each source written assource-N.txt). Any source carrying a high-severity finding is quarantined: dropped from the set handed to the synthesizer, with a warning, so a poisoned page never shapes the skill (the finding is still recorded and still trips the gate — quarantine is mitigation, not amnesty; if every source is quarantined the run fails loudly). The synthesizer only ever sees the sanitized, surviving sources, and a liveProgressBarreports per-source progress. Lower-severity findings are reported and recorded inrun.json, not silently dropped; the synthesis prompt additionally fences all source text as untrusted DATA. - Output (after synthesis):
SelfCorrectionLoopscans the draft, feeds findings to aReviser(a local-LLM revision viaChatModel), and rescans — up to a bounded number of iterations. Residual findings are surfaced, never hidden.
Run gate. LearnCommand pools the high-severity findings (HIGH/CRITICAL, plus
SARIF error) from both scans via Result.blockingFindings(); if any remain, the skill
is still written but learn exits 3. --no-fail-on-findings disables the gate, and
advisory MEDIUM/LOW findings never gate. When SkillSpector is unavailable the scans are
skipped, so nothing is gated — absence of findings is never asserted, only observed.
"Clean" means safe and well-formed, not necessarily accurate — hence the separate accuracy gate above.
Network and model access sit behind interfaces (SearchClient, PageFetcher,
ChatModel), so the pure logic (query parsing, scoring, freshness, consensus, date
extraction, name sanitization, cutoff resolution, post-processing, verification wiring)
is unit-tested with fakes and HTML fixtures — no live network or model. Concurrency in
RetrievalService and SkillSpectorRunner is stress-tested with async-test-lib.
- Scraped page text and
--input-filecontent are untrusted data: both flow through the same path, are secret-redacted and SkillSpector-scanned before synthesis (input vetting), delimited in the synthesis and verification prompts with the model constrained to the supplied context, and the output is vetted by SkillSpector again. - External processes (
git, the venvskillspector) run viaProcessBuilderwith explicit argument lists — never a shell string. - Network egress is discovery (Brave + page scraping) and the synthesis provider:
loopback by default (
local), or a hosted endpoint when--llm-provider openaioranthropicis chosen explicitly. With--input-file, discovery makes no network calls at all. The Brave key and any provider key are treated as secrets (@AIPrivacy) and never logged.
- Missing/unknown
--target-modelwith no--cutoff-time→ clear error. - A hosted provider with no key (
openai/anthropic) → clear error before any run. - No usable sources, or everything filtered out under
--strict-cutoff→ loud error. - An unreadable or malformed
--input-file(no=== SOURCE ===blocks, or a source missing itsurl:header) → clear error before any model call. - Per-URL fetch failures (403/timeout/parse) are skipped; discovery is best-effort.
- Future-dated sources are dropped by the freshness upper bound.
- SkillSpector not installed → both input and output vetting skipped with a warning (skill still emitted; nothing is gated).
- A source carries a high-severity finding → it is quarantined (dropped before synthesis) with a warning; if every source is quarantined the run fails loudly.
- High-severity findings remain after vetting → skill is written, but
learnexits3(override with--no-fail-on-findings). - Synthesis/verification provider errors surface as
IOExceptionand fail the run.