A Hermes Agent plugin that
ports the useful parts of
Prime Agent's RLM model
onto Hermes' own tool stack: a persistent Python kernel per session,
subagents as function calls (blocking or handle-only), Python-backed
skills, crash-surviving checkpoints, and a continual harness that carries
lessons across sessions. Stdlib-only, no core patch — it survives every
hermes update.
Measured impact (details below): 40–48% less conversation context on follow-up questions over the same data, a 21× cheaper marginal question, 2.5× faster parallel delegation, 11.9× faster tool-free subagents — and with 0.3.0, subagent output stays out of the parent context entirely.
git clone https://github.com/jarnodevries-byte/hermes-rlm ~/.hermes/plugins/hermes-rlm
hermes plugins enable hermes-rlmThen restart your gateway (launchctl kickstart -k gui/$(id -u)/ai.hermes.gateway
or however you run it) — plugins load at startup. Stdlib-only: no pip installs.
Optional but recommended: add a routing line to agent.environment_hint in
~/.hermes/config.yaml telling the model to prefer rlm_exec for repeated
questions over the same dataset — measured to be the load-bearing adoption step.
The kernel executes model-generated Python with your own permissions. It is a durable control environment, not a security sandbox. Do not enable it on safe-mode profiles.
execute_code is stateless: every script starts empty, so a large dataset is
re-loaded for each follow-up question and every intermediate result must pass
through the conversation to survive. This plugin adds a stateful lane beside
it. execute_code stays the safe, ephemeral, secret-scrubbed default.
| Tool | Purpose |
|---|---|
rlm_exec |
Run Python in the session's persistent kernel |
rlm_vars |
List what is in the namespace (names, types, sizes, previews) |
rlm_reset |
Kill the kernel and discard all state |
rlm_skills |
List Python-backed skills importable in the kernel |
rlm_checkpoint |
Save/restore the namespace to disk so state survives a crash |
rlm_refine |
Backward-compatible durable harness: distil and immediately apply reversible entries |
rlm_eval |
Stage, evaluate, promote/reject, and regression-rollback refinements |
rlm_improve |
Durable observations, isolated builds, deterministic gates, manual promotion/rollback plans, and signed state export/import |
rlm_improve stores observations and candidate history under
$HERMES_HOME/state/rlm/improvement, independent of plugin reinstalls. Code
candidates use the existing coder isolated-worktree primitive. Evaluation
requires explicit baseline/candidate metrics, build/test/review gates, protected
path checks, and an optional secret-stripped canary with an isolated
HERMES_HOME. Export/import archives are integrity checked and state files are
atomically written with mode 0600 under mode 0700 directories.
Promotion and rollback return review commands only. The controller never merges, pushes, publishes, restarts Hermes, or edits the active plugin, evaluator, security policy, deployment, release, or approval surfaces.
Preloaded, no import needed:
- Hermes tools —
read_file,write_file,search_files,patch,terminal,web_search,web_extract rlm(goal, context)— run one real Hermes subagent, blockingrlm_many([{goal, context}, ...])— up to 9 subagents in parallelrlm_spawn(goal)/rlm_wait([ids])/rlm_children()— handle-only subagents: children run detached, results deliver via files and a disk registry that survives kernel and gateway restartsharness_store()— CRUD access to the durable harness (see below)
Subagent children inherit the parent's model (whatever Hermes itself
runs on) but do their legwork at low reasoning effort with a tight turn
cap, passed as explicit CLI flags — profile-level agent: settings do not
survive Hermes' config merge, so flags are the reliable route. Tune via:
HERMES_RLM_CHILD_REASONING=low # any hermes reasoning level; "0" omits the flag
HERMES_RLM_CHILD_MAX_TURNS=25 # tool-loop cap per child; "0" omits the flag
HERMES_RLM_LEAF_PROFILE=rlm-leaf # minimal child profile; "0" inherits parent profile# One rlm_exec call:
trades = my_skill.load("/path/to/trades.db")
# A separate rlm_exec call — no reload:
my_skill.summarise(trades)
my_skill.by_field(trades, "exit_reason")A skill becomes importable by shipping a package under python/:
<skill>/SKILL.md
<skill>/python/<module>/__init__.py
Those python/ directories go on the kernel's sys.path at boot, so
import <module> works with no install step. tests/test_skills.py
fabricates a complete worked example.
rlm_exec (plugin handler)
└── KernelHandle one long-lived python process per session
├── stdin/stdout newline-delimited JSON, one namespace
└── Unix socket RPC token-authenticated tool bridge
├── ALLOWED_TOOLS → handle_function_call (normal Hermes dispatch)
└── __rlm_delegate__ → hermes chat -q (a real subagent)
delegate_task is an agent-loop tool: it needs the live parent agent object,
which plugin handlers never receive. Rather than fork Hermes core, rlm()
spawns a real Hermes CLI subagent — same isolation, no core patch. Depth is
capped at MAX_RLM_DEPTH = 2 via the HERMES_RLM_DEPTH env var.
Lessons should outlive the session that earned them. The harness stores
small entries — prompt-notes, memories, skills, subagent recipes — per
session and globally, and injects a compact overview into future system
prompts. rlm_refine distils the recent transcript into at most 4
evidence-backed edits via one model call; every refinement is validated,
snapshotted in a ledger, and reversible with rollback. From inside the
kernel, harness_store() gives direct CRUD access. This is a deliberate
lite port of Prime Agent's continual harness: same invariants (immutable
base prompt, evidence-backed edits, rollback), no daemon required.
rlm_refine keeps its immediate-apply behavior for compatibility. Safer new
work uses rlm_eval: stage persists a bounded proposal without changing
active entries; evaluate consumes a deterministic JSON suite shaped as
{"cases":[{"baseline":0.7,"candidate":0.8}],"thresholds":{"min_candidate_mean":0.75,"min_mean_improvement":0.0,"max_case_regression":0.05}};
promote requires the latest evaluation to pass; reject closes a staged
candidate. Evaluating a promoted candidate again automatically rolls its
refinement back when any configured threshold fails. Candidates and immutable
evaluation records survive restarts under the existing HERMES_HOME-aware,
atomic mode-0600 harness state. This lifecycle can only edit harness entries;
it cannot modify base prompts, policy, source, permissions, releases or skills.
The refine model call and the subagent fast path share one provider config
(any OpenAI-compatible endpoint), set in the environment or ~/.hermes/.env:
HERMES_RLM_FAST_BASE_URL=https://your-endpoint/v1
HERMES_RLM_FAST_MODEL=your-model
HERMES_RLM_FAST_KEY_ENV=YOUR_KEY_ENV_VAR # default HERMES_RLM_FAST_API_KEYUnset means: fast path unavailable, everything falls back to the full agent.
2,846 real trading round-trips via a Python-backed skill, six follow-up questions:
| Stateless | Persistent | |
|---|---|---|
| Context bytes | 1,116 | 581 |
| Seconds | 0.42 | 0.69 |
48% less context. Speed went the other way here — and that is the honest result, not a rounding artefact.
Per follow-up question the gap is what changes behaviour:
| Marginal cost per question | |
|---|---|
| Stateless | 0.0210 s |
| Persistent | 0.0010 s — 21× cheaper |
A follow-up becomes effectively free, so you actually ask it. In practice that verification instinct caught a real data fault during development: a "free" check revealed that 77% of a trades table was a structurally different record type silently corrupting every exit-lifecycle figure — the kind of check that gets skipped when it costs a full reload.
Context saving is the reliable win. It holds at any dataset size, because the stateless lane must re-send the loading preamble with every question.
Speed is conditional. The kernel costs ~0.23s to boot and ~0.0005s per
warm call. It only wins when load_seconds × questions exceeds that boot
cost. Measured:
| Workload | Speedup |
|---|---|
| 2 columns, 2.8k rows, 6 questions | 0.61× (slower) |
| 2 columns, 12k rows, 6 questions | 2.26× |
select *, 12k rows, 5 questions |
3.4× |
Rule of thumb: reach for rlm_exec when the load is expensive or you have
several questions. One cheap question is better served by execute_code.
Keep the reply envelope lean. An early version returned kernel
diagnostics on every call and measured −79% context: the metadata cost more
than the kernel saved. Successful replies now carry only ok, value and
any output; diagnostics appear on failures, where they help.
Measured before these existed: 12 concurrent kernels held 206 MB with no cap, and a single kernel allocated 1.5 GB unchallenged on a 16 GB laptop.
| Limit | Value | Behaviour on breach |
|---|---|---|
| Kernels | 12 | Evicts least-recently-used; evicted session still works |
| RSS per kernel | 2048 MB | warning field with a concrete remedy |
| Idle lifetime | 1 hour | Reaped |
| Parallel subagents | 9 | Queued |
| Delegation depth | 2 | Refused with an explicit error |
| Checkpoint store | 14 days / 500 MB | Pruned on every save |
A subagent costs ~30s, so rlm_many is not a convenience — it is the
difference between 4 goals taking 30s and taking 75s (measured 2.5×).
The kernel runs model-generated Python with your own permissions. It is a
durable control environment, not a security sandbox. The RPC socket is
mode 0600 in a mode 0700 temp dir with a per-kernel token, and the tool
allow-list is enforced parent-side, so the kernel cannot widen its own reach.
State persists in memory — anything you load stays until rlm_reset, an
hour of idleness, or session end.
cd ~/.hermes/plugins/hermes-rlm
for t in tests/test_*.py; do ~/.hermes/hermes-agent/venv/bin/python "$t"; done
~/.hermes/hermes-agent/venv/bin/python tests/benchmark_context.pyAll thirteen suites print ok and exit 0 on any machine — tests fabricate
their own fixtures; the two benchmarks skip cleanly unless you point them
at a dataset (HERMES_RLM_BENCH_DB, HERMES_RLM_IMPACT_*). CI runs the
self-contained suites across Python 3.11–3.13 with both RSS policies in an
empty HERMES_HOME; suites that exercise a live Hermes CLI run in the
release environment.
launchctl kickstart -k (and systemctl restart) start the replacement
while the old instance's sockets can still be in TIME_WAIT. Hermes' API
server treats a bind failure as permanent (retryable=False in
gateway/platforms/api_server.py), so the gateway then runs for hours
without its API — messaging platforms keep working, which is why nobody
notices. Observed in the wild: 9 hours down.
scripts/hermes-restart.sh does it properly: stop → wait until the port is
genuinely free (killing only leftover gateway listeners) → start → verify
/health, with one automatic second pass if the API did not come up.
scripts/hermes-restart.sh # 25s on the reference machine
scripts/hermes-restart.sh --detached # when an agent restarts itself
scripts/hermes-restart.sh ai.hermes.gateway.leaf 8643--detached matters inside an agent turn: SIGTERM propagates to the process
group, so a foreground restart kills the very process that issued it. The
script uses setsid to step out of that group. It supports launchd,
systemd-user, and containers where the gateway is PID 1.
Two real bugs found by a fleet operator running this on Hermes v0.19.0 (thanks!):
- Harness sync missed writes on coarse-mtime filesystems.
_sync()comparedst_mtimeonly, so several consecutive writes sharing one timestamp made a second writer's changes invisible. Identity is now(st_mtime_ns, st_size);tests/test_harness.pyreproduces the five-rapid-writes scenario. --reasoningdoes not exist on older Hermes CLIs, and argparse rejected it with exit 2 — every subagent failed silently in 0.2s. Child flags are now capability-probed againsthermes chat --help(cached, fails open), so the plugin only sends flags the installed CLI accepts.
Also new:
coder_many([{goal, repo, test_cmd}, ...])— independent coding tasks in parallel worktrees (default 3 concurrent,HERMES_RLM_CODER_PARALLEL). Same call shape doubles as best-of-N: submit one goal N times, pick the strongest diff. Never give two tasks overlapping edit scopes.- Automatic adversarial review: every
coder()result carries areviewfield from a second agent that attacks the diff (advisory — the test gate still ownsok). Disable per call withskip_review=Trueor globally viaHERMES_RLM_CODER_REVIEW=0.
New kernel primitive coder(goal, repo, test_cmd="", context=""): repo
work becomes a fixed-discipline pipeline instead of hand-rolled terminal
calls. It creates an isolated git worktree, runs a full Hermes agent
in it (high reasoning, own tools/skills/rules — Hermes codes itself, no
external CLI required), gates the result on test_cmd and feeds one
bounded retry with the failure output, then returns the diff + test
report. Nothing is merged automatically: the orchestrator reviews the
diff and decides; the worktree is kept for inspection.
Knobs: HERMES_RLM_CODER_REASONING (high), HERMES_RLM_CODER_MAX_TURNS
(40), HERMES_RLM_CODER_TIMEOUT (900), HERMES_RLM_CODER_RETRIES (1),
HERMES_RLM_ENABLE_CODER=0 to disable, and HERMES_RLM_CODER_CMD as a
{workdir}/{prompt} template to substitute a different worker CLI.
Hermes already ships the daemon prime-agent needed a subsystem for: the
gateway. Run a second gateway on the minimal leaf profile with only its
API server bound (e.g. port 8643) and it becomes a resident worker:
agent-loop init is paid once at its start, so a blocking rlm() call
costs roughly model time.
# ~/.hermes/.env
HERMES_RLM_WARM_URL=http://127.0.0.1:8643Measured: 8.6–12s cold CLI spawn → 1.8–5.1s per full agent run. Any
failure (gateway down, empty reply) falls back to the CLI spawn silently,
so the warm lane can only make delegation faster, never less reliable.
rlm_spawn keeps using detached CLI children (surviving restarts is the
point there). Setup notes for the leaf gateway live in the section below;
supervision is plain launchd/systemd KeepAlive — no custom daemon code.
Caveats: the warm gateway needs its own provider config (the gateway path
does not inherit the main config's providers), its own .env WITHOUT your
messaging tokens (a second Telegram poller would conflict with your main
gateway), and a separate API port.
- Hard feature flags (the last fleet-review blocker): setting
HERMES_RLM_ENABLE_CHECKPOINT=0,HERMES_RLM_ENABLE_REFINE=0,HERMES_RLM_ENABLE_SUBAGENTS=0orHERMES_RLM_ENABLE_PYTHON_SKILLS=0makes that capability absent, not hidden: the tool is never registered, kernel helpers refuse with an operator message, skills paths stay off the kernel'ssys.path, and with checkpoints off even autosave/salvage (a pickle path) is a no-op.tests/test_feature_flags.pyproves both modes. - GitHub Actions CI: Python 3.11–3.13 × RSS policy warn/stop in an empty
HERMES_HOME.
A minimal pilot (exec/vars/reset only) is now one line of configuration:
HERMES_RLM_ENABLE_CHECKPOINT=0 HERMES_RLM_ENABLE_REFINE=0 \
HERMES_RLM_ENABLE_SUBAGENTS=0 HERMES_RLM_ENABLE_PYTHON_SKILLS=0Driven by an external fleet review (thanks!). Safe defaults everywhere; previous behaviour stays available behind explicit opt-ins.
HERMES_HOMErespected in every state path (harness, checkpoints, spill, subagent registry, env, skills, profiles) — isolated profiles and migrated agents no longer leak state into the wrong home. Covered by a dedicatedtests/test_isolation.pyproving two homes stay disjoint.- Cross-session checkpoint restore is now opt-in (
allow_cross_sessionparameter, default false). A session can never silently load another session's pickle. - Children load AGENTS.md/SOUL.md by default. The
--ignore-rulesspeedup (~30%) is now an explicit operator choice:HERMES_RLM_CHILD_IGNORE_RULES=1. - Resource ceilings are configurable and can be made hard:
HERMES_RLM_MAX_KERNELS,HERMES_RLM_MAX_RSS_MB,HERMES_RLM_IDLE_SECONDS, andHERMES_RLM_RSS_POLICY=stop(autosave + kill instead of a warning — recommended on shared machines). - Harness prompt injection can be disabled (
HERMES_RLM_HARNESS_INJECT=0) for fleets that already have centrally governed memory layers. - Fast-path live speed checks skip cleanly when no provider is configured, and the speedup threshold is environment-aware (>1.2×).
Suggested pilot posture for a fleet:
HERMES_RLM_MAX_KERNELS=2 HERMES_RLM_MAX_RSS_MB=512 HERMES_RLM_IDLE_SECONDS=900 HERMES_RLM_RSS_POLICY=stop HERMES_RLM_HARNESS_INJECT=0 plus the 0.4.1 feature flags to reduce the
surface to rlm_exec/rlm_vars/rlm_reset.
- Handle-only subagents:
rlm_spawn(goal)returns a handle immediately; children run detached (they survive kernel and even gateway restarts) with output to a file, never into the parent context. Collect withrlm_wait([ids]), inspect withrlm_children(). A disk registry finalises orphans visibly (finished-unverified), never silently. - Autosave + auto-restore: LRU eviction and the idle reaper now save the
namespace first, and the next
rlm_execfor that session restores it automatically with a one-line note. Resource management no longer costs state. - Continual-harness-lite: durable entries (prompt-notes, memories,
skills, subagent recipes) per session + global, injected compactly into
future system prompts.
rlm_refinedistils the recent transcript into at most 4 evidence-backed edits via one model call — validated, snapshotted in a refinements ledger, reversible withrollback. Kernel-side access viaharness_store(). - Leaf profile for subagents: children route to a minimal
rlm-leafprofile when present (no MCP spawns, no memory, no plugins) and are marked as delegated contexts (kanban mutations blocked). Configure viaHERMES_RLM_LEAF_PROFILE(set0to inherit the parent profile). - Eleven test suites, all green.
- Oversized
rlm_execoutput now spills to disk instead of losing its middle: the reply keeps the tail plus the spill-file path (mode 0600, newest 40 kept), so the full text stays reachable viaopen()in the kernel. Mirrors prime-agent's output-accumulator design. - Tool description now teaches variable-binding discipline: assign results to named variables, print only the summary — printed text is permanent conversation context, variables are free.
plugin.yamllists all five tools (checkpoint and skills were missing).
Initial release: persistent kernel, tool bridge, subagents, checkpoints, Python-backed skills, resource limits.
The three big adoption candidates from the prime-agent study shipped in 0.3.0 (harness-lite, handle-only subagents, autosave/auto-restore). Still open, in rough order of value:
- Per-variable pickling with a manifest — the autosave currently reuses whole-namespace checkpoints; per-variable saves would make partial failures more granular.
- Compaction truncation of old rlm results — needs hermes-core cooperation; out of reach for a plugin.
Deliberately rejected: prime-agent's fork-server (Linux-only; fork without
exec is unsafe on macOS) and the full daemon/supervisor (a whole subsystem;
only worth it if detach/reattach becomes a hard requirement — rlm_spawn
covers the async need without it).
MIT — see LICENSE.