Skip to content

Commit e4ea6a6

Browse files
chore(release): v0.2.0
Highlights since v0.1.0: - feat: SkillOpt-Sleep engine — nightly offline self-evolution (harvest -> mine -> replay -> consolidate behind a validation gate), with multi-objective reward, experience replay + dream rollouts, slow-update long-term memory, and secret redaction in cycle diagnostics. Shipped as the `skillopt-sleep` CLI. - feat: cross-tool backends & plugin shells — Claude, Codex (+Desktop harvest), Copilot, Devin, and OpenClaw. - feat: SearchQA split materialization + rollout fail-fast. - fix: Windows robustness for claude/codex backends, hardened JSON fallback, Qwen timeout/thinking gating, Codex failure surfacing. Packaging: - Bump pyproject / skillopt / skillopt_sleep to 0.2.0. - Restore skillopt_webui to the packaged wheel. See CHANGELOG.md for the full changelog and contributor acknowledgements. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent 5487e2c commit e4ea6a6

6 files changed

Lines changed: 122 additions & 6 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 100 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,100 @@
1+
# Changelog
2+
3+
All notable changes to SkillOpt are documented here. This project adheres to
4+
[Semantic Versioning](https://semver.org/) and the format is based on
5+
[Keep a Changelog](https://keepachangelog.com/).
6+
7+
## [0.2.0] — 2026-07-02
8+
9+
The headline of this release is **SkillOpt-Sleep**: a nightly offline
10+
self-evolution engine that harvests a coding agent's real session
11+
transcripts, mines recurring tasks, replays them offline, and consolidates
12+
short-term experience into long-term memory and skills — all behind the same
13+
held-out validation gate that keeps SkillOpt training honest. It ships as a
14+
decoupled top-level package (`skillopt_sleep/`, zero dependency on the
15+
research code) and as the new `skillopt-sleep` CLI.
16+
17+
### Added
18+
- **SkillOpt-Sleep engine** — nightly offline self-evolution cycle
19+
(harvest → mine → replay → consolidate) behind a validation gate, exposed
20+
as the `skillopt-sleep` console script and `python -m skillopt_sleep`.
21+
- Multi-objective reward (accuracy / tokens / latency) with user preferences.
22+
- Multi-rollout contrastive reflection under a token/time budget.
23+
- Experience replay + controllable dream rollouts (opt-in).
24+
- Slow-update long-term memory field (runs even with the gate off).
25+
- 3-way train/val/test split with `gate_mode on|off`.
26+
- Verifier-discipline validation gate, with a stress-test suite
27+
(thanks @Tanmay9223, #87).
28+
- **Cross-tool backends & plugin shells** for Claude Code, Codex, Copilot,
29+
Devin, and OpenClaw:
30+
- Codex Desktop transcript harvesting, skill-first Codex integration, and a
31+
reviewed task-file flow (thanks @Kirchberg, #48, #49, #60).
32+
- GitHub Copilot backend (`CopilotCliBackend`) + research-engine MCP plugin
33+
(thanks @Dongbumlee, #50).
34+
- Devin plugin: MCP server + ATIF-v1.7 harvest (thanks @xerxes-y, #88).
35+
- OpenClaw shell for SkillOpt-Sleep (thanks @Elzlxx, #59).
36+
- **SearchQA** split materialization helper and fail-fast on systemic rollout
37+
failures, with a `searchqa` install extra (thanks @summerview1997,
38+
#63, #64, #65).
39+
- WebUI environment loading and backend preflight (thanks @summerview1997, #63).
40+
41+
### Changed
42+
- Decoupled the Sleep engine into a standalone top-level `skillopt_sleep/`
43+
package with zero dependency on the research code.
44+
- Made `EnvAdapter.reflect` a shared default so reflect kwargs are no longer
45+
dropped (thanks @imshunsuke, #44).
46+
- English-only pass across the engine, plugins, and docs.
47+
48+
### Fixed
49+
- Windows robustness for the Claude/Codex backends, plus a hardened JSON
50+
fallback path (thanks @Yif-Yang, #79).
51+
- Reject prose pseudo-JSON wrapped in single quotes/backticks (#82).
52+
- Surface Codex auth/model/version failures instead of silently scoring 0
53+
(thanks @dmmdea, #92).
54+
- Redact secrets before persisting cycle diagnostics.
55+
- Configure the `qwen_chat`/`minimax` backends so local LLM endpoints work
56+
(thanks @imrehg, #85).
57+
- Forward the Qwen target timeout and gate `enable_thinking` for vLLM targets
58+
(thanks @mvanhorn, #40).
59+
- Make `--bare` conditional on `ANTHROPIC_API_KEY` (#68), add a
60+
`SKILLOPT_SLEEP_PYTHON` override with a lookback-hours first-run fallback
61+
(#74), and fix ALFWorld gamefile paths relative to `ALFWORLD_DATA`.
62+
63+
### Packaging
64+
- Bump `skillopt`, `skillopt.__version__`, and `skillopt_sleep.__version__`
65+
to `0.2.0`.
66+
- Restore `skillopt_webui` to the built wheel (it was dropped when the
67+
`packages.find` include list was made explicit).
68+
- Add the `searchqa` extra and include `json_repair` in the `claude`, `qwen`,
69+
and `all` extras.
70+
71+
### Acknowledgements 🙏
72+
v0.2.0 landed thanks to our community contributors — thank you!
73+
74+
- @Kirchberg — Codex Desktop harvesting, skill-first Codex integration,
75+
reviewed task-file flow (#48, #49, #60)
76+
- @Dongbumlee — GitHub Copilot backend + research-engine MCP plugin (#50)
77+
- @summerview1997 — SearchQA materialization, rollout fail-fast, WebUI
78+
preflight (#63, #64, #65)
79+
- @xerxes-y — Devin plugin: MCP server + ATIF-v1.7 harvest (#88)
80+
- @Elzlxx — OpenClaw shell for SkillOpt-Sleep (#59)
81+
- @imshunsuke — shared `EnvAdapter.reflect` default + docs fixes (#43, #44)
82+
- @mvanhorn — Qwen timeout forwarding + `enable_thinking` gating (#40)
83+
- @dmmdea — surface Codex auth/model/version failures (#92)
84+
- @Tanmay9223 — verifier-discipline stress test (#87)
85+
- @imrehg — `configure_qwen_chat` for local LLM endpoints (#85)
86+
- @samuelgoofus-boop — community contributions
87+
88+
Special thanks to @Yif-Yang for driving the SkillOpt-Sleep engine.
89+
90+
**Full changelog:** https://github.com/microsoft/SkillOpt/compare/v0.1.0...v0.2.0
91+
92+
## [0.1.0] — 2026-06-02
93+
94+
Initial public release: the full training loop (rollout → reflect →
95+
aggregate → select → update → evaluate), multi-backend support
96+
(OpenAI / Azure / Claude / Qwen / MiniMax), six built-in benchmarks, and the
97+
WebUI dashboard.
98+
99+
[0.2.0]: https://github.com/microsoft/SkillOpt/releases/tag/v0.2.0
100+
[0.1.0]: https://github.com/microsoft/SkillOpt/releases/tag/v0.1.0

‎README.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -14,6 +14,7 @@
1414
---
1515

1616
## News 🔥🔥🔥
17+
- **[2026-07-02]** 🚀 **SkillOpt [v0.2.0](https://github.com/microsoft/SkillOpt/releases/tag/v0.2.0) is out on [PyPI](https://pypi.org/project/skillopt/)!** Headline feature: **SkillOpt-Sleep**, a nightly offline self-evolution engine (harvest → mine → replay → consolidate, all behind a held-out validation gate) with multi-objective reward, experience replay + dream rollouts, and long-term memory — now shipped as the `skillopt-sleep` CLI. This release also adds cross-tool backends and plugin shells for **Claude, Codex, Copilot, Devin, and OpenClaw**, SearchQA split materialization, Windows robustness, and hardened JSON parsing. See the [release notes](https://github.com/microsoft/SkillOpt/releases/tag/v0.2.0) for the full changelog and contributor acknowledgements.
1718
- **[2026-06-15]** 😴 **SkillOpt-Sleep (preview)** — a nightly offline self-evolution companion for local coding agents (Claude Code / Codex / Copilot): review past sessions, replay recurring tasks, and consolidate validated skills behind a held-out gate. See **[`docs/sleep/README.md`](docs/sleep/README.md)** for what it is, how to use it, and results.
1819
- **[2026-06-03]** 🎉 **[gbrain](https://github.com/garrytan/gbrain), [gbrain-evals](https://github.com/garrytan/gbrain-evals/blob/main/docs/benchmarks/2026-06-03-skillopt.md), and [darwin-skill](https://github.com/alchaincyf/darwin-skill) have all integrated SkillOpt.**
1920
- **[2026-06-02]** 🎉 **SkillOpt [v0.1.0](https://github.com/microsoft/SkillOpt/releases/tag/v0.1.0) is now available on [PyPI](https://pypi.org/project/skillopt/)!** Install with `pip install skillopt`. This initial release includes the full training loop (rollout → reflect → aggregate → select → update → evaluate), multi-backend support (OpenAI / Azure / Claude / Qwen / MiniMax), six built-in benchmarks, and WebUI dashboard.

‎docs/sleep/README.md‎

Lines changed: 14 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -28,6 +28,20 @@ experience → long-term competence).
2828

2929
## How to use it
3030

31+
### Quickest path: the `skillopt-sleep` CLI (pip)
32+
33+
```bash
34+
pip install skillopt # installs the engine + the `skillopt-sleep` command
35+
skillopt-sleep dry-run # harvest + mine + replay, report only (changes nothing)
36+
skillopt-sleep run # a full nightly cycle; the proposal is staged for review
37+
skillopt-sleep status # show state + the latest staged proposal
38+
skillopt-sleep adopt # apply the latest staged proposal
39+
skillopt-sleep schedule # install a nightly cron entry for this project
40+
```
41+
42+
The per-agent plugin shells below (Claude Code / Codex / Copilot) still come from the
43+
repo; the CLI above is the standalone, pip-only way to run a cycle.
44+
3145
One engine, thin per-agent shells (see [`plugins/`](../../plugins)):
3246

3347
| Platform | Folder | Install |

‎pyproject.toml‎

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
44

55
[project]
66
name = "skillopt"
7-
version = "0.1.0"
7+
version = "0.2.0"
88
description = "SkillOpt: Agentic Skill Optimization via Reflective Training Loops"
99
readme = "README.md"
1010
license = {text = "MIT"}
@@ -68,9 +68,10 @@ Repository = "https://github.com/microsoft/SkillOpt"
6868
Issues = "https://github.com/microsoft/SkillOpt/issues"
6969

7070
[tool.setuptools.packages.find]
71-
# skillopt* = the research package; skillopt_sleep = the open-source Sleep tool
72-
# (decoupled, zero dependency on the research code).
73-
include = ["skillopt", "skillopt.*", "skillopt_sleep", "skillopt_sleep.*", "scripts*"]
71+
# skillopt* = the research package
72+
# skillopt_sleep = the open-source Sleep tool (decoupled, zero research dep)
73+
# skillopt_webui = the Gradio dashboard (installed via the `webui` extra)
74+
include = ["skillopt", "skillopt.*", "skillopt_sleep", "skillopt_sleep.*", "skillopt_webui", "skillopt_webui.*", "scripts*"]
7475

7576
[tool.ruff]
7677
line-length = 120

‎skillopt/__init__.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -12,7 +12,7 @@
1212
6. Evaluate — validate candidate skill, accept/reject
1313
"""
1414

15-
__version__ = "0.1.0"
15+
__version__ = "0.2.0"
1616

1717
from skillopt.types import ( # noqa: F401
1818
BatchSpec,

‎skillopt_sleep/__init__.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -17,4 +17,4 @@
1717
from __future__ import annotations
1818

1919
__all__ = ["__version__"]
20-
__version__ = "0.1.0"
20+
__version__ = "0.2.0"

0 commit comments

Comments
 (0)