Skip to content

fix(sleep): skip Claude Code isMeta records when harvesting user prompts - #295

Merged
Yifan Yang (Yif-Yang) merged 1 commit into
microsoft:mainfrom
JayOfTheKeyboard:fix/harvest-skip-claude-meta-messages
Sep 30, 2026
Merged

Yifan Yang (Yif-Yang) merged 1 commit into
microsoft:mainfrom
JayOfTheKeyboard:fix/harvest-skip-claude-meta-messages

Conversation

@JayOfTheKeyboard

Copy link
Copy Markdown

Problem

Claude Code writes some text it injects on the user's behalf into the transcript as a type: "user" record with message.role: "user", marked isMeta: true. The most common case is the full SKILL.md body of each skill the agent loads ("Base directory for this skill: …"). Others are messages relayed from another session, usage-limit notices, and "Continue from where you left off." The user typed none of it.

digest_transcript counts every such record as a user prompt and runs _detect_feedback over it. Skill documents are full of words the feedback heuristic reads as complaints ("wrong", "revert", "did not", "broken"). A session that loaded a skill is then mined as fail whatever the user actually said. A session opened with a slash command (the <command-name> message is already filtered) takes the skill body itself as the task intent.

Measured over the 400 most recent sessions in one real ~/.claude/projects (Claude Code 2.1.x, heavy skill use), with harvest.digest_transcript and mine.heuristic_mine:

main (79124b3) this PR
harvested user prompts 2846 2231
… of which a SKILL.md body 320 0
… of which a relayed cross-session message 167 0
feedback signals 1834 609
mined tasks whose intent is a SKILL.md body 57 0
mined tasks labelled fail 264 of 385 122 of 362

Every one of the 320 skill-body records carries isMeta: true. The 23 tasks that disappear are sessions whose only "user" text was injected (a skill body, a relayed message, a /context report).

End to end, skillopt-sleep harvest --backend mock --source claude on one of those sessions writes [val/fail] on main and [val/unknown] with this change. The difference in the task draft is the skill body leaving context_excerpt, and the outcome no longer being decided by it.

Fix

In digest_transcript, skip user records whose isMeta is true, before the text is counted as a prompt or scanned for feedback. The one exception is an injected body that contains an _AGENT_SESSION_MARKERS marker. It is kept, so a session driven by the plugin's own /skillopt-sleep command body ("You are driving **SkillOpt-Sleep**") is still dropped by _is_agent_session, as it is on main. Skill attribution does not change: skills_used comes from the assistant's Skill tool call, so the fan-out's skill hint is kept.

This follows #99, which filtered sub-agent transcripts and expanded slash-command bodies for the same reason: prompts the user never wrote were being mined as their tasks.

Not changed

  • Only the Claude Code source. isMeta is a Claude Code field; the other harvesters read other formats.
  • _is_meta_prompt's text heuristics are left as they are. The isMeta flag is the structural signal, so no new text marker is added.
  • A session whose only user records are injected now has n_user_turns == 0 and is dropped like any other empty session. In the sample above that was 20 sessions: a skill slash command with no typed follow-up, a session driven only by relayed messages, and a /context-only session. I think dropping them is right (there is no user intent to mine), but it is your call.
  • The ARGUMENTS: tail Claude Code appends to an injected skill body is dropped with the body. For a user-typed /skill args, the arguments are also in the <command-args> record, which _is_meta_prompt already skipped before this change.
  • No CHANGELOG entry; happy to add one under ### Fixed if you want it in the PR.

Test

Three tests in tests/test_sleep_engine.py::TestHarvest. The first two use a transcript shaped like the real one: a typed prompt, a Skill tool call, the injected isMeta skill body, and an isMeta relayed message.

  • test_digest_skips_claude_injected_meta_messages: only the typed prompt is harvested, no feedback signals, skills_used still names the skill.
  • test_injected_skill_body_does_not_decide_the_mined_outcome: with a closing "perfect, thanks", the mined task is success with the typed intent.
  • test_injected_agent_marker_body_still_drops_the_session: a /skillopt-sleep session whose command body arrives as an isMeta record is still excluded by harvest(). This one passes on main and guards the existing behaviour. It fails if isMeta records are skipped without the marker exception.

On main the first two fail: the first with the skill body and the relayed message in user_prompts, the second with 'fail' != 'success'.

Mutations, each run against the three tests:

Mutation Result
skip removed 2 failed
marker exception removed 1 failed (agent-marker test)
skip only skill bodies ("Base directory" in text) 2 failed
skip only relayed messages 2 failed
drop the prompt but still scan it for feedback 2 failed

Validation

  • python -m pytest -q: 1500 passed, 11 skipped on Python 3.10 and 3.14 (main: 1497 passed, 11 skipped).
  • ruff check tests/test_sleep_engine.py clean. ruff check skillopt_sleep/harvest.py reports one I001 that is already on main; the repo-wide count is 149 on both.
  • Real-tool run: skillopt-sleep harvest --backend mock --source claude --scope all --lookback-hours 0 under an isolated HOME, as above.

Claude Code writes text it injects on the user's behalf as role "user"
records marked isMeta: true. The most common is the full SKILL.md body
of every skill the agent loads; others are messages relayed from other
sessions and usage-limit notices. The harvester counted all of them as
user prompts and ran the feedback heuristic over them, so words a skill
document uses ("wrong", "revert", "did not") labelled the session a
failure, and a session opened with a slash command mined the skill body
itself as the task intent.

Skip isMeta user records in digest_transcript, except one that carries
an _AGENT_SESSION_MARKERS marker, so a session driven by the plugin's
own /skillopt-sleep body is still dropped as an agent session. The skill
stays attributed through the assistant's Skill tool call.
@Yif-Yang
Yifan Yang (Yif-Yang) merged commit 94ebddd into microsoft:main Sep 30, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants