A Claude Code skill that turns your Obsidian wiki into a self-evolving learning loop: generates open-ended questions from your concept pages, grades your answers strictly, tracks mastery over time, and surfaces wiki gaps as actionable patches.
中文简介:把你的 Obsidian 知识库变成一个会自我进化的学习闭环——基于你的概念页出题、严格评分、跟踪熟练度,并把答题暴露的 wiki gap 转成可操作的补 wiki 行动项。
Most flashcard / quiz tools are static:
- You write what you know → tool tests what you wrote → score goes up → wiki never changes
- Result: your max possible score = your wiki's completeness. Gaps stay invisible.
concept-quiz inverts this:
- A grader scores against an ideal target (a
target_roleyou define — e.g. "AI PM at company X") - When you miss a point that's not in your wiki but should be, it surfaces as a
🔧 Wiki gap - You can immediately have Claude draft a patch to your concept page — the wiki co-evolves with you
- You can override the grader's judgment — those overrides accumulate and calibrate the grader over time (lightweight RLHF)
Two loops:
- Inner: study wiki → answer → score up
- Outer: answer → wiki gap exposed → wiki improved → broader coverage next time
1. Orient Read _progress.md, ask which time bucket (10 / 30 / 60 min)
2. Plan Pick concepts: mix of priority slots + random slots
3. Execute Generate question → you answer → 2x grader (self-consistency) → write back mastery
4. Wrap up Aggregate session, log to _log.md, regenerate _progress.md
Each scored answer separates feedback into 3 tags:
| Tag | Meaning | Action |
|---|---|---|
[页内有] (in-wiki) |
You forgot what your own wiki says | Re-read your concept page |
[页内无 / 应补强] (wiki gap) |
Wiki doesn't have it but ideal target requires it | Patch the wiki |
[超纲] (out-of-scope) |
Niche / research front, not required for your target_role |
Just FYI, no penalty |
This skill is opinionated. It assumes:
- You use Obsidian (or any wiki with
[[wikilinks]]) + YAML frontmatter on.mdfiles - Concept pages are organized in folders under one root directory (the skill globs them)
- Each concept page is roughly one topic (definition + key points + relationships + self-test, optional)
- You're OK with the skill adding two fields to each concept page's frontmatter:
mastery: 0-100andlast_reviewed: YYYY-MM-DD
Not required but works best with:
- A layered structure (e.g.
L0-foundations/,L1-classical-ml/, ...) — gives meaningful "各层进度" breakdown - Cross-linking via
[[wikilinks]]— feeds the "relationships" feedback
If your wiki structure differs significantly, you may need to tweak the path globs in SKILL.md.
-
Clone into your Claude Code skills directory:
cd ~/.claude/skills git clone https://github.com/dingdugan/concept-quiz.git
(Or copy the 4
.mdfiles into~/.claude/skills/concept-quiz/.) -
Edit
SKILL.md— change one line to point to your wiki:VAULT_WIKI = <your-vault-path>/_wikiExample:
VAULT_WIKI = /Users/yourname/Documents/MyVault/Knowledge/_wiki -
Restart Claude Code (skills are scanned at startup).
-
First run: type
/quiz. It will detect that_progress.mddoesn't exist, bootstrap one from your concept pages, ask you to set atarget_role, then offer time buckets.
The single anchor that determines what counts as [应补强] vs [超纲]. Be specific:
- ✅ "AI PM candidate at frontier-lab company, 3 months from interview"
- ✅ "ML researcher specializing in long-context attention mechanisms"
- ✅ "Backend engineer learning enough AI to integrate LLM APIs into production"
- ❌ "AI learner" (too vague — grader will drift)
Stored in _progress.md frontmatter. Edit anytime, takes effect on next session.
| Bucket | Quota |
|---|---|
| 10 min | 1 question (80% priority / 20% random) |
| 30 min | 3 priority + 1 random + 1 deep-read |
| 60 min | 4 priority + 2 random + 1-2 deep-reads |
Modifiers (append after time arg):
norandom— all priorityallrandom— all randomrandom=N— set random slot count
Example: /quiz 30 random=2
priority = (1 - mastery/100) × log(days_since_reviewed + 1)
- Concepts with
mastery >= 90excluded from priority pool (but appear in random pool — catches "thought you knew it" decay) - Random pool includes high-mastery concepts (the whole point — surface forgetting)
When the grader marks an [应补强] tag you disagree with, you can override:
override (1) to 超纲
This:
- Reverses the deduction (you get the points back)
- Logs the override into
_progress.md用户判别偏好section - Next grader run reads this log first — items overridden 2+ times are auto-applied without re-asking
Over time, the grader becomes your personal examiner instead of a generic LLM judge.
_progress.md— single source of truth for mastery / priority / patterns / overrides (auto-regenerated each session)_log.md— append-only history of all sessions (your event log)- Two frontmatter fields per concept page touched:
mastery: N,last_reviewed: YYYY-MM-DD
Nothing else is modified unless you explicitly choose to patch a wiki gap during a session.
Three principles drove the design:
-
The wiki is alive, not ground truth. Grading against a static wiki creates a closed loop. Grading against a target role with feedback to update the wiki creates a learning loop.
-
LLM-as-judge needs a specific anchor. "Common knowledge" is too abstract — graders drift and hallucinate. A concrete
target_roleties every judgment to something verifiable. -
The user always wins ties. When grader and user disagree, the user's override is ground truth — and accumulated overrides shape future grading. This is RLHF without training.
concept-quiz/
├── README.md # this file
├── LICENSE # MIT
├── SKILL.md # main entry — Claude reads this on /quiz
├── grader-prompt.md # strict examiner system prompt + tagging rubric
├── question-prompt.md # question generator with target_role + mastery calibration
├── state-schema.md # _progress.md format spec
└── examples/
├── concept-page-template.md # what a good concept page looks like
└── progress-template.md # what _progress.md looks like after a few sessions
- Grader can still drift despite all the constraints — LLMs aren't perfectly reliable judges. The
overridemechanism is your insurance. - Self-consistency uses min of 2 scores, which trends harsher. If you find it too punishing, edit
SKILL.mdStep 3.3 to usemeanor justscore_1. - Costs: each question = 1 generation + 2 grader calls (self-consistency) + occasional wiki patch drafts. At Sonnet pricing, a 30-min session ≈ a few cents.
- No spaced-repetition scheduling (SR algorithms like SM-2 assume daily use; this targets irregular weekly cadence). Priority formula handles forgetting via the
log(days)term. - Hard-coded for layered Obsidian vaults. Other wiki formats need glob adjustments.
MIT — do whatever you want, no warranty. See LICENSE.
@dingdugan — built this to learn AI concepts for a Claude Code skill. The full design conversation that produced this skill is itself a case study in iterative product design with Claude as thinking partner.
If you use this and have ideas / find bugs / want to share what you learned, open an issue.