← the whole session plugin/skills/knowledge-atom-extraction/SKILL.md
Reference design for extracting and consuming typed knowledge atoms (corrections, mental models, findings, decisions, patterns, open questions, workflows) from session transcripts. Superseded in the example project this pattern was drawn from by the lighter-weight `instinct-harvest` skill — kept here as a reference design, not a default setup step.
Before anything else: this is a historical reference, not a live default
In the example project this pattern was drawn from, this exact pipeline was built, used for a few months, and then explicitly abandoned — the store shows real records through early March 2026 and nothing after. The reason: the way agents were recording information was losing all the context, losing all the conditions, and confusing agents more than it was adding value. It was replaced by a lighter format, the instinct-harvest skill (situation, lean, reason, source — an "instinct" the agent still has to weigh, not a rule it obeys blindly), which is the piece this harness actually ships as a working default.
If you're trying to capture what you've corrected agents on before, use instinct-harvest first. This page stays because the underlying discipline it describes — typed, dated, verbatim-quoted extraction with weighted retrieval — is genuinely reusable, and someone building a fuller structured knowledge store than instinct-harvest provides may want to build from this design instead of starting over. Nothing below is wired into this plugin's setup; every endpoint and path in it is illustrative of a design you'd build yourself.
What Are Knowledge Atoms?
Self-contained typed assertions meant to prevent agent confusion by capturing corrections, mental models, findings, decisions, patterns, workflows, and open questions as first-class searchable objects.
Illustrative example of the failure mode this was meant to solve: an agent asks whether to hide a metrics column because it looks like it's measuring only organic traffic — when actually that data source receives all traffic sources, and the real distinction between the two things being compared is something else entirely (a different property, not the one the agent guessed). A correction atom, if it existed and got retrieved, would have prevented the wrong guess. This is the general shape of the problem: agents re-guessing at distinctions a person already explained, because the explanation only lived in one past conversation.
8 Atom Types
| Type | When to Extract | Signal Phrases |
|---|---|---|
| correction | Human says agent is wrong | "You're confused", "No, actually", "That's wrong" |
| mental_model | Human explains how a system works | "The way this works is", "Think of it as", "There are two..." |
| finding | Human shares empirical data | "Yesterday we spent $150 and got", "The data shows" |
| decision | Human states a choice + reasoning | "We decided to", "Let's go with", "The reason we chose" |
| pattern | Human describes recurring behavior | "Every time X happens, Y", "This always means" |
| open_question | Human flags unresolved uncertainty | "I'm not sure yet whether", "We need to figure out" |
| workflow | Natural sequence to realize an objective | "First X, then Y, then Z", "The process is..." |
| gestalt | Pre-composed domain briefing | Created manually or by grouping related atoms |
Verbatim Quotes
Every atom should capture the exact human words (verbatimQuote field) — this is the cleanest signal. In a working implementation, the extractor captures these automatically, and consuming code should show verbatim quotes alongside the statement, not a paraphrase.
Temporal Awareness
Every atom has an established date = the session date when the knowledge was expressed. In the original design this was enforced with zero exceptions — anything without dates breaks semantic search.
- Atoms >14 days old: confidence downgraded to "inferred", time-sensitive data flagged
- Batch checkpoints track session age range for quality monitoring
Consuming Atoms (Reading)
In the original design, atoms surfaced inside search results under a KNOWLEDGE ATOMS section, weighted by type:
Type multipliers (example weighting): correction=1.15, mental_model=1.10, gestalt=1.10, workflow=1.05, open_question=0.85.
Source weights (example weighting): a project's own default source=1.3, company-level source=1.0, a specific person's own statements=1.2.
Proposing Atoms (Writing)
If you build your own version of this, the shape of a write is:
Single atom:
curl -X POST "$YOUR_API_HOST/knowledge-atoms/propose" \
-H "Content-Type: application/json" \
-H "x-api-key: $YOUR_API_KEY" \
-d '{"type":"correction","statement":"...","domain":["tag1"],"wrongAssumption":"...","whyConfused":"...","verbatimQuote":"exact human words"}'
Batch:
curl -X POST "$YOUR_API_HOST/knowledge-atoms/propose-batch" \
-H "Content-Type: application/json" \
-H "x-api-key: $YOUR_API_KEY" \
-d '{"atoms":[{"type":"finding","statement":"...","domain":["tag1"],"verbatimQuote":"exact human words","established":"2026-03-01"}]}'
Running an Extractor
If you build one, the shape of the CLI is:
# Mine last 5 sessions
node your-extraction-script.js --recent 5
# Dry run (preview only — no writes)
node your-extraction-script.js --recent 10 --dry-run
# Specific session
node your-extraction-script.js --session <UUID>
# Force re-extraction (ignores whatever "already processed" state file you keep)
node your-extraction-script.js --recent 20 --force
Multi-Channel Output
A useful extractor produces two channels per session:
- Atoms → your knowledge-atom store
- Intents → your intent store (see
intent-db, if installed)
Batch checkpoints every 10 sessions should report: signal-to-noise ratio, duplicate rate, type distribution, domain coverage, session age range — this is how you'd notice the pipeline degrading the way the example above did.
Extraction Statement Quality Rules
Each atom's statement MUST be:
- Self-contained — readable without conversation context
- Typed correctly — corrections must include wrongAssumption + whyConfused
- Domain-tagged — semantic tags for retrieval (lowercase, hyphenated)
- Non-obvious — skip routine instructions, approvals, small talk
- Dated —
establishedfield = session date, zero exceptions - Quoted —
verbatimQuotecaptures exact human words when available
Architecture (reference shape, not shipped infrastructure)
- Store: a
knowledgeatomscollection/table with an auto-ID, correction chains, gestalt clusters - Dual-write for search: a flat JSONL log alongside the structured store, for vector search
- Search integration: a source type
atomwith type multipliers as above - Extraction: a cheap/fast model doing multi-channel classification of human messages from session transcripts
- Dedup: hash of statement text, plus a state file tracking which sessions were already processed
- Temporal: session mtime = established date, staleness detection at 14 days
Calibration Protocol (if you build this)
After adding new atoms or adjusting weights, verify retrieval quality against a few of your own known corrections — pick 2-3 real corrections a person gave an agent in the past, phrase them as a query the way a future agent might ask, and confirm the correction atom ranks in the top 2 results. If it doesn't, the weighting or the statement wording needs adjustment — tune atom weight, correction multiplier, and retrieval threshold against your own known-correction test queries.