← the whole session plugin/skills/instinct-harvest/SKILL.md
Mine INSTINCTS (not facts) from a person's confirmed sanity checks and coherence checks in session transcripts, and maintain a per-domain instinct ledger that agents read before entering a situation. General tier — domain-agnostic, works standalone with a built-in "general" domain, or a specialized domain layer can supply its own ledger path and situation taxonomy. Use when seeding a domain's instinct ledger, when harvesting a session's confirmed checks, when generating the ledger index, or when checking whether the loop is healthy.
Instinct Harvest — general tier
Print ## INSTINCT_HARVEST on its own line before any output produced under this skill.
What this is in service of
When an agent finishes reasoning about a task, it writes a check — a coherence check or a sanity check — that restates, in the person's own terms, what it thinks they asked for. The person confirms or corrects it. That confirmed material is the most valuable thing a session produces, and by default it dies in the chat thread the moment the session ends.
This skill keeps it. It captures every check and reply verbatim, has a mining pass turn the confirmed ones into short "instincts," and has later agents read the relevant instincts before they act — visibly weighing them, never blindly obeying them.
Example setup (opt-in context — the method is general)
This method was built for SEO work where many agents worked concurrently and needed to compound what got confirmed session to session. The reasoning behind every rule in this file:
Instincts are a stronger unit than facts because if a sanity check says "our intention is to only notify search engines when we push to main," that shouldn't get written down as a hard fact — the context could shift later in a way that would hold an agent back from doing its job. Written as an instinct instead, it steers agents toward the correct thing while leaving the spaciousness for them to use their own judgment about whether the instinct actually applies.
Why the sanity check is the source: sanity checks are the heavy cognitive lifting of determining whether an agent's programmatic solution addresses the original intent of the person — they're effectively the bridge.
The non-negotiable on every item: each item MUST specify the exact date it was captured and the exact context in which it's relevant, and have any undefined variable names unpacked to their semantic meaning.
An earlier attempt at this same goal — a system that mined "knowledge atoms" as plain facts — failed for a specific, identified reason: the way agents were recording information was losing all the context, losing all the conditions, and confusing agents more than it was adding value.
Everything below follows from that reasoning. If a step in this skill would produce a fact stripped of its conditions, a rule that removes the agent's judgment, an item without its capture date, or an item carrying a bare code name — the step is wrong.
What an instinct is
A fact is a proposition: true or false. When the world moves, a stale fact is a lie the next agent cannot detect, and a fact written as a rule ("always ping on push to main") pigeonholes the agent into obeying it in a situation where it no longer holds.
An instinct is a prior. It carries the moment it came from. It says: when you're in this situation, we lean this way, because of this, and here's where that came from. The agent still has to think. A stale instinct is a prior the agent weighs against what is actually in front of it. That is why facts breed hallucination and instincts do not.
Four parts, always together, one to three sentences total, in the person's own voice:
- The situation that triggers it — recognition, not a code condition. "When you're about to tell search engines about pages."
- The lean — direction, never a command. "Lean toward doing it only once they're really live for a human, on main." Never "always" or "never." Strength lives in the words: hard lean, lean, slight lean, we've gone back and forth.
- The why — the value or the past pain that produced it. "Pinging anything else teaches crawlers we serve broken pages."
- Where it came from — exact date, session, what the agent's check said, what the person replied, verbatim.
Plus, where known: when it might not apply. That is the spaciousness. And every variable unpacked: if the source check names a code constant, a flag, or an internal label, the instinct states what that is in plain meaning with its exact condition, so a cold reader needs no follow-up question.
Corrections are the strongest instincts. The person had to intervene; the pain travels with the instinct. An agent's self-discovered correction is second strongest — someone paid for it.
The pipeline
transcripts ──extract──▶ harvest inbox ──mine──▶ ledger items ──index──▶ ledger README ──consume──▶ agent's task-start check
(script, pure capture) (agent judgment) (one file each) (generated) (weighs instincts visibly)
└──health──▶ failures surfaced to the person
Two rules govern the whole thing:
- Extraction never summarizes. The script captures the check block, the person's reply, and the surrounding task context verbatim, with timestamps. It makes no judgment.
- Mining is judgment, not extraction. The mining agent reads a captured pair and asks "what lean would have produced this decision, and why?" That is a reasoning question. First batches should run on a smaller/cheaper model so quality can be measured cheaply; if instincts come back as restated facts or rules, escalate mining to a stronger model — do not lower the bar.
The scripts
This skill ships its own scripts (in scripts/ next to this file — self-contained, no other part of the plugin required):
scripts/extract-check-pairs.js— Step 1, extraction, and--healthfor Step 6scripts/mining-status.js— Step 0, what's already donescripts/verify-instinct-quotes.js— Step 3b, the mandatory quote guardscripts/build-instinct-index.js— Step 4, regenerate the ledger README
They default to reading transcripts from ~/.claude/projects/*/*.jsonl (respecting CLAUDE_CONFIG_DIR) and storing the harvest inbox and ledgers under whatever alignment-harness records instincts-harvest / alignment-harness records instincts-ledger print — or, if that CLI isn't installed, a plain .instinct-harvest/ folder under the current directory. Nothing about them depends on any one repo or product; point --project <dir> and --session-filter <substring> at whatever this domain's sessions actually look like.
There is no bundled automatic hook wiring in this skill (that lives in the plugin's own hook configuration, if it has one for this). Until/unless that's wired up for you, run the extractor yourself at a natural point — end of a session, start of the next one, or on whatever cadence you like:
node <this-skill-folder>/scripts/extract-check-pairs.js --since 7
Step 0 — Find out what is already done (never skip, never guess)
node <this-skill-folder>/scripts/mining-status.js --domain <id>
It prints every harvested pair, whether it's already mined, whether it has a reply yet, and the next free instinct id. Mine only rows marked TO MINE.
Three rules that follow from it:
- A pair already mined is finished. Re-mining it creates duplicate instincts under fresh ids, inflates the counts, and leaves two versions of one lean for the next agent to reconcile. If you believe a mined pair holds something that was missed, do not re-mine it — strengthen or amend the existing item, which is one of the four moves in Step 2.
- Take ids from
next free idand claim a range up front when several agents mine in parallel, so two agents never write the same instinct id. - If you read a pair in full and there is genuinely nothing to take, record that (a line in a
_examined-no-yield.mdfile next to the ledger) with the date, who read it, and why. "Examined and empty" is a real result. Without the record it looks exactly like "not yet looked at" and the next run pays to open it again. Never write a weak instinct to avoid an empty row — a thin batch of real instincts beats a full batch of invented ones. - A row outside this domain's
--session-filteris not yours. Leave it; whichever domain it belongs to will pick it up.
Extraction itself is idempotent — a pair already in harvest.jsonl is skipped by content hash — so re-running the extractor is always safe. Mining is not idempotent. That is why this step exists.
Step 1 — Extract
node <this-skill-folder>/scripts/extract-check-pairs.js [--since <days>] [--project <dir>] [--rebuild]
Output: one markdown file per captured pair under the harvest inbox, plus an append-only harvest.jsonl at the inbox root. Idempotent: a pair already captured (same content hash) is skipped.
A block ends at the next check heading or the end of the assistant turn — never at an ordinary ## subheading. Checks routinely use ## for internal sections, and a naive "end at any H2" rule can silently truncate a long check down to a fraction of its real content, dropping most of its concerns with nothing noticing. When an extractor correctness fix like this matters, re-run with --rebuild to re-capture everything, then re-mine any pair whose item was produced from an old, partial capture.
A pair is: an assistant message containing a check heading (SANITY_CHECK, COHERENCE_CHECK, AGENT_REFLECTION, SCOPE_DECLARATION, or GAP_ANALYSIS — bare or with a leading ## ), plus the next genuine human-typed message before and after it (the task context and the reply). Each carries its ISO timestamp, session id, project, and git branch when the transcript has one.
Verdict is a heuristic label only — confirmed-ish, corrected-ish, unanswered, ambiguous — set by the script from the reply text, and always re-judged by the mining agent reading the actual reply.
Step 2 — Mine
Run the how-to-talk-like-the-founder skill's voice-calibration step first (its own name is historical; it calibrates to whichever person you're working with). Every instinct is written in that register: intent leads, code rides inside as the anchor.
For each pair marked TO MINE for this domain:
- Read the person's reply as the verdict. Confirmed: the check's intent lines become candidate instincts. Corrected: the correction is the instinct, and the corrected line is recorded as the wrong belief with why it looked reasonable. Unanswered or ambiguous: do not mine into the ledger. Log it in a waiting file next to the ledger, so the person can confirm in a batch. Instincts enter only through their yes.
- Ask of each confirmed or corrected line: what lean would have produced this decision, and why? Which situation does it fire in? What was the alternative it was chosen over? What would make it not apply?
- Check the ledger before writing. Search existing items by situation and by wording. Then one of four moves:
- Same instinct → strengthen: add the new source under "Where it came from", update
last-confirmed, raise strength wording if warranted. Never duplicate. - Tension with an existing instinct → keep both, name the tension in each ("in tension with INST-...: …"). The person holds tensions; so does the ledger.
- Override — the person confirmed an agent deviating from an existing instinct with a reason → add "When it might not apply" to the existing item with the date and the reason.
- New → write a new item.
- Same instinct → strengthen: add the new source under "Where it came from", update
- Write the item using the format below. Then run the cold-read test: with zero session context, would a reader ask any clarifying question? If yes, the item is incomplete — fold the answer in, do not append it.
- Never write a rule. If the draft contains "always", "never", "must", or reads as a command, rewrite as a lean with a why.
Step 3 — Item format
One file per instinct at <ledger>/items/INST-<domain>-<NNNN>-<slug>.md. Never edit another agent's item except through the four moves above; never a shared file two agents write at once.
The frontmatter is machine metadata — it drives the generated index, so its keys are fixed:
---
instinct-id: INST-GENERAL-0007
domain: general
situation: about-to-tell-search-engines-about-pages
strength: hard-lean # hard-lean | lean | slight-lean | contested
captured: 2026-09-07T14:22:03Z # timestamp of the CONFIRMING or CORRECTING reply
last-confirmed: 2026-09-07T14:22:03Z
mined: 2026-09-09
mined-by: sonnet | opus
status: proposed # proposed | confirmed
person-verdict-on-source: confirmed | corrected
source-session: 5a85b6d3
source-pair: 5a85b6d3/003-coherence-check.md
in-tension-with: []
---
The body is prose. Not fields. Agents must answer the questions below and communicate like a person, not a structured record.
Below the frontmatter comes an H1 — one specific sentence naming the situation and the lean, never a category label. Everything after it is written the way a person tells another person something they learned.
Read the example item before you write one — examples/example-item.md in this skill's own folder. It's an illustrative, generic worked example (not real transcript data), showing the shape: the paragraph opens with the situation, the provenance arrives as a short story, code vocabulary is explained inside the sentence that needs it, and the limits of the lean are stated plainly at the end without a header announcing them. Once the person has their own first confirmed batch, replace this pointer with their own items as the reference standard — their own voice is the real target, not the example's.
Your item must ANSWER all of the following, and must NOT label any of them:
- What the lean actually is, concretely enough to act on.
- Why we hold it — the value or the past pain underneath it.
- Where it came from: when, in what session, what the agent had claimed, and the person's exact words. Open this differently every time.
- Where it might not apply, or what would make us revisit it.
- Any code name, branch, file, flag, constant or internal label, unpacked into plain meaning with its exact condition, woven into the sentence that needs it.
- If the person's quoted words contain dictation typos, say once, in passing, that they are reproduced exactly on purpose.
Banned in the body: a bold label followed by a period introducing a paragraph (**The lean.**, **Why.**, **Where it came from.**, **Variables unpacked.**), any Key: value line, and any structure repeated identically across items. Bold is for emphasis inside a sentence, never as a header.
Quotation marks mean the person said it. Do not put your own phrasing in quotes — the guard in Step 3b will catch it and it should. Reproduce their words exactly, typos and all. Never tidy their dictation.
Step 3b — Verify the quotes (MANDATORY, before the index)
node <this-skill-folder>/scripts/verify-instinct-quotes.js --domain <id> [--quarantine]
Every quoted span in an item's Where it came from section must match the person's words, in order, from that item's own source transcript pair. Punctuation and case are normalized; changed, added, removed or reordered WORDS fail.
Reproduce dictation exactly, typos and all. A cleaned-up paraphrase presented as a quote reads better and is exactly the thing that destroys the ledger's trust: it is only worth something if the words are really theirs. Per this project's own writing rule: "Never restate someone's words in 'better' form. Paraphrase closely or quote."
If a quote fails, restore the exact wording or drop the quotation marks and paraphrase. Never present cleaned-up dictation as a quote. An item that cannot pass this does not enter the ledger — run with --quarantine and it will be moved out automatically, with a note of what didn't match.
Step 3b-2 — An unverified item is not a finished item
A mining agent that stops before it runs the quote guard has not produced an instinct; it has produced an unverified draft sitting in the canonical ledger, indistinguishable from the verified ones. So whoever orchestrates a mining run owns the guard, not just the agent that was asked to run it — run verify-instinct-quotes.js --domain <id> --quarantine yourself after every batch, including batches whose agent never reported back. An instinct whose provenance cannot be trusted is worse than a missing instinct, because the next agent has no way to tell it apart from a real one.
Step 3c — Consolidate, when more than one agent mined at once
Parallel miners cannot see each other's items, so two of them will independently write the same lean from different sessions. After all miners finish, one agent reads every headline in the ledger and looks for pairs saying the same thing. For each pair, the person decides — do not merge silently. Bring them the two headlines and a recommendation. A pair that is genuinely one lean firing in two situations usually becomes one item whose headline names the shared move, with both sources under it and both situation slugs; a pair that only sounds alike stays as two items that reference each other.
Also re-run the quote guard once at the end, over the whole ledger rather than one agent's slice, because a merge rewrites prose and prose is where quotes live.
Step 4 — Index
node <this-skill-folder>/scripts/build-instinct-index.js --domain <id>
Regenerates <ledger>/README.md from item frontmatter: grouped by situation, each line = headline + strength + captured date + link. This is the map to substance an agent reads first. Never hand-edit the README — fix the item and regenerate.
Step 5 — Consume
The domain layer names the consumers (see below). Minimum for any domain:
- The domain's canonical entry point says, before anything else: read
<ledger>/README.md, then the items for the situation you are about to enter. - The task-start check for that domain gains a section: "Instincts I'm weighing" — which instincts apply to what I'm about to do, how I'm weighing them, and any I think don't apply here and why. The person sees whether the instincts fired right. Their reply is itself a new confirmed pair, which the next harvest mines. That is how it compounds without pigeonholing: the agent weighs, visibly; it does not obey.
- If the ledger for this domain is empty or doesn't exist yet, say so plainly — "no ledger yet (collecting)" — rather than omitting the section or inventing something to weigh.
Step 6 — Health (surfacing failure)
A loop that silently stops is worse than none — the person assumes it is working.
node <this-skill-folder>/scripts/extract-check-pairs.js --health [--write-health]
It reports on: sessions with checks vs. harvested, inbox age, the awaiting-verdict backlog, and capture integrity (a captured check body under 400 characters is a truncated capture, not a short check). --write-health writes the result where build-instinct-index.js will read it and prepend a warning line to the ledger README if anything's failing, so any agent that opens the ledger sees it.
A couple of the checks named in the original design of this loop need something this skill doesn't ship by default and should be treated as optional extras, not silently skipped as passing:
- Ledger growth / consumption (zero new items in 14 days, or few recent sessions containing "Instincts I'm weighing") — compute these yourself by comparing item
captureddates and grepping recent transcripts, if you want them tracked. - Findability — the original design checked this with a private semantic-search tool. If this project has institutional-memory search configured, use it to confirm a few item headlines are findable; otherwise, note that findability wasn't checked rather than reporting a pass.
Run health regularly (each session, or on whatever cadence you set) — a health check that's never re-run is exactly as useless as no health check, because the warning it shows will describe an old, smaller problem instead of the current one.
Domain layer contract
By default, everything above runs against a single built-in domain called general — pass no --domain flag, or --domain general explicitly, and it just works with no further setup. A general domain's situations are organized by what an agent is about to do, not by system component: change something a user will see, decide what to work on next, claim something is done, touch money or sign-in, write something to the person.
If you want more than one domain (say, one ledger for "customer support work" and a separate one for "infra work"), a domain layer supplies:
domain: short id used in instinct ids and--domainsession-filter: how to recognize the domain's sessions — pass it to--session-filteron the scripts above (matched against the project folder name and session id; extend the scripts yourself if you need matching by cwd or git branch instead)situations: the taxonomy of trigger situations for the domain, as slugs with one-line meaningsseed-sources: non-transcript files that already hold confirmed material for a one-time seed (verbatim intent logs, decision records, scar tissue)consumers: the entry-point files and skills that must say "read the instinct ledger first"
This can be as simple as a short markdown note next to the ledger listing those five things — there's no required file format.
Model tiers
Extraction: script, no model. Mining: a smaller/cheaper model for the first measured batch; escalate to a stronger one if the person judges the first batch's instincts insufficiently nuanced or for seeding from dense sources. The person: reads items and says yes/no; confirmed items move status to confirmed.
Starting with nothing yet
On a fresh install, there are no captured checks and no ledger. That's expected, not broken:
- Run the extractor once over existing history in case earlier check-style exchanges are already there:
node <this-skill-folder>/scripts/extract-check-pairs.js --since 90. - If nothing turns up, say so plainly: "there's nothing to mine yet — this fills in as you use checks and confirm or correct them." The "Instincts I'm weighing" section reports "no ledger yet (collecting)" rather than being silently dropped, and health reports "waiting for input" rather than "healthy."
- After a handful of sessions produce confirmed or corrected checks, run a first mining pass and bring the proposed instincts to the person a few at a time for a yes or no. Their approved batch becomes the reference standard for how future items should read — more useful than any outside example.
What would make this fail
- Items written as facts or rules. Test: does the item tell the agent what to do, or what we lean toward and why?
- Mining unanswered checks. The agent's claim without the person's reply is not an instinct.
- A summarizing extraction step. If anyone proposes "just have a cheap model pull the key points," that is the exact failure this method was built to avoid.
- The ledger nobody routes to. The consumers list exists so this is checked, not assumed.
- Silent loop death. The health check exists so a stopped loop is visible quickly — but only if it's actually re-run; a health check nobody re-runs is the same failure with extra steps.