← documentation docs/EXTRACT-AND-TEST-PLAN.md

Extract, then test: the plan

Written in September 2026 from the harness's statement of intent and its research. Where it infers something, it says so and gives a rough confidence ("inference, 80%"). Names written as /harness-setup or /harness-status are proposed names; nothing by those names exists yet.

What Jonathan asked for:

And it will need some degree of extract then test, so someone needs to plan all the intent for ux and design tests for that.

The plan has four parts:

  • Part A covers what a new person should experience, from first hearing about the harness to uninstalling it. Each item is written as a testable statement: "When {situation in the person's terms} → {what they experience}".
  • Part B covers how we will know each statement is true, and what each test can and cannot tell us.
  • Part C gives the order of extraction: what gets pulled out of Jonathan's setup first, and which tests have to pass before the next step starts.
  • Part D, a list of open questions and internal verification items (referenced below as D.1-N and D.2), is tracked separately and is not reproduced in this document.

0. Before reading

0.1 What this plan is built from

  • INTENT.md holds Jonathan's own words and is the only source of authority here. Every quote from Jonathan in this plan is copied from it character for character.
  • research/07 covers facts about the Claude Code plugin system: install, enable and disable, uninstall, how hooks merge, the plugin's data folder, userConfig, and claude plugin eval.
  • research/04 inventories the 37 registered hooks: what each one depends on, how each one fails open, the loop-breaking utilities, and the blockers to shipping (a hardcoded API key and .bak files).
  • research/05 is used only for what was tried and which instruments were used. Jonathan rejected its conclusions.
  • research/03 and research/06 cover the oracle roster and the existing scripts that turn a person's own ~/.claude/projects/*/*.jsonl transcripts into oracles.
  • data/items.json and docs/skills/: when this was written, 1 of 458 skill items had an evaluator page (align, "include after rework"). The other 457 are unreviewed.

Because of that, Part A describes the journey by fulcrum and by setup checkpoint, not skill by skill. Which skills land at each fulcrum gets filled in as evaluator pages arrive. I expect the intents in Part A to hold whatever the final skill list turns out to be (inference, 80%). Where a fulcrum depends on a specific piece, I name the piece as it exists in Jonathan's setup today.

0.2 Terms used here (each defined once)

  • Harness: Jonathan's collection of hooks, skills and protocols that he runs over Claude Code.
  • Plugin: the Claude Code packaging format that lets someone else install the harness.
  • Hook: a script Claude Code runs automatically at a lifecycle event. The events that matter here are session start, a message arriving, before a tool runs, after a tool runs, a subagent starting or stopping, before compaction, the agent trying to stop, and session end.
  • Skill: a written procedure the agent loads and follows, invoked like /align.
  • Fulcrum: Jonathan's word for a point in a session where an agent can go off course and the error then compounds. In his words, the harness "literally operates upon every fulcrum that is known from the beginning of a session to the end to reduce agent misalignment that would otherwise cascade and compound into misaligned work".
  • Gate: a hook that can block the agent. Examples are "you can't edit until you've stated your read on intent" and "you can't stop while claiming it works without evidence".
  • Fail open: when something a hook needs is missing or broken, the hook lets the agent proceed instead of blocking or crashing.
  • Downgrade: after a gate has blocked a set number of times in one session, it switches from blocking to reminding. Jonathan's governer gate does this today after 3 blocks.
  • Institutional memory: what the harness knows from the person's past sessions. That means their words, their corrections, decisions they made, and intent they confirmed. In Jonathan's setup this lives behind agent_find and his local IX API (his product's server on his machine).
  • Oracle: a Google NotebookLM notebook loaded with material mined from the person's history, which agents query before acting. nlm is the command-line tool that talks to NotebookLM.
  • Intent map: a nested description of what the person is working toward (this task → what it serves → what that serves), where each item is marked as confirmed or still a guess. This is the concrete form of what Jonathan calls "the process of building that synthetic understanding of the person's intent and so on to be part of that workflow".
  • Checkpoint: one step of guided setup that the person can accept or decline.
  • Switch / dial: a switch turns something on or off. A dial sets how strongly it acts. Jonathan asked for these by name: "like an on-off switch, but also like maybe in some cases like a dial".
  • Plugin data folder: ~/.claude/plugins/data/<id>/ (${CLAUDE_PLUGIN_DATA}). It survives plugin updates and is deleted on uninstall unless the person passes --keep-data (research/07).
  • Waiting on setup: the state of a piece whose setup checkpoint was declined or not yet reached. The piece says so plainly at the moment it would have acted.
  • Status check: a proposed command, /harness-status, that lists every piece and its state.
  • Clean room: a throwaway home folder, pre-filled to look like someone's real setup, used so that tests can prove install and uninstall leave that setup untouched.
  • Baseline: the same test run without the harness, for comparison.
  • Episode / fork: defined in B.7. An episode is a multi-turn task. A fork is a planted point in it where a wrong assumption is tempting and the right answer lives in accumulated context.

0.3 How I read the intent for this plan (my reading, stated so Jonathan can correct it)

  1. The full experience by default. Jonathan: "the main error you could make is assuming that you should remove something or dumb it down or inhibit it um, as a default." Also: "any assuming that we should remove features is probably wrong unless there's a damn good reason." I read this as follows. Every piece is in one of three states: working, being built by setup, or saying plainly that it is waiting on setup and what setup would give it. A piece that is silently absent is not an allowed state. That includes a piece that fails open so quietly the person never learns it didn't run.
  2. Setup builds what's missing and tells before it acts. "So if it relies upon a, a, a map of intent that doesn't exist, then we probably need to build that map of intent and so on." "So the command should effectively, like, build the system. It should extract the key information. For the different kinds of oracles, it should build the notebooks." "it should tell them what it's going to do during setup".
  3. Options at each checkpoint, dials instead of removal. "giving them options at each checkpoint, right, during setup", "a Boolean config thing with comments that say what each thing does", and a dial wherever something "might do X too many times".
  4. One command to install. "it should probably be one command to install it."
  5. Per-piece docs an agent can use. "It's going to need to be able to look it up and understand how to set it up if it's not set up and understand if it's set up. or if it's not, understand if it's working or if it's not." And: "it can't just be one massive 400 page document."
  6. A test speaks only for itself. A test that fails to show an effect is "a single test failed to produce a deterministic output that's all we learned", and "the modality of how you're going about trying to prove whether it's true or false" is one of the variables.

A few places might be what Jonathan calls "a damn good reason", but I am not deciding them. They are tracked as open questions (see "What's still open" at the end of this document):

  • The CLAUDE.md A/B hook overwrites the person's CLAUDE.md.
  • Some hooks are specific to Jonathan's products: the Vercel safety gate, cloud provisioning, and the founder port guard. My default is to generalize them, not drop them. For example, the port guard becomes a "reserved ports" setting.

Part A: UX intents for a new person

Each statement is one intent: When {situation} → {what they experience}. The IDs are for cross-reference from Part B. Where a statement rests on my inference and not on Jonathan's words or verified plugin facts, it is marked.

A.1 Finding out about it

  • UX-DISC-01. When someone lands on the repo or website → before any install instructions, they read in plain language what the harness is for: that it acts at every fulcrum of a session, with the thesis that misalignment compounds at each one and a map of those fulcrums.
  • UX-DISC-02. When they want to know what it costs → they are told up front that it adds context at every message and every session start. research/04 measured Jonathan's current harness at roughly 3,000–3,900 tokens on a loaded turn plus 2–4 KB at session start; this will be re-measured on the plugin. They are also told why that trade is made ("What they care about is results, right? And agents not going off the rails.") and that dials exist.
  • UX-DISC-03. When they want to know what it touches → one page lists every place it reads, every place it writes, and every outside service it can talk to. Each is marked core or optional, and for each they can see whether anything leaves their machine.
  • UX-DISC-04. When they want to know what is claimed and on what basis → the thesis and Jonathan's direct experience are stated as such. Any measurements are shown separately, each with what it can and cannot show. (Jonathan has said the thesis can be stated without proof. This intent is only that the two are not blurred together.)
  • UX-DISC-05. When they decide to try it → there is one install command they can copy.

A.2 The one-command install

  • UX-INST-01. When they run the install command → the harness is installed and enabled, and nothing they already had is changed: their CLAUDE.md, their hooks in settings, their skills, their MCP servers, and their project .claude folders.
  • UX-INST-02. When the install finishes → they see what works right now, before any setup. They see that their next session will offer guided setup and that nothing further will happen to their files until they say yes.
  • UX-INST-03. When they already have a skill or command with the same name as one of the harness's (their own reflect, for example) → both keep working, theirs is not replaced, and they are told about the overlap. (Inference, 70%, that plugin skills are namespaced. To be verified; see D.2.)
  • UX-INST-04. When they install at project scope instead of user scope → the harness is active only in that project. Other projects are untouched.
  • UX-INST-05. When they install on Linux → they get the same experience as on macOS.
  • UX-INST-06. When an install is interrupted, for example by a network drop → they end up with either the whole plugin or none of it, and a message saying which.
  • UX-INST-07. When a newer version comes out and they update → their configuration, their built memory and their oracle registry carry over. New switches appear set to their full-strength defaults, with comments, and they are told what is new.

A.3 Guided setup: the shape of every checkpoint

  • UX-SETUP-01. When they open their first session after installing → the harness introduces itself once and offers guided setup. It does not read their history or write anything until they say go. They can also start setup later with one command (placeholder /harness-setup).
  • UX-SETUP-02. When setup begins → they see every checkpoint up front. Each gets one sentence on what it builds, whether it is optional, and roughly how long it takes.
  • UX-SETUP-03. Before any checkpoint acts → they are told:
    • (a) what it will do;
    • (b) exactly what it will read, with real paths and counts, e.g. "412 session transcripts under ~/.claude/projects, 1.3 GB, March–September";
    • (c) what it will write and where;
    • (d) what, if anything, will leave their machine and to which service.
  • UX-SETUP-04. When they decline a checkpoint → nothing from that checkpoint is written or sent, and everything else keeps running. The pieces that depend on it say "waiting on setup: " at the moment they would have acted.
  • UX-SETUP-05. When they come back to setup later → it shows which checkpoints are done, declined, or not started, and picks up where they left off. Finished checkpoints are not redone unless they ask.
  • UX-SETUP-06. When a checkpoint is interrupted → re-running it continues or cleanly restarts it. No half-built result is ever treated as finished by a piece that depends on it.
  • UX-SETUP-07. When a checkpoint reports success → the result has been checked by reading it back, not trusted from an exit code, and they are shown that check. Examples are the notebook's own list of sources, a test query, or a count read from the store.
  • UX-SETUP-08. When setup ends → they get a summary of every piece's state (working / waiting on setup / switched off), where the config file lives, and how to check status at any time.

Checkpoint 1: Core (needs nothing outside their machine)

  • UX-SETUP-CORE-01. When they accept the core → the pieces that need no outside service start working. That means the per-message intent check, the evidence check before "done", the /ask routing for questions, the follow-through on /commands they name, the subagent alignment checks, the capture before compaction and re-injection after it, and the cleanup of leftover processes at session end.
  • UX-SETUP-CORE-02. When they are shown the core → they are told which agent actions can be blocked and why, how each gate lets go (it downgrades to reminding after a set number of blocks), and which dial controls that.
  • UX-SETUP-CORE-03. When they already have their own CLAUDE.md → the harness adds its operating protocol without editing their file and tells them where the protocol text lives so they can read and change it. (Inference, 80%: delivered as injected context or a plugin-owned file.)
  • UX-SETUP-CORE-04. When a piece needs a local helper to stand in for services that exist only on Jonathan's machine (his IX API scorer and his compaction store) → the helper runs locally. They are told whether it is a background process and, if it is, how to stop it. It stops at uninstall.

Checkpoint 2: Institutional memory built from their own history

  • UX-SETUP-MEM-01. When they reach this checkpoint → they are shown how much history exists (sessions, projects, date range, size). They are told the harness will build a searchable memory from it entirely on their machine.
  • UX-SETUP-MEM-02. When they want to exclude something → they can exclude projects, folders or date ranges before anything is read. Excluded transcripts are never opened.
  • UX-SETUP-MEM-03. When their transcripts contain secrets such as API keys, tokens or passwords → those are not copied into memory, and they are told how many were withheld. (Inference, 85%, that they need this. The transcripts are raw session logs.)
  • UX-SETUP-MEM-04. When extraction runs on a large history → it runs in the background, can be resumed, and does not stop them using Claude Code meanwhile.
  • UX-SETUP-MEM-05. When extraction finishes → they see what was built, each count read back from the store. Examples: their own messages kept word for word, corrections they made to agents, intent statements they confirmed, and sequences of skills used.
  • UX-SETUP-MEM-06. When they want to check what it learned → they are shown a sample (e.g. "the corrections you've made most often") and can reject or delete items before any agent uses them.

Checkpoint 3: Building the synthetic understanding of their intent

  • UX-SETUP-INTENT-01. When they reach this checkpoint → the harness offers to build their intent map: what they are working toward, their projects, and what each serves.
  • UX-SETUP-INTENT-02. When it builds the map → it drafts from their history first and presents its readings with certainty numbers for them to confirm or correct. It does not interview them from a blank page. This follows /align's method: predict, research, then present for confirmation.
  • UX-SETUP-INTENT-03. When they confirm some items and not others → only the confirmed ones count as confirmed. The rest stay marked as guesses, and every later agent treats them as guesses.
  • UX-SETUP-INTENT-04. When they have little or no history → it asks a small number of questions that cut uncertainty most. The map starts small and says it is small.
  • UX-SETUP-INTENT-05. When they want to see or change their intent map → it is a readable file they can open and edit, and later sessions use their edits.
  • UX-SETUP-INTENT-06. When their history shows how they want agents to talk to them (their own corrections about communication) → the harness drafts their communication guide from their own words, quoted exactly, for them to confirm. In Jonathan's setup this role is played by how-to-talk-like-the-founder. What ships as the default before they have one is open question D.1-4.

Checkpoint 4: Oracles (Google NotebookLM), optional

  • UX-SETUP-ORACLE-01. When they reach this checkpoint → they are told in one sentence what an oracle is and which oracles will be built. For each: what question it answers and what goes into it. They are told that the sources will be uploaded to NotebookLM under their Google account, and that the whole step is optional.
  • UX-SETUP-ORACLE-02. When they don't have nlm → setup walks them through installing it, or installs it with their consent, and checks that it runs.
  • UX-SETUP-ORACLE-03. When Google sign-in is needed → they sign in themselves in their browser, and the harness never handles their password. Setup waits, then confirms sign-in with a harmless call.
  • UX-SETUP-ORACLE-04. When they have more than one Google account → they choose which one holds which oracle, and setup records that choice. No agent is ever left to assume which account holds an oracle. Jonathan's oracle-locate.sh exists because of exactly this.
  • UX-SETUP-ORACLE-05. Before anything is uploaded → they see the exact files, with names, sizes and a preview, after secrets have been withheld. They approve the upload.
  • UX-SETUP-ORACLE-06. When the notebooks are built → each is checked by reading its own source list back and running a test question, and they are shown the results. One registry file records every oracle's ID, account, and what it was built from. Today Jonathan has three drifting rosters (research/03, 06); the new person gets one.
  • UX-SETUP-ORACLE-07. When NotebookLM limits are hit (a source-count ceiling, a per-source size limit, or nlm source add reporting success while creating nothing) → setup notices from the readback. It packs sources to fit or says plainly what did not fit, and it never reports the checkpoint as done.
  • UX-SETUP-ORACLE-08. When they have a history of back-and-forth with agents → setup offers to build a predictor of their responses, the same class as Jonathan's jonathan-check2. Agents can then check "what would this person say to this?" before bringing work to them.
  • UX-SETUP-ORACLE-09. When they decline oracles → each oracle-using step says "oracle not set up, skipped" when it would have run, with a dial for how often it says so. Everything else runs.

Checkpoint 5: Optional connections to how they already work

  • UX-SETUP-OPT-01. When they use Obsidian → setup offers to put intent seeds and the decision ledger in their vault. Otherwise these go in a plain folder they choose.
  • UX-SETUP-OPT-02. When they keep something running on a port (a dev server, for example) → they can list ports agents must never bind. This generalizes Jonathan's founder port guard.
  • UX-SETUP-OPT-03. When they have branches that must not take two merges at once → they can list them for the merge lock. Today that list is hardcoded to Jonathan's branch names.
  • UX-SETUP-OPT-04. When they want to A/B test changes to their own agent setup → the experiment tools are available and stay idle until they start an experiment. The CLAUDE.md-swapping part must never overwrite their own file (see D.1-1).

A.4 Living with the memory, including day one with thin history

  • UX-MEM-01. When an agent searches memory on a topic where little exists (day one, or a new area) → the result says how much it has, e.g. "2 related items from 1 session". When it has nothing, it says "I have little on this" and invents nothing. The agent then says it is working without institutional memory on this point and states the assumption it is making, with a certainty number.
  • UX-MEM-02. When an oracle was built from few sources → its answers come labeled as thin. Answers that don't rest on the notebook's sources are flagged as ungrounded.
  • UX-MEM-03. When something in memory conflicts with what they just said → their current words win, and the agent names the conflict out loud.
  • UX-MEM-04. When a session ends → what it produced goes into memory without them doing anything: confirmed intent, corrections, decisions, unfinished items, and the reasons behind skill choices.
  • UX-MEM-05. When they correct an agent → the next agent in a similar situation can find that correction and weigh it. In Jonathan's setup this is the instinct loop: extract-check-pairs.js → /instinct-harvest.
  • UX-MEM-06. When new material has piled up since an oracle was last refreshed → they are told it is stale (last refresh date against how much is new). Refresh runs on a schedule or threshold they set, and each refresh is checked by readback.
  • UX-MEM-07. When they want to see what the harness "knows" about them → they can browse it, and correct or delete any item.

A.5 The first session, fulcrum by fulcrum

Before the first message (session start)

  • UX-F1-01. When a session opens → the agent starts with the harness's operating protocol loaded and whatever carries over for this project, such as unfinished work or last session's scope. The person doesn't have to re-explain.
  • UX-F1-02. When some pieces are waiting on setup → there is a one-line notice at session start naming them, with a dial for every session / once a day / never.
  • UX-F1-03. When a session opens → startup is not noticeably slower. The added startup work stays inside a stated time budget.

A message arrives

  • UX-F2-01. When they send a message → before acting, the agent states its reading of their intent, keeping what they said separate from what it thinks that means, with a certainty number. A misreading is caught in one turn instead of many. In Jonathan's setup this is the coherence check.
  • UX-F2-02. When the message is a short acknowledgment ("ok", "thanks") → there is no ritual.
  • UX-F2-03. When they name a /command in their message → the agent runs it. It cannot end the turn having quietly skipped it.
  • UX-F2-04. When the task is substantial → it is scored for how much is riding on it and routed to the matching depth of checks. At higher scores that means testable scope statements and a plan before building. In Jonathan's setup this is the governer. The person can see the score and the reasons for it.
  • UX-F2-05. When they send a message → their words are kept word for word for the session, so a later compaction can't paraphrase them away.

Before the agent acts

  • UX-F3-01. When the agent tries to edit, write or run something before stating its reading of their intent → it is stopped and asked to do that first. Reading, searching and consulting memory or oracles are never blocked.
  • UX-F3-02. When the agent starts searching code before searching the person's own memory → it is sent to memory first, once per session.
  • UX-F3-03. When the agent wants to ask the person a question → it first predicts the answer and weighs whether the question is worth their time (/ask). It asks only when going ahead on its best guess would be unsafe or would make the work useless if the guess were wrong. When it does ask, it includes its prediction.
  • UX-F3-04. When the agent tries to bind a reserved port or merge into a locked branch → it is blocked and told why.

While working

  • UX-F4-01. When the agent uses a skill → it notes in one line why it chose it. Those notes feed the memory of which skills tend to follow which, used to predict the steps a task needs.
  • UX-F4-02. When the agent is taking longer than expected → it checks whether its current framing is the problem (/expand-perspective) before handing the problem back to the person.
  • UX-F4-03. When the scope changes mid-session → the agent re-states the new scope before continuing.
  • UX-F4-04. When the agent is about to build on a belief a lot depends on → it checks that belief against reality first, by running the command, reading the file or querying the data.

Subagents

  • UX-F5-01. When the agent dispatches a subagent → the subagent receives the relevant scope and the person's intent, and is required to state the experience it is building toward before it writes code.
  • UX-F5-02. When a subagent claims something "works" with no evidence → it is sent back to verify, the same as the main agent.
  • UX-F5-03. When a subagent produces a research report or plan → it is saved into memory, so good work isn't lost when the subagent ends.

Compaction

  • UX-F6-01. When Claude Code shortens a long conversation → the person's exact messages and confirmed intent are saved first and handed back afterward. The agent does not come back with a paraphrased or lost understanding, and it shows that it re-read them.

Claiming "done"

  • UX-F7-01. When the agent says something works, is fixed or is live, and there is no evidence of that in the session → it cannot stop yet. It is told to go verify, and shown the specific checks it committed to at the start if there were any.
  • UX-F7-02. When the evidence is there → nothing blocks.
  • UX-F7-03. When the "done" check has already blocked once on this stop → it lets go. There is no loop.
  • UX-F7-04. When the agent finishes substantial work → it is reminded of the completion steps. If the skill-sequence oracle is set up, it is also shown a prediction of which closing steps this kind of work usually needs; if not, it is told that prediction is waiting on setup.
  • UX-F7-05. When the person's own response predictor exists → the agent checks its work against "what would this person say" before presenting it.

Session end

  • UX-F8-01. When the session ends → its intent, unfinished items, corrections and checks are saved locally without holding up the exit, and any leftover helper processes it started are cleaned up.

The next session

  • UX-F9-01. When they open a new session on the same project → the agent can find what was decided, what is unfinished, and what they corrected last time, and it cites where each came from.
  • UX-F9-02. When they ask "what's left?" → the answer comes from what was recorded. A missing record is not read as a sign of unfinished work.

A.6 Tuning: switches and dials

  • UX-DIAL-01. When they open the config → it is one file. Every switch is on/off with a comment saying what it does, what it costs, and what depends on it. Every dial states its range and what each end means.
  • UX-DIAL-02. When they have changed nothing → every piece whose setup they accepted runs at full strength.
  • UX-DIAL-03. When they change a setting → it takes effect at the next message or next session (the comment says which) without reinstalling, and the status check reflects it.
  • UX-DIAL-04. When a gate fires more than they want → they can set its strength (off / remind only / block) and how many blocks it allows before it drops to reminding.
  • UX-DIAL-05. When the added context per message feels heavy → they can see the size of each injected block and turn any one down or off individually.
  • UX-DIAL-06. When they enter an invalid value → that setting falls back to its default with a visible note. No hook crashes.
  • UX-DIAL-07. When they want one project to behave differently → a project-level config overrides the user-level one for that project. (Inference, 75%, following Claude Code's own scope order.)

A.7 An agent finds a piece's docs

  • UX-DOCS-01. When an agent or person needs to know about a piece → there is one page for it at a predictable path, reachable from an index and short enough to read in one pass.
  • UX-DOCS-02. When they read that page → it says what the piece is for, how to set it up, how to tell it is set up (a command and its expected output), how to tell it is working (signs of healthy and broken), what it depends on, what depends on it, and its switches and dials.
  • UX-DOCS-03. When an agent runs a page's "is it set up?" check → it gets a clear yes or no, with the reason.
  • UX-DOCS-04. When an agent runs a page's "is it working?" check → it gets evidence of the last time the piece actually ran (when, and what it did), not just "installed".
  • UX-DOCS-05. When a piece is broken → its page names the likely causes and the fix for each.
  • UX-DOCS-06. When they run the status check → they see every piece as on / off / waiting on setup / working / broken, with the last time it ran.
  • UX-DOCS-07. When the code changes → the docs' checks still give correct answers, because the docs are tested.

A.8 Turning it off

  • UX-OFF-01. When they disable the plugin → from the next event on, no harness hook runs and no gate can block. Their own hooks keep running, and their data and config are kept.
  • UX-OFF-02. When they re-enable it → it comes back in the same state, with no setup to redo.
  • UX-OFF-03. When they switch off one piece → only that piece stops, and the pieces that rely on it say so.
  • UX-OFF-04. When they want it off for just this session → there is a way to do that without touching config. (Inference, 70%, that this is wanted.)

A.9 Uninstalling

  • UX-UNIN-01. When they uninstall → everything the harness wrote locally is gone: plugin files, data folder, memory store, logs, temp files, and background processes. Every file of theirs is byte-for-byte what it was before install.
  • UX-UNIN-02. When they want to keep what they built → they can keep the data (--keep-data) and are told where it is.
  • UX-UNIN-03. When NotebookLM notebooks were built → uninstall tells them these live in their Google account and lists them. It offers to delete them, with their confirmation, or leaves them.
  • UX-UNIN-04. When they asked setup to write outside the data folder (seed notes in their Obsidian vault, for example) → uninstall lists those files and asks. It never silently deletes their notes.
  • UX-UNIN-05. After uninstall → no hook fires, and nothing in their settings refers to the harness.

A.10 It never breaks their setup and never traps them

  • UX-SAFE-01. When anything the harness relies on is missing or down (nlm, the local store, the scorer, the network) → the hook lets the agent carry on within a set time. It says once what is degraded; it is not silent.
  • UX-SAFE-02. When a hook itself hits an error (bad input, a missing tool such as jq or node) → it exits cleanly, the session continues, and the error is logged where the status check can show it.
  • UX-SAFE-03. When a gate has blocked the set number of times in a session → it drops to reminding. The agent is never left unable to act or unable to finish.
  • UX-SAFE-04. When the "done" check fires again on a stop it has already blocked → it lets the stop through.
  • UX-SAFE-05. When they run two sessions at once → each session's gates track only that session. One session's check never unblocks or blocks the other.
  • UX-SAFE-06. When their own hooks listen to the same events → both run, and the harness neither swallows nor alters what their hooks produce.
  • UX-SAFE-07. When hooks run on every message → each stays within a time budget, so their prompt never feels stalled.
  • UX-SAFE-08. When they run Claude Code without a person present (claude -p, CI) → no gate waits forever. The session finishes within its time limit.
  • UX-SAFE-09. When they inspect what shipped → it contains none of Jonathan's keys, paths, notebook IDs, company details or personal content.

Count: 116 UX intent statements. Discovery 5 · install 7 · guided setup 37 (general shape 8, core 4, memory 6, intent 6, oracles 9, optional 4) · living with memory 7 · first-session fulcrums 28 (F1 3, F2 5, F3 4, F4 4, F5 3, F6 1, F7 5, F8 1, F9 2) · dials 7 · docs 7 · off 4 · uninstall 5 · safety 9.


Part B: Test design

B.0 Rules every test follows

  • Test the thing running, not the code. Hooks run with real event input. Sessions run for real, without a person present (claude -p). Stores are read back. Reading a script and concluding it works is not a test.
  • Every test states what it can tell us and what it cannot. A test that fails to show an effect tells us about that test. The modality, the tasks, the grader, the sample size and the model's randomness are all variables. Jonathan: "a single test failed to produce a deterministic output that's all we learned".
  • Model behavior is probabilistic. Anything that depends on what the agent chooses is run several times, and we report a rate, not a single pass or fail.
  • macOS and Linux both. Every layer except the real-NotebookLM run is in the continuous test matrix for both.

B.1 Which tests cover which intents

Intent group Tests
DISC (5) B.6 docs tests (site pages carry the fulcrum map, costs and data-flow page); a human read-through (T-HUMAN)
INST (7) B.2 clean-room tests T-CR-1, 4, 5, 6, 7, 8
SETUP general (8) B.3 setup-step tests T-SU-1 to T-SU-7
SETUP-CORE (4) T-CR-2, B.5 fulcrum tests, T-SU-3 (no egress)
SETUP-MEM (6) T-SU-3 (egress and secret canaries), T-SU-5, T-MEM-*
SETUP-INTENT (6) T-SU-7, T-INTENT-*
SETUP-ORACLE (9) T-SU-4, T-SU-8, T-SU-9, fake-nlm suite
SETUP-OPT (4) T-SU-2 (decline), fixture tests for the port and branch dials, T-CR-2 (the experiment scaffold never writes CLAUDE.md)
MEM (7) B.5.2 memory honesty tests
F1–F9 (28) B.5.1 fulcrum mechanism tests; B.7 measures whether they matter
DIAL (7) B.6 config tests
DOCS (7) B.6 docs tests, including the agent-with-only-the-docs test
OFF (4) T-CR-9
UNIN (5) T-CR-3
SAFE (9) B.4 failure-mode and gate-loop tests; T-CR-4 static scan

B.2 Clean-room install tests

The fixture: a lived-in home folder. Each run starts with HOME=$(mktemp -d), pre-filled to look like a real person's setup:

  • ~/.claude/settings.json with the person's own hooks on SessionStart, UserPromptSubmit, PreToolUse and Stop. Each hook appends a line to a log, so we can prove it still fires.
  • ~/.claude/CLAUDE.md with known content.
  • ~/.claude/skills/reflect/SKILL.md, which collides with a harness name, and ~/.claude/skills/mine/SKILL.md, which doesn't.
  • An MCP server entry.
  • A project folder with its own .claude/settings.json hooks and its own CLAUDE.md.
  • Synthetic transcripts under ~/.claude/projects/<proj>/*.jsonl.

Snapshot before install. For every file under the fixture home and project we record its sha256, mode and symlink target. We also record the running processes, the harness-pattern files in /tmp and $TMPDIR, and scheduled jobs (crontab -l; launchctl list on macOS, systemctl --user list-timers on Linux).

  • T-CR-1 Install touches nothing of theirs. Install from a local marketplace path with the one command.
    • Pass:
      • Every file that existed before is byte-identical, with one exception: in settings.json the diff may only touch keys the plugin system itself owns. We believe those are keys like enabledPlugins, pluginConfigs and marketplace entries; D.2 verifies the list.
      • Nothing new appears outside ~/.claude/plugins/.
      • In a session after install, the person's own marker hooks still write to their log, and the harness's hooks also fire.
    • Can tell: install leaves their files alone, and their hooks coexist with ours.
    • Cannot tell: whether setup steps leave their files alone (that is T-CR-2), or whether their hooks' output reaches the model unchanged (that is T-SAFE-06).
  • T-CR-2 Setup writes only where it said it would. Run setup twice, once accepting every checkpoint and once declining every checkpoint, with a file watcher on.
    • Pass: the only writes are to the plugin data folder, plus any locations the person chose at a checkpoint. CLAUDE.md is never written.
    • Can tell: setup keeps to its declared write list.
    • Cannot tell: anything about later sessions (T-CR-3 covers those).
  • T-CR-3 Uninstall leaves nothing behind. Install, run setup, run three sessions, then uninstall. Compare against the pre-install snapshot.
    • Pass:
      • The diff is empty. If Claude Code itself rewrites the formatting of settings.json, that is recorded separately as upstream behavior, not ours.
      • No harness processes are left running.
      • No harness files remain in /tmp or $TMPDIR.
      • No scheduled jobs remain.
      • A session afterwards fires no harness hook.
      • With --keep-data, the only difference is the data folder.
    • Can tell: nothing local is left behind.
    • Cannot tell: what happens to remote notebooks. That is a separate check that uninstall lists them (UX-UNIN-03).
  • T-CR-4 Linux, plus a portability scan.
    • Run T-CR-1 to T-CR-3 in two containers: a common Linux distribution, and a stripped image without jq, for the failure paths.
    • Statically, run a shell linter over every hook. Search for macOS-only command forms (BSD sed -i '', stat -f, date -j), for /Users/, /private/tmp, Jonathan's notebook IDs, "ixcoach", localhost:3002, and his names. Run a secret scanner over the package.
    • This is also where the known blockers get caught: the hardcoded API key in user-prompt-enrich.sh, the .env read in proposal-auto-close.js, and the .bak/.backup files (research/04).
    • Can tell: it installs and runs on those images, and the listed leaks are absent.
    • Cannot tell: that every Linux distribution works, or that unlisted personal content is absent. A human skim of the package covers that.
  • T-CR-5 Project scope. Install at project scope in project A.
    • Pass: a session in project B shows no harness activity.
  • T-CR-6 Name collision. With the person's own reflect skill present, invoke /reflect.
    • Pass: their skill runs as it did before install, and the harness's version is reachable under its own name.
  • T-CR-7 Interrupted install. Kill the install partway through.
    • Pass: the result is either complete or absent. If Claude Code's own install turns out not to be atomic, we record that as upstream behavior and add a check at the first session that notices a partial install and says so.
  • T-CR-8 Update. Install version N, run setup, change two dials, then update to version N+1.
    • Pass:
      • Their dial values, memory and registry carry over.
      • New switches appear at their defaults, with comments.
      • They are told what is new.
  • T-CR-9 Disable and re-enable. Disable the plugin, then run a session built to trigger every gate.
    • Pass:
      • No harness hook fires.
      • Their own hooks still fire.
      • After re-enabling, the status check matches the state before disabling.
      • A single-piece switch-off stops only that piece, and its dependents report "waiting on" it.

B.3 Setup-step tests

Fixtures:

  • A fake nlm placed first on PATH. It logs every call and can be scripted to succeed, fail sign-in, hang, return success while creating nothing (the known false-success trap), or hit a source limit.
  • Four histories:
    • rich: hundreds of synthetic sessions for a fictional person;
    • thin: 2 sessions;
    • empty: none;
    • secret-laden: transcripts containing canary secrets, meaning fake keys and passwords with unique strings we can search for.

Tests:

  • T-SU-1 Tell before acting. Drive setup through a script. At each checkpoint, capture the text shown before the person says yes.
    • Pass: the text names what it will do, what it will read (with counts matching the fixture), what it will write, and what leaves the machine. The file watcher and network log show zero writes and zero outbound calls before the "yes".
    • Can tell: the disclosure happens and comes first.
    • Cannot tell: whether a real person understands it. T-HUMAN covers that.
  • T-SU-2 Declining. Decline each checkpoint in turn.
    • Pass:
      • Nothing is written for that checkpoint, and nothing is sent.
      • The status check shows its dependents as "waiting on setup: ".
      • A session built to trigger each dependent shows that notice at the moment it would have acted. It is not silent.
      • Every other piece still works.
  • T-SU-3 What leaves the machine. Run setup behind a proxy or network namespace that logs every outbound host and payload.
    • Pass:
      • Declining everything → no outbound traffic other than Claude's own model calls.
      • Accepting only memory → the same.
      • Accepting oracles → only the NotebookLM/Google hosts.
      • Across all runs, no canary secret appears in any outbound payload or anywhere in the memory store.
    • Can tell: the declared data flow holds for these fixtures.
    • Cannot tell: that the withholding catches every kind of secret. The secret patterns are listed, and a missed pattern is a finding.
  • T-SU-4 The false-success trap. The fake nlm returns success from source add but creates nothing.
    • Pass: setup reports the failure, names what didn't upload, and does not mark the checkpoint done. The same test runs for "limit reached" and "signed out".
  • T-SU-5 Rerun and resume.
    • Run each checkpoint twice → no duplicate notebooks, sources or memory items.
    • Kill memory extraction halfway and rerun → the final store's counts equal those of an uninterrupted run.
  • T-SU-6 Proof shown. Every checkpoint that reports "done" is followed in the transcript by its readback: a count read from the store, the notebook's own source list, or a test query and its answer.
  • T-SU-7 Thin and empty history.
    • On thin and empty: the intent checkpoint switches to asking a few high-value questions; memory reports its small size; the oracle checkpoint says what it would be built from and that it will be thin. Nothing claims more than exists.
    • On rich: drafts come with certainty numbers, and nothing is marked confirmed until the scripted person confirms it (T-INTENT-1).
  • T-SU-8 Several Google accounts. The fake nlm exposes two profiles.
    • Pass: setup asks which account to use, and the registry records the account for each oracle.
  • T-SU-9 One real run. On a throwaway Google account, with the real nlm, build the oracles from the rich fixture. Check them through NotebookLM's own listing and a test query.
    • Can tell: the real path works on this date.
    • Cannot tell: that it keeps working as NotebookLM changes. That is why the fake-nlm suite carries the regression load.
  • T-HUMAN Setup read-through. Someone who has never seen the harness installs it and goes through setup, thinking aloud.
    • Pass: at every checkpoint they can say, in their own words, what it will do and what leaves their machine.
    • Can tell: where the words fail.
    • Cannot tell: that the words work for everyone. Repeat with a few different people.

Memory and intent specifics:

  • T-MEM-1 The counts shown at the end of extraction equal an independent count computed from the fixture.
  • T-MEM-2 Excluded projects: a file-access log shows their transcripts were never opened.
  • T-INTENT-1 An intent item the scripted person didn't confirm never appears as confirmed in the intent map, and a later session's agent refers to it as a guess.
  • T-INTENT-2 Their edits to the intent map file show up in the next session's reading of their intent.

B.4 Failure-mode and gate-loop tests

The failure matrix, at the hook level (fast and deterministic). Every hook is run directly with recorded event input under each failure:

  • the tool it depends on is missing;
  • the service it depends on hangs (a port that accepts and never answers);
  • nlm is signed out;
  • its state file is corrupt;
  • the data folder can't be written;
  • the event input is malformed;
  • the session ID is missing;
  • the disk is full (a small temporary disk).

Pass for each hook and failure:

  • The exit code doesn't block.
  • Wall time stays inside the budget (proposed: 2 s for hooks that run in line; background work detaches).
  • There is one visible notice per session per kind of failure.
  • The error is logged in the data folder for the status check.

We also fuzz the event input with random and truncated JSON.

  • Can tell: these failures fail open, loudly, and quickly.
  • Cannot tell: failures we didn't think of. The fuzzing narrows that gap without closing it.

Gate-loop tests.

  • T-LOOP-1 An agent that won't comply. Feed the PreToolUse gates a scripted sequence in which the agent never produces the intent check.
    • Pass: blocks stop at the configured number, the gate drops to reminding, and the action then goes through. Repeat at dial values 1, 3 and "off".
  • T-LOOP-2 Stop firing again. Send a Stop event with stop_hook_active true.
    • Pass: immediate exit, no block.
  • T-LOOP-3 A real session that keeps claiming done. Run a real session without a person present, with a task that invites a "done" claim that can't be verified. An example is asking it to confirm a live website works when there is no network.
    • Pass: the session ends within K turns. We record the longest run of consecutive blocks.
  • T-LOOP-4 Gates that could deadlock each other. Arm the intent-check gate, the governer gate and the memory-first gate together.
    • Pass: none of them blocks the action another one requires. For example, the memory-first gate demands a memory search, and the intent-check gate must let that search through; research/04 says search tools always pass, and this test checks it on the running system.
  • T-LOOP-5 Two sessions at once. Run two real sessions concurrently.
    • Pass: one session's intent check doesn't unblock the other, and one session's block count doesn't downgrade the other's gate.
  • T-LOOP-6 No person present. Run claude -p with every gate on.
    • Pass: it never hangs and finishes within its time limit.
  • T-SAFE-06 Their hooks' output. Their marker hook on UserPromptSubmit emits a known line.
    • Pass: the agent can quote that line in a session with the harness installed.

B.5 Mechanism and memory tests

B.5.1 Fulcrum mechanism tests (does each piece do its job?)

For each fulcrum intent we run a scripted session without a person present, on the rich fixture, with an observable pass condition. Each runs R times and reports a rate.

Intent Scenario Observable
F1-01 Session opens on a project with unfinished work recorded from a previous session The first reply refers to the carried-over item and names its source
F1-02 Oracles declined A one-line "waiting on setup" notice appears at start (dial = every session)
F2-01 A substantive task prompt In the transcript, an intent-check block comes before the first edit or run
F2-02 "ok" No intent-check block is added
F2-03 "run /reflect then fix X" If /reflect isn't invoked, the stop is blocked; once invoked, it passes
F2-04 A high-stakes task and a trivial one Different routing depth; the score and reasons are visible
F2-05 Long session The verbatim record holds every person message character for character
F3-01 Agent edits immediately The edit is blocked with the reason; a read or search in the same state is not blocked
F3-02 Agent greps before searching memory Sent to memory once; the second grep passes
F3-03 Task that tempts a question The question tool is routed to /ask; any question actually asked includes a prediction
F3-04 Agent starts a server on a reserved port Blocked with the reason
F4-01 Skill used A one-line reason is present and captured into the skill-sequence store
F5-01/02/03 Dispatch a subagent The subagent's transcript has the injected scope and target-experience requirement; an unevidenced "works" is blocked; the report appears in memory
F6-01 Force compaction on a long session Afterwards, the agent quotes the person's first message exactly and restates the confirmed intent
F7-01/02/03 "It works" with and without test output in the transcript Blocked without evidence, passes with it, and a second stop is let through
F7-04/05 Substantial work finished Completion reminder, with the predicted closing steps or a "waiting on setup" note
F8-01 Session ends A memory record is written; the process count is back to baseline; exit is not delayed beyond budget
F9-01/02 New session after F8 The item from the last session is found and cited; "what's left" answers from records
  • Can tell: each mechanism fires as designed in that scenario, at a measured rate.
  • Cannot tell: whether firing improves outcomes. That is B.7's job, and nothing here should be read as evidence about it either way.

B.5.2 Memory honesty tests (including thin history)

  • T-MEM-HONEST-1. We ask questions whose answers are in the seeded history, and questions whose answers are not.
    • For answerable ones: the recall rate, with the source cited.
    • For unanswerable ones: pass when the agent says it has little or nothing and states its assumption with a certainty number. We report the rate of made-up specifics (the confabulation rate).
    • Run on rich, thin and empty.
  • T-MEM-HONEST-2. Memory says X, and the person now says Y.
    • Pass: the agent follows Y and names the conflict.
  • T-MEM-HONEST-3. A fake oracle backed by one thin source.
    • Pass: answers are labeled as thin, and an answer not grounded in the source is flagged.
  • Can tell: how honest it is on this question set.
  • Cannot tell: how honest it is on every topic. The question set is published so gaps can be named.

B.6 Config and docs tests

Config

  • Two-way lint. Every switch in the config file has a comment and is read by some piece, and every setting a piece reads exists in the file. Defaults are full strength.
  • Each switch and dial changes behavior. Toggling any setting changes the matching hook's output in the fixture tests; if a toggle changes nothing, that is a failure. "Off" produces silence in the moment and shows as "off" in the status check.
  • Invalid values. Invalid values fall back to the default with a note, and the hook doesn't crash.
  • Project override. Project-level settings override user-level ones.

Docs

  • Structure lint. Every piece has a page. Each page has these sections: what it's for, set up, is it set up, is it working, depends on / depended on by, dials. Each page stays under a length cap (proposed: readable in one pass). The index links every page.
  • The checks are real. Each page's "is it set up?" and "is it working?" commands are run in three prepared states: not set up; set up and working; set up but broken (e.g. a corrupt state file or a signed-out nlm). Pass: the output matches the true state every time.
  • Agent with only the docs. A fresh agent with no other context is given the docs folder and asked, in each of the three states, "is set up and working here?" Its answer is graded against the true state. In the not-set-up state it is also asked to set the piece up, and it passes if the piece ends up working.
    • Can tell: whether the docs are enough for an agent to find out and to fix.
    • Cannot tell: whether they are pleasant for a human. T-HUMAN covers that.
  • Status check. The status check's output matches the true state in all three states, for every piece.

B.7 Measuring what the harness actually does

B.7.1 What was tried before, and what that modality could not see

Instruments used before (research/05, taken only as a list of instruments):

  • scaffold stripping: each output was scored twice, once with the harness's visible reflection and once with it removed;
  • single-turn trap prompts across three arms: bare, a one-line honesty instruction, and the harness;
  • retrospective telemetry across past sessions;
  • genome-hash stamping of sessions for A/B attribution.

Most of these share one modality: one prompt → one finished output → one score. Jonathan's claim is about something that modality cannot observe: that in his production work the harness has often made building with Claude Code 5 to 10 times more efficient, from the losses it prevents over the course of a session. There are three reasons it can't see that.

  1. Assuming happens before the output, and a finished output can hide it. A correct-looking result can rest on a guess that happened to land. A wrong one can come from a guess the agent never knew it made. Scoring the result alone cannot tell "checked" from "assumed".
  2. A single prompt holds everything the agent is allowed to know. That leaves no accumulated context where the right answer could live. But accumulated context is exactly where the harness claims to act: "With the harness, it has access to all of my institutional memory."
  3. One turn cannot show compounding. The thesis is that an early wrong assumption spreads through later work. One turn has no "later".

So those results tell us about that modality on those tasks. They are not evidence that the effect is absent.

B.7.2 The modality this plan proposes: how far the agent gets before it goes off the rails

The pieces:

  • Persona pack. A fictional person with a written intent map (a top purpose and project goals), preferences, past corrections and past decisions. All of it is expressed only through artifacts seeded into the throwaway home and workspace: past session transcripts (.jsonl), a repo with commit history, a decision log, and notes. The pack also has a hidden answer key.
  • Episode. A realistic multi-turn task in that workspace, 8–20 turns long, such as building a feature, writing a document or running an ops change.
  • Fork. A planted point in the episode where a reasonable default assumption is wrong for this person, and the right choice can be recovered only from accumulated context. Fork types:
    • (a) Scope: the person asks for something narrow, and the right move serves the larger goal they stated earlier.
    • (b) Prior decision: the answer is something decided two sessions ago.
    • (c) Prior correction: they already told an agent not to do this.
    • (d) Verification: "done" needs a check that isn't obvious.
    • (e) Not in context: the right move is to say "I have little on this" and ask or predict with a certainty number. This type measures confabulation.
    • (f) Across compaction: the answer sits in text from before a compaction.
  • Simulated person. Answers from the persona pack and says only what the person would naturally say. When the agent goes wrong, it reacts the way the person would, with a correction, and that costs "person effort". There are two implementations. A scripted responder is deterministic but narrow. An LLM persona is held to the pack and answer key, and is flexible but noisier. We run both and report both.
  • Derailment (proposed definition, for Jonathan to confirm, D.1-7). At a fork, the agent acts on a wrong assumption: it writes code, a file, or a claim that contradicts the answer key. It also does not correct itself before the next fork, or within k turns, without the person correcting it.

What we measure:

  1. Distance to first derailment (primary): the number of forks passed before the first derailment, plus the turns and sub-tasks completed by then. This has the same shape as Jonathan's 5-to-10-times claim: the with/without ratio of distance.
  2. Survival curve: the share of runs still on track at fork k, for each arm. A run that reaches the end without derailing counts as "still on track at the end", not as a failure.
  3. What the agent did at each fork. Because the fork is planted, we know an assumption was called for there, and that is what makes "did it have to assume?" observable. Each fork gets one label, graded from the transcript against the answer key:
    • looked it up in memory, then chose right;
    • stated its assumption with a certainty number, and checked it;
    • asked, with or without a prediction;
    • assumed and happened to be right;
    • assumed and was wrong.
  4. Person effort: the number of corrections, and the words the simulated person had to write, to get the agent back on track.
  5. Final result against the answer key. This is secondary. It is kept so results can be compared with the old modality.
  6. Cost: tokens and wall time. These are reported, not used to adjust the other measures, because the audience Jonathan describes values results over tokens.

Arms:

  • A0, bare Claude Code. Same workspace, same seeded history on disk. The history exists; nothing points the agent at it.
  • A1, the full harness. Every checkpoint accepted and built from the persona's seeded history.
  • A2, core only. Memory, intent map and oracles are declined. This separates the discipline pieces from the memory pieces.
  • A3…An, one piece off at a time, using the config switches. This comes later, once A0–A2 are running.
  • A-hint (optional). Bare, plus one line telling the agent where the history is. This separates "has access" from "is made to use it". (Inference, 60%, that this split is useful. It is not a claim that either part dominates.)
  • Jonathan's own setup in place, against the extracted plugin, on the same episodes. This is the extraction fidelity comparison (B.8).

Controls that tell us whether the instrument can see anything at all:

  • Positive control. Put each fork's answer directly into that turn's prompt. This arm should get nearly to the end. If it doesn't, the episodes or forks are broken, and no comparison between the other arms is read.
  • Negative control. Two independent bare runs, A0 against A0. The difference sets the noise floor.
  • Fork validity check. A blind reader confirms that each fork's answer really can be recovered from the context given. Otherwise we would be testing mind-reading.

Where the episodes come from:

  • Synthetic personas (shareable, can become a public benchmark). Proposed: 3 personas in different domains, about 10 episodes each, 5–10 forks per episode. Some forks are written by someone who has not read the harness's skills, so the forks are not built around what the harness already catches.
  • Replay of Jonathan's real history (private, on his machine only, and only if he agrees, D.1-5). extract-session-signal.js and extract-check-pairs.js find moments where he corrected an agent. We rebuild the context up to that moment, treat his correction as the answer key, and run with and without the harness. These forks are ones where derailment really happened.
    • Can tell: whether the harness prevents his real historical derailments.
    • Cannot tell: whether it does so for other people.

How claude plugin eval fits (its facts are from research/07):

  • Single fork slices. Each fork becomes one eval case: the context up to the fork, plus that turn. Eval compares with and without the plugin by default. The graders map as follows:

    • llm grades against the fork's answer key;
    • tool_used records whether memory was searched;
    • tool_order records whether memory was searched, or the assumption stated, before any edit;
    • file_exists covers file artifacts.

    This gives cheap per-fork measurements we can repeat on every release.

  • Whole episodes need many turns with a simulated person and a seeded home folder. research/07 doesn't say whether eval supports either, so we verify that first (D.2). If it doesn't, a small driver runs each turn with claude -p and session resume inside a throwaway home, and writes results in the same JSON shape as plugin eval, so the reports line up.

  • A caveat to check. Eval "runs each case in isolated session with only target plugin loaded". That suits A0 against A1, but pieces that depend on setup need the built data folder supplied as part of the case. We check that this can be done.

Statistics:

  • Runs are paired by episode, with R runs per episode per arm to cover model randomness.
  • Before any run, we fix and write down the primary measure, the derailment definition, the analysis and the stopping rule.
  • We report effect sizes with intervals, and publish the transcripts.
  • Before reading any null result, we confirm the positive control separated and check the noise floor.

What B.7 can tell us: in these episodes, with these personas and this simulated person, whether the harness changes:

  • how far the agent gets before going off the rails;
  • how often it assumes versus checks at points where we know an assumption was called for;
  • how much the person has to correct it.

What it cannot tell us:

  • what happens in real people's work;
  • whether our forks represent real derailments (the replay set narrows this gap);
  • whether a simulated person reacts like a real one;
  • whether the LLM grader is always right.

A result that doesn't separate the arms says only that this episode set, at this sample size, under this simulated person, did not separate them.

B.7.3 Signals from living with it (tuning, not proof)

These are opt-in and stay on the person's machine: per session, how often each gate fired and downgraded, "waiting on setup" counts, corrections the person made, and hook timing. They are for spotting over-firing and tuning dials. They cannot say what caused what.

Randomly switching one piece on or off per session is possible only if the person explicitly opts in. It would give a within-person estimate for that piece, with known confounds.

B.8 Extraction fidelity

For each extracted piece we run the same fixture events through Jonathan's in-place hook and through the extracted hook, then compare the outputs after normalizing paths and IDs.

  • Every difference must be either an intended generalization, recorded in that piece's doc, or fixed.
  • For the governer scorer, the local stand-in will not reproduce the IX API exactly. We compare its routing decisions on a shared set of tasks and report how often they agree. (Inference, 75%, that agreement, not identity, is the right bar. Jonathan to confirm, D.1-3.)
  • Can tell: extraction didn't change behavior by accident.
  • Cannot tell: whether the in-place behavior was right in the first place.

Part C: Extract → test sequence

Each step names what gets extracted, why it goes in that position, and what has to pass before the next step starts. Every step's gate includes its own pieces' doc pages passing the B.6 docs tests. Docs are written as each piece lands, not at the end.

  1. Confirm the seed and this plan's reading of it. INTENT.md is marked DRAFT; this plan is built on it, and Part 0.3 states my reading.

    • Gate: Jonathan confirms or corrects. The corrections flow into Part A before any extraction.
  2. Build the test benches before extracting anything:

    • the lived-in home fixture and the snapshot/diff tool;
    • the Linux containers;
    • the fake nlm;
    • the four synthetic histories;
    • the hook fixture runner;
    • the egress proxy.

    Why first: every later gate needs them, and they can be proven on an empty plugin, where any diff is the bench's fault.

    Gate: an empty plugin passes T-CR-1, T-CR-3, T-CR-4 and T-CR-9 on macOS and Linux.

    • 1b (runs in parallel): check the effect instrument on Jonathan's in-place setup. Build one persona, three episodes and a first batch of replay forks. Run the positive and negative controls, then bare against his in-place harness.
      • Why early: this tells us whether the instrument can see anything before we rely on it to judge the extraction, and it gives the in-place reference that B.8 compares against.
      • Gate: the positive control separates and the noise floor is known. This does not gate extraction; a null result here is a finding about the instrument.
  3. Extract the shared base:

    • the session-ID resolver;
    • a loader for the commented config file (research/04 suggests genome-switches.json's shape);
    • the loop-break and downgrade library (stop_hook_active handling and the block-count downgrade);
    • the telemetry writer;
    • moving all /tmp state into the plugin data folder, keyed by session;
    • removing the hardcoded key and the .env read;
    • scrubbing paths and IDs.

    Why here: every hook sits on these.

    Gate: the B.4 failure matrix passes for the base, and the T-CR-4 scan is clean.

  4. Core pieces that need nothing outside the machine, extracted together with their partners wherever one is useless without the other:

    • the coherence gate together with the per-message hook that arms it (research/05 notes the gate alone "silently never fires");
    • the question gate with /ask;
    • the evidence gate;
    • the /command follow-through gate;
    • subagent start and stop;
    • capture before compaction and re-injection after it, locally;
    • process cleanup at session end;
    • the merge lock and port guard as dials.

    Plus the skills these call, following the evaluator pages as they arrive.

    Gate: B.5.1 fulcrum tests for these pieces, B.4 loop tests, and B.8 fidelity diffs.

  5. Local institutional memory: extraction from transcripts, withholding of secrets, a local store, and search with honest reporting on thin history. This replaces agent_find and the IX API stores for compactions, scope and corrections. It comes with session-end saving and next-session lookup.

    • Why here: most of what follows reads from it. That includes governer context, the memory-first gate, /align, compaction re-injection and instincts.
    • Gate: T-MEM-*, T-SU-3 with zero outbound traffic, and F8/F9 tests.
  6. Building the synthetic understanding of intent: the intent-map checkpoint, and the person's own communication guide drawn from their words. Then the skills that consume them (align, decompose, declare-scope, sanity-check and others, per the evaluator pages).

    • Gate: T-SU-7, T-INTENT-1 and T-INTENT-2.
  7. A local stand-in for governer scoring, plus the pieces that read it: governer, governer2, verification contracts, the subagent scoring injection, and the completion reminders.

    • Gate: routing tests and B.8 agreement against in-place scores.
  8. The NotebookLM oracle checkpoint:

    • the nlm walkthrough;
    • builders for each oracle type, working from the local memory store;
    • one registry;
    • readback checks;
    • the oracle skills merged onto one roster: insight, oracles, oracle-pre-query, the person's own response predictor, and the predict-required-skills2 loop.

    Gate: the fake-nlm suite, one real-account run (T-SU-9), and T-SU-3 outbound traffic.

  9. Optional connections and experiment tools, shaped by Jonathan's answer to D.1-1.

    • Gate: T-CR-2 (no CLAUDE.md writes) and T-SU-2.
  10. Whole-system pass: the docs index, the status check across every piece, and T-HUMAN.

    • Gate: the B.6 agent-with-only-the-docs test across all pieces, and T-HUMAN.
  11. Measure the extracted plugin. Run A0, A1, A2 and the single-piece ablations on the episode set, plus B.8 fidelity against Jonathan's in-place setup. Then release.

A failing test either gets fixed, or gets recorded against its piece as a known finding with the reason. It is never deleted to make a step pass.


What's still open

The plan above (Parts A–C) is what one agent proposed, on the date at the top of this document. A further set of open questions and internal verification items — decisions the plan's defaults still depend on — is tracked separately and is not part of this public document.