A harness for Claude Code

A harness that acts at every moment of a Claude Code session where an agent can drift from what the human meant.

The Alignment Harness is the result of applying systems theory to real-world, scalable workflows that run fleets of 500+ agents: tracing the gaps in their coherence back to their origins, and surgically testing solutions that extend the coherence they operate with at scale.

The thesis: misalignment compounds at each fulcrum

When an agent misreads what a person meant, it rarely fails at that moment. It fails later, and bigger. A misread message becomes a wrong plan. The plan becomes code, and tests written to the same wrong target. The tests pass, so the agent reports success. The report goes into a commit and a session summary, and the next session reads that summary as settled fact and builds on it. Each step inherits the error and adds its own, so misalignment compounds rather than adds up.

There are specific, recurring moments in a Claude Code session where this drift gets in. This map calls them fulcrums, because a small shift at one of them moves everything that comes after. A single check at the end can't undo what compounded along the way, and a single check at the start can't foresee what drifts later.

So the harness intervenes at each fulcrum, from before the first message to the next session, and tries to catch drift while it is still cheap. At a fulcrum the fix is one sentence of 'here is what I think you mean'. Later it is a day of unwinding work built on the wrong reading.

What it has made possible

In daily production use, the harness has made building with Claude Code dramatically more efficient — by the author's own estimate from that use, on the order of 5 to 10 times, and a session run without it gets roughly a fifth as much done before it goes off the rails, because it has to assume too much about what's wanted. The gain isn't the agent doing more; it's the absence of the losses that hallucination, miscalibration and compounding misalignment otherwise create at each point where a session can drift from what was meant.

Who it is for

This is explicitly for research purposes, for any party working towards greater alignment in AI.

A session, fulcrum by fulcrum

A fulcrum is a moment in a Claude Code session where the agent's work can drift from what the human meant — from before the harness knows the person at all, through every turn, to the next session and to changes in the harness itself. Each stop below is one of them. Open any stop to see what goes wrong there, how it compounds, and which pieces of the harness act at that moment.

Higher = further from what the human meantNothing catches itCaught at each fulcrum123456789101112131415161718Session startSession end, and afterHigher = further from what the human meantNothing catches itCaught at each fulcrumSession startSession end, and after

An illustration of the thesis, not measured data. Height is the distance between the work and what the human meant. Both traces assume the same fresh drift enters between every pair of fulcrums. In red, nothing catches it, so each step builds on the drift already there and it multiplies. In green, the harness catches it at each fulcrum and pulls it back — not to zero. At fulcrum 1 the two are the same; everything that separates them afterwards is compounding.

  1. Fulcrum 1: Before the harness knows the person

    Installation and the first sessions of someone new, before the harness has any record of what this person wants.

    6 pieces of the harness act here

  2. Fulcrum 2: When a session opens

    Session start, before the person has typed anything.

    8 pieces of the harness act here

  3. Fulcrum 3: The moment a message arrives

    Each time the person sends a message, before the agent has read it.

    5 pieces of the harness act here

  4. Fulcrum 4: Between understanding and the first action

    After the agent has read the message, before it edits, writes or runs anything.

    5 pieces of the harness act here

  5. Fulcrum 5: Before the agent rebuilds from scratch what is already known

    When the agent needs to understand something (a system, a past decision, a preference) and is about to work it out from the code or from its own guess.

    9 pieces of the harness act here

  6. Fulcrum 6: Turning intent into scope and tasks

    When the agent translates what the person wants into a plan, a task list or a handoff.

    8 pieces of the harness act here

  7. Fulcrum 7: Deciding how much care the work deserves

    At the start of any task, before implementation.

    6 pieces of the harness act here

  8. Fulcrum 8: When the agent is about to stop and ask

    The moment the agent thinks 'I should check with the person', whether mid-task, at an ambiguity, or before a decision.

    6 pieces of the harness act here

  9. Fulcrum 9: While it works

    During long stretches of work between messages, and whenever several agents share one machine.

    8 pieces of the harness act here

  10. Fulcrum 10: Handing work to a subagent

    When an agent delegates a piece of work to another agent (a subagent), and when that subagent reports back.

    9 pieces of the harness act here

  11. Fulcrum 11: When the conversation is compacted

    When a long conversation fills the context window and Claude Code replaces the earlier part with a summary (compaction).

    5 pieces of the harness act here

  12. Fulcrum 12: When the person corrects the agent

    Any time the person says, in whatever words, that the agent got something wrong.

    8 pieces of the harness act here

  13. Fulcrum 13: Calling it finished

    When the agent is about to say something works, is fixed, passes or is live, and end its turn.

    9 pieces of the harness act here

  14. Fulcrum 14: Reporting back to the person

    At the end of a turn that produced something, when the agent tells the person what happened.

    14 pieces of the harness act here

  15. Fulcrum 15: Committing the work

    When changes are turned into a commit.

    5 pieces of the harness act here

  16. Fulcrum 16: When the session ends

    When the session closes, whether deliberately or not.

    7 pieces of the harness act here

  17. Fulcrum 17: The next session reaching what was already settled

    The start of any later session, and any moment in it where a past decision matters.

    10 pieces of the harness act here

  18. Fulcrum 18: When the harness itself changes, or quietly stops working

    Whenever a skill, hook or protocol is edited. It also applies all the time, because a hook can stop firing without anyone noticing.

    7 pieces of the harness act here

The core loops

The harness isn't one check. It runs a few loops that repeat through a session, each passing through several fulcrums.

Every turn: understand before acting, prove before calling it done

Each message runs the same small cycle. The person's words are put first and relevant context is pulled in. The agent shows its reading before it acts, looks up what is known before guessing, and predicts before asking. At the end of the turn, whatever the agent says is finished is checked against evidence, and the report is written so a person can read it in seconds. What this keeps true is that the person can see, every turn, whether the agent understood them, and a misreading costs one sentence to fix instead of a day.

Passes through3. The moment a message arrives4. Between understanding and the first action5. Before the agent rebuilds from scratch what is already known8. When the agent is about to stop and ask13. Calling it finished14. Reporting back to the person

Every task: intent, scope, rigor, work, proof, record

A task moves from intent to confirmed scope, then to a rigor score, the work itself, evidence, a report, and a commit that records why it was done. What this keeps true is that what gets checked at the end is what the person meant at the start, and that the depth of checking matches what a mistake would cost.

Passes through6. Turning intent into scope and tasks7. Deciding how much care the work deserves9. While it works13. Calling it finished14. Reporting back to the person15. Committing the work

Every delegation: the why travels with the work, and proof comes back

Each time work is handed to a subagent, the intent goes with it: the target experience is stated before any code, along with a verification contract that could fail. What comes back is checked against evidence rather than taken on the subagent's word, and anything the subagent learned is saved. What this keeps true is that splitting work across many agents doesn't multiply the misreadings.

Passes through6. Turning intent into scope and tasks10. Handing work to a subagent13. Calling it finished

Every session: open informed, survive compaction, close without losing anything

A session opens with the operating protocol and standing trigger rules already loaded. It keeps the person's exact words and the confirmed scope through compaction. When it ends, it files what it learned and what is unfinished. What this keeps true is that the session that ends knows more than the one that started, and the next one can reach that knowledge.

Passes through2. When a session opens11. When the conversation is compacted16. When the session ends17. The next session reaching what was already settled

Every correction: fix the source, not just the instance

When the person corrects the agent, the correction is traced to the instruction or premise that produced the error, and that source is fixed. The correction is captured so it can be injected the next time the same situation comes up. What this keeps true is that the person teaches something once, not every week.

Passes through12. When the person corrects the agent9. While it works17. The next session reaching what was already settled

Across sessions: the learning flywheel

As skills work, they print blocks that are easy to recognise: COHERENCE_CHECK, SANITY_CHECK, reflections, scope declarations, and a one-line note each time a skill is used saying why. At session end, hooks and scripts extract those blocks word for word, together with the person's replies. What the person confirmed or corrected becomes instincts, principles and oracle notebooks. Those are read the next time a similar situation comes up, and the person's reaction to the next check becomes new material. What this keeps true is that alignment improves with use instead of resetting every session. As of September 2026, some parts of this loop run and some are designed but not yet built. For example, according to an earlier reading of the skill files, the combined notebook that jonathan-check3 is meant to query has not been created.

Passes through4. Between understanding and the first action5. Before the agent rebuilds from scratch what is already known7. Deciding how much care the work deserves12. When the person corrects the agent16. When the session ends17. The next session reaching what was already settled1. Before the harness knows the person

Continuously: the harness checks itself

Each session is stamped with a fingerprint of the setup that ran it. Alignment scores are watched against a baseline, and diagnostic skills show which hooks actually fired. What this keeps true is that a change to the harness, or a piece of it that has silently stopped working, can be noticed instead of trusted blindly.

Passes through18. When the harness itself changes, or quietly stops working2. When a session opens