← the whole session

Fulcrum 13 of 18

Calling it finished

When the agent is about to say something works, is fixed, passes or is live, and end its turn.

What goes wrong here

The agent reports success it hasn't observed. It read the code and concluded the thing works, or it ran a check that tests something nearby rather than the thing itself. Saying so sounds exactly like reporting a verified result — the danger isn't the guess itself, it's treating an unverified guess as settled reality. That gap is where hallucination comes from.

How it compounds if nothing catches it

A false 'done' is the most expensive kind of drift, because it closes the loop and nobody looks again. The person finds out in front of a user, or the next session builds on it as settled.

What the harness does here

A stop gate (stop-evidence-gate.sh) reads the agent's last message. If it says something works and the transcript holds no verification output, the gate blocks the stop and names what to verify, including any verification contract written at the start of the task. It has a loop-break so it can't trap the agent. verify makes the agent define the outcome before gathering evidence, so a nearby stand-in can't pass for it. validate-load-bearing-claims-against-reality requires measuring reality for anything a decision will rest on. verification-gate checks that every file-and-line citation exists before an answer is published. complete-seed and sanity-check compare the finished work with the confirmed intent. A second stop gate is meant to block the stop when the person asked for a /command the agent never ran, but its feed is currently switched off.

The pieces that act here

  • stop-evidence-gate.sh

    A hook: a script Claude Code runs on its own at a set moment.

  • /verification-contracts

    Task-specific verification contracts — written before work starts, recorded with `alignment-harness contract add`, and enforced by the Stop hook at task end. The full reference for the per-message contract prompt.

  • /verify

    Outcome-first verification protocol. Forces the agent to define what the OUTCOME looks like in measurable terms BEFORE gathering any evidence. Blocks the failure pattern of substituting mechanism-evidence (hook is configured, file exists, command runs) for outcome-evidence (the thing the human wants to be true is actually happening). Use whenever asked to check if something is working.

  • /validate-load-bearing-claims-against-reality

    Use when you are about to report or act on a load-bearing conclusion — especially one a governer scored 90+ — that you reached by reading code, docs, or any single source of inference. This skill is the discernment that turns operating-certainty into validated truth by measuring the actual reality your claim describes. Triggers — high-leverage conclusion, "the code shows", "the DB does X", "this resolves to", "the gate fires", reporting a finding you have not measured against the thing it is a claim ABOUT, any 90+ governer task before the completion claim.

  • /verification-gate

    Final-pass auditor that verifies every code citation in a research/synthesis answer against actual files BEFORE publishing. Use at the end of any code-grounded research task. Catches the retrieval-augmented confabulation pattern where agents invent specific file/line citations that feel grounded but don't exist.

  • /complete-seed

    Agent-invoked completion gate. When an agent believes its work is done, this skill forces it to collect evidence — load the page, grab screenshots, run verification commands from the seed's success criteria. The agent invokes this itself when it's ready. No hooks, no interrupts.

  • /sanity-check

    Confirms that work matches intent — before starting, after completing, before committing, or when reviewing a plan or diff.

  • /anthropic-proof-its-fixed

    Produce a full evidence package proving something was broken and is now fixed. Use when an agent finishes a fix and needs to prove before/after with data, screenshots, and optionally video. Invoke after completing any fix that touches user-facing behavior.

  • stop-slash-command-gate.sh

    A hook: a script Claude Code runs on its own at a set moment.

Loops that pass through here