← the whole session plugin/skills/verification-contracts/SKILL.md
Task-specific verification contracts — written before work starts, recorded with `alignment-harness contract add`, and enforced by the Stop hook at task end. The full reference for the per-message contract prompt.
Verification Contracts — Task-Specific Completion Evidence
The Problem This Solves
Agents hallucinate completion. They see "no errors" and report "done." The generic evidence gate (/verify, /verification-gate) catches vague runtime claims, but it can't answer: did you verify the RIGHT thing?
A curl to /health is not evidence that /api/users/signup works.
Passing unrelated tests is not evidence that YOUR changes work.
TDD in isolation is not proof of production behavior.
What a Verification Contract Is
A verification contract is a specific, falsifiable statement of what constitutes proof of completion for THIS task — written by you before the work starts, recorded locally, and read back by the Stop hook when you try to finish.
Each item has two parts:
- Intent statement: "When {who} does {what}, given {ALL conditions}, then {testable outcome}." Include negative statements too ("When X, people do NOT see Y"). No unconditional statements.
- Verification command: a runnable action — a command, a request, a test, or a browser step whose output proves the statement true.
How It Works — Two Gates, One Contract
START (whenever the task is scored 20 or more)
The per-message hook already reminds you of this when it applies — this skill is the full version of that reminder.
- The person assigns work.
- Decompose it into experience statements, per the format above. Every condition that matters — user state, feature flags, fail-open/fail-closed behavior — appears in the statement described as what the person experiences, not as variable names. Variable names are pointers in parentheses, not substitutes for meaning.
- Pair each statement with a runnable verification command.
- Print the contract to the person before acting — this is the accountability mechanism; they see exactly what you committed to check.
- Record each item so the finish-line gate can hold you to it:
alignment-harness contract add --intent "WHEN <conditions> THEN <outcome>" --verify "<runnable command>" - Keep working — recording the contract doesn't block you, it just makes the promise checkable later.
END (Stop hook)
- You claim completion (or Claude Code tries to end your turn).
- The Stop hook reads this session's verification contract from local state.
- If the task's score is at or above the contracts threshold (default 60, see
alignment-harness config show→evidenceGate.contractsAtScore) and any item is unverified:- The block message lists every unverified item, with its intent statement and its command.
- Run each command, look at the output, and only then mark it:
alignment-harness contract verify <n> # one item alignment-harness contract verify all # all at once, only if you actually ran every check - Only mark an item verified after its command's output actually shows the intended outcome — marking it to get past the gate without running it is the exact hallucination this mechanism exists to catch.
- If there's no contract for this task, the generic evidence gate still applies (did you run something that shows your edit works).
Check the contract at any point: alignment-harness contract list. Clear it if scope changed enough that the old items no longer apply: alignment-harness contract clear (then write a new one — don't silently modify old items to match new work).
When Contracts Are Required
Governer score determines contract depth (a starting point — adjust to your own judgment of the task):
| Score range | Contract requirement |
|---|---|
| < 20 | No contract — trivial task, generic evidence gate is sufficient |
| 20–59 | Lightweight contract — 1-2 verification items for the primary change |
| 60–79 | Standard contract — verification item per intent statement |
| ≥ 80 | Full contract — verification item per intent statement + regression check on existing behavior |
If your own risk table (alignment-harness risk-table show, confirmed via /alignment-harness:harness-setup) marks a category as always-critical for your product (payment, auth, and whatever else you named), treat those as always requiring a full contract regardless of score.
Writing Good Verification Contracts
The Intent Statement
BAD (unconditional — you'll verify the wrong thing):
Users see a preferences section on /settings
GOOD (fully conditional — you know exactly what to verify):
When a paying user visits /settings, given they have an active subscription,
then they see a Preferences section with 3 toggles.
When a user visits /settings who is NOT a paying subscriber, they do NOT see the
Preferences section (fail-open: non-payers see the page without the section, never an error).
The Verification Command
Must be a specific, runnable action — not a description.
BAD: "Check that the preferences section works" GOOD: "Open localhost:3000/settings with a paying-user session, verify 3 toggles render. Then with a non-paying session, verify the section is absent."
BAD: "Run the tests" GOOD: "npx jest src/components/tests/PreferencesSection.test.js — expect 0 failures"
BAD: "Verify the API endpoint" GOOD: "curl -s -X PATCH localhost:3000/api/user/preferences -H 'Cookie: token=...' -H 'Content-Type: application/json' -d '{"emailNotifications":false}' | jq '{status: .status, emailNotifications: .data.emailNotifications}'"
Template shapes for critical paths
Adapt these to your own product's actual endpoints and flows — they're shapes, not literal commands to copy:
Payment changes:
intentStatement: "When a user with {subscription_state} visits {page}, given {conditions}, then {payment UX outcome}"
verificationCommand: "curl -s <your API>/{payment endpoint} -H '...' | jq '{status, plan, subscriptionStatus}'"
regressionCheck: "Verify existing payment flow still works: curl -s <your API>/{whoami endpoint} -H '...' | jq '.user.subscriptionStatus'"
Auth changes:
intentStatement: "When a {user_type} attempts to {auth_action}, given {auth_conditions}, then {auth_outcome}"
verificationCommand: "curl -s <your API>/{auth endpoint} -H '...' -d '{...}' | jq '{status, user.email, authenticated}'"
regressionCheck: "Verify existing login flow: open the login page, complete it, verify redirect to the authenticated home"
Example (opt-in illustration):
intentStatement: "When a user sends a message during a coaching session, given {session_conditions}, then {coaching_outcome}"
verificationCommand: "curl -s <their API>/coaching/send -H '...' -d '{message: \"test\"}' | jq '{status, response}'"
regressionCheck: "Open the app, start a conversation, verify response quality"
Integration with Existing Systems
- The per-message hook reminds you to write a contract once a task is scored 20+ (
hooks/user-prompt.jsinjects the reminder; this skill is the full version). - The Stop hook (
hooks/stop.js) reads the contract from this session's local state and blocks on unverified items once the score passes the contracts threshold. alignment-harness statusshows how many contract items exist and how many are verified.- TaskCreate — each intent statement can map to a task, if you're also tracking tasks.
Rules
- Contracts describe a fixed target once written — don't quietly edit an item's meaning during execution. If scope changes,
alignment-harness contract clearand write new items reflecting the new scope. - Every contract item must have a runnable verification command — "check the UI" is not runnable. "Open localhost:3000/settings with a paying-user session" is.
- Evidence must match the specific contract — generic evidence (curl to /health, an unrelated test passing) does not satisfy a specific contract item.
- The contract is printed to the person — they see exactly what you committed to verify. This is the accountability mechanism.
- Contracts persist in local session state — they survive across turns. The Stop hook reads them every time it fires, for this session.