← the whole session plugin/skills/verification-contracts/SKILL.md

Task-specific verification contracts — written before work starts, recorded with `alignment-harness contract add`, and enforced by the Stop hook at task end. The full reference for the per-message contract prompt.

Verification Contracts — Task-Specific Completion Evidence

The Problem This Solves

Agents hallucinate completion. They see "no errors" and report "done." The generic evidence gate (/verify, /verification-gate) catches vague runtime claims, but it can't answer: did you verify the RIGHT thing?

A curl to /health is not evidence that /api/users/signup works. Passing unrelated tests is not evidence that YOUR changes work. TDD in isolation is not proof of production behavior.

What a Verification Contract Is

A verification contract is a specific, falsifiable statement of what constitutes proof of completion for THIS task — written by you before the work starts, recorded locally, and read back by the Stop hook when you try to finish.

Each item has two parts:

  • Intent statement: "When {who} does {what}, given {ALL conditions}, then {testable outcome}." Include negative statements too ("When X, people do NOT see Y"). No unconditional statements.
  • Verification command: a runnable action — a command, a request, a test, or a browser step whose output proves the statement true.

How It Works — Two Gates, One Contract

START (whenever the task is scored 20 or more)

The per-message hook already reminds you of this when it applies — this skill is the full version of that reminder.

  1. The person assigns work.
  2. Decompose it into experience statements, per the format above. Every condition that matters — user state, feature flags, fail-open/fail-closed behavior — appears in the statement described as what the person experiences, not as variable names. Variable names are pointers in parentheses, not substitutes for meaning.
  3. Pair each statement with a runnable verification command.
  4. Print the contract to the person before acting — this is the accountability mechanism; they see exactly what you committed to check.
  5. Record each item so the finish-line gate can hold you to it:
    alignment-harness contract add --intent "WHEN <conditions> THEN <outcome>" --verify "<runnable command>"
    
  6. Keep working — recording the contract doesn't block you, it just makes the promise checkable later.

END (Stop hook)

  1. You claim completion (or Claude Code tries to end your turn).
  2. The Stop hook reads this session's verification contract from local state.
  3. If the task's score is at or above the contracts threshold (default 60, see alignment-harness config show → evidenceGate.contractsAtScore) and any item is unverified:
    • The block message lists every unverified item, with its intent statement and its command.
    • Run each command, look at the output, and only then mark it:
      alignment-harness contract verify <n>       # one item
      alignment-harness contract verify all        # all at once, only if you actually ran every check
      
    • Only mark an item verified after its command's output actually shows the intended outcome — marking it to get past the gate without running it is the exact hallucination this mechanism exists to catch.
  4. If there's no contract for this task, the generic evidence gate still applies (did you run something that shows your edit works).

Check the contract at any point: alignment-harness contract list. Clear it if scope changed enough that the old items no longer apply: alignment-harness contract clear (then write a new one — don't silently modify old items to match new work).

When Contracts Are Required

Governer score determines contract depth (a starting point — adjust to your own judgment of the task):

Score range Contract requirement
< 20 No contract — trivial task, generic evidence gate is sufficient
20–59 Lightweight contract — 1-2 verification items for the primary change
60–79 Standard contract — verification item per intent statement
≥ 80 Full contract — verification item per intent statement + regression check on existing behavior

If your own risk table (alignment-harness risk-table show, confirmed via /alignment-harness:harness-setup) marks a category as always-critical for your product (payment, auth, and whatever else you named), treat those as always requiring a full contract regardless of score.

Writing Good Verification Contracts

The Intent Statement

BAD (unconditional — you'll verify the wrong thing):

Users see a preferences section on /settings

GOOD (fully conditional — you know exactly what to verify):

When a paying user visits /settings, given they have an active subscription,
then they see a Preferences section with 3 toggles.
When a user visits /settings who is NOT a paying subscriber, they do NOT see the
Preferences section (fail-open: non-payers see the page without the section, never an error).

The Verification Command

Must be a specific, runnable action — not a description.

BAD: "Check that the preferences section works" GOOD: "Open localhost:3000/settings with a paying-user session, verify 3 toggles render. Then with a non-paying session, verify the section is absent."

BAD: "Run the tests" GOOD: "npx jest src/components/tests/PreferencesSection.test.js — expect 0 failures"

BAD: "Verify the API endpoint" GOOD: "curl -s -X PATCH localhost:3000/api/user/preferences -H 'Cookie: token=...' -H 'Content-Type: application/json' -d '{"emailNotifications":false}' | jq '{status: .status, emailNotifications: .data.emailNotifications}'"

Template shapes for critical paths

Adapt these to your own product's actual endpoints and flows — they're shapes, not literal commands to copy:

Payment changes:

intentStatement: "When a user with {subscription_state} visits {page}, given {conditions}, then {payment UX outcome}"
verificationCommand: "curl -s <your API>/{payment endpoint} -H '...' | jq '{status, plan, subscriptionStatus}'"
regressionCheck: "Verify existing payment flow still works: curl -s <your API>/{whoami endpoint} -H '...' | jq '.user.subscriptionStatus'"

Auth changes:

intentStatement: "When a {user_type} attempts to {auth_action}, given {auth_conditions}, then {auth_outcome}"
verificationCommand: "curl -s <your API>/{auth endpoint} -H '...' -d '{...}' | jq '{status, user.email, authenticated}'"
regressionCheck: "Verify existing login flow: open the login page, complete it, verify redirect to the authenticated home"

Example (opt-in illustration):

intentStatement: "When a user sends a message during a coaching session, given {session_conditions}, then {coaching_outcome}"
verificationCommand: "curl -s <their API>/coaching/send -H '...' -d '{message: \"test\"}' | jq '{status, response}'"
regressionCheck: "Open the app, start a conversation, verify response quality"

Integration with Existing Systems

  • The per-message hook reminds you to write a contract once a task is scored 20+ (hooks/user-prompt.js injects the reminder; this skill is the full version).
  • The Stop hook (hooks/stop.js) reads the contract from this session's local state and blocks on unverified items once the score passes the contracts threshold.
  • alignment-harness status shows how many contract items exist and how many are verified.
  • TaskCreate — each intent statement can map to a task, if you're also tracking tasks.

Rules

  1. Contracts describe a fixed target once written — don't quietly edit an item's meaning during execution. If scope changes, alignment-harness contract clear and write new items reflecting the new scope.
  2. Every contract item must have a runnable verification command — "check the UI" is not runnable. "Open localhost:3000/settings with a paying-user session" is.
  3. Evidence must match the specific contract — generic evidence (curl to /health, an unrelated test passing) does not satisfy a specific contract item.
  4. The contract is printed to the person — they see exactly what you committed to verify. This is the accountability mechanism.
  5. Contracts persist in local session state — they survive across turns. The Stop hook reads them every time it fires, for this session.