← the whole session

Fulcrum 18 of 18

When the harness itself changes, or quietly stops working

Whenever a skill, hook or protocol is edited. It also applies all the time, because a hook can stop firing without anyone noticing.

What goes wrong here

A change meant to improve alignment makes agents worse, and nobody can tell. Or a hook silently stops firing (for example, after a Claude Code update renames a field it reads), and the protection it gave disappears without a trace. Or a skill exists but no agent can find it.

How it compounds if nothing catches it

A broken safeguard is worse than none, because the person still trusts it. Every session after the break inherits the gap, and without measurement nobody can connect the drift to its cause.

What the harness does here

genome-experiment takes a fingerprint of the whole agent setup (skills, hooks, CLAUDE.md) and stamps each session with it, so sessions can be compared across versions of the setup. A circuit breaker warns, but never reverts, when recent alignment scores fall below the baseline. Some start hooks assign sessions to different versions of a prompt so those versions can be compared. agent-debug shows which hooks fired this turn and what they injected. alignment-doctor checks one thing: whether institutional context is being injected. discoverability-audit-learnings-findable-by-agents checks that new tools and skills can actually be found by future agents.

The pieces that act here

  • /genome-experiment

    A/B test changes to the agent operating system (skills, hooks, CLAUDE.md) by measuring alignment KPIs via Statsig. Use when restoring skills, changing hooks, editing CLAUDE.md, or modifying governer config and you want to know if it helped or hurt.

  • session-start-genome-stamp.sh

    A hook: a script Claude Code runs on its own at a set moment.

  • genome-circuit-breaker.sh

    A hook: a script Claude Code runs on its own at a set moment.

  • session-start-coherence-gate-ab.sh

    A hook: a script Claude Code runs on its own at a set moment.

  • /agent-debug

    Print a diagnostic overview showing what harness hooks fired, what got injected into context, what's silent, and what's shaping agent behavior — before every response. Use when the person wants to see what's actually happening behind the scenes.

  • /alignment-doctor

    Check whether institutional context (a memory-search result) was actually injected into the last turn, with a one-line reason when it wasn't. Reads the harness's own telemetry log rather than scanning raw Claude Code internals, so it works the same on any install.

  • /discoverability-audit-learnings-findable-by-agents

    Audit whether tools, skills, and workflows you created are actually findable by future agents — and fix gaps before they become invisible, duplicated work.

Loops that pass through here