← the whole session plugin/skills/genome-experiment/SKILL.md

A/B test changes to the agent operating system (skills, hooks, CLAUDE.md) by measuring alignment KPIs via Statsig. Use when restoring skills, changing hooks, editing CLAUDE.md, or modifying governer config and you want to know if it helped or hurt.

Agent Genome Experiment System

What this does

When you change the agent operating system — restore archived skills, edit hooks, modify CLAUDE.md, update governer config — this system measures whether the change improved or degraded agent alignment (how well the agent's work matched what you actually wanted). The scripts below are a reference implementation of the mechanism, built for one specific machine — they are not shipped inside this plugin, so if they aren't present in your project yet, see "Building this yourself" below before the quick start will run as written. Statsig (a third-party A/B testing service) is an optional add-on for a dashboard and its Autotune feature; the local comparison described here works without it.

How it works

  1. A "genome hash" fingerprints the entire agent config (skills + hooks + CLAUDE.md + governer state)
  2. Every session is stamped with the current genome hash at startup (SessionStart hook)
  3. After sessions, alignment scores are logged locally as 10 separate KPI events per session, and — if you have Statsig set up (see /alignment-harness:harness-setup) — pushed there too
  4. The local comparison (or Statsig, if set up) compares sessions under different genome hashes and tells you which variant wins
  5. A circuit breaker warns if alignment drops >15% from baseline (Stop hook)

Quick start — wrap any change in an experiment

# 1. Before your change: capture baseline
.claude/scripts/genome-experiment-start.sh "Description of the change"

# 2. Make your change (restore skills, edit hooks, etc.)

# 3. Use Claude Code normally for 5-10 sessions

# 4. See results
.claude/scripts/genome-experiment-report.sh

# 5. Push all KPIs to Statsig for Autotune (optional — skip if you haven't set up Statsig)
.claude/scripts/genome-experiment-report.sh --push

# 6. When done
.claude/scripts/genome-experiment-stop.sh

For agents — one function does everything

If you have wired a Statsig-bridge script into your own project (one example lives at api/scripts/misalignment-pipeline/log-alignment-to-statsig.js — private to that setup, not shipped with this plugin), call its runGenomePipeline() export the same way it's used here:

const { runGenomePipeline } = require('<path-to-your-bridge-script>/log-alignment-to-statsig');

// Score all sessions and log to Statsig
await runGenomePipeline({ scoreFirst: true });

// Preview without logging
await runGenomePipeline({ dryRun: true });

// Score and log a single session
await runGenomePipeline({ sessionId: 'abc-123' });

Otherwise, use the harness's own local scoring instead — no server or bridge script required: run alignment-harness score --session <id> ... per session to get a governer score, and append that score plus the 10 KPIs below to a local JSONL file in the folder printed by alignment-harness records genome-experiments. That local file is what genome-experiment-report.sh should read to build its comparison table.

10 alignment KPIs logged to Statsig per session

Event name What it measures Direction
agent_alignment_score Composite alignment (0-100) Higher = better (PRIMARY)
agent_corrections Count of corrections from the person you're working with Lower = better (GUARDRAIL)
agent_confirmations Count of confirmations from the person you're working with Higher = better
agent_re_explanations Times the person had to re-explain something Lower = better (GUARDRAIL)
agent_directed_frustration Frustration directed at agent Lower = better (GUARDRAIL)
agent_frustration_ratio Frustration signals per message Lower = better
agent_directed_frustration_ratio Directed frustration per message Lower = better
agent_first_third_score Alignment in first third of session Higher = better
agent_last_third_score Alignment in last third of session Higher = better
agent_session_length Message count Context metric

Files

File Purpose
.claude/scripts/genome-hash.sh Compute deterministic config fingerprint
.claude/hooks/session-start-genome-stamp.sh Tag sessions with genome hash (SessionStart)
.claude/hooks/genome-circuit-breaker.sh Warn on alignment regression (Stop)
.claude/scripts/genome-experiment-start.sh Capture baseline, start experiment
.claude/scripts/genome-experiment-report.sh Compare control vs treatment
.claude/scripts/genome-experiment-stop.sh Archive experiment
<your-bridge-script>/log-alignment-to-statsig.js (optional, your own — one private copy lives in a separate repo) Bridge scores to Statsig, exports runGenomePipeline()

Building this yourself (if the scripts above don't exist yet)

The first six scripts are a reference implementation. They are not shipped in this plugin, so if they don't already exist in your project, build them from the mechanics described in this file (a hash script, two hooks, three helper scripts) — or, for a faster start, run alignment-harness score per session and keep your own comparison log, skipping the shell-script layer entirely.

Circuit breaker

The genome circuit breaker fires at session end (Stop hook). If average alignment across the last 5 sessions dropped >15% from the experiment baseline, or directed frustration count exceeds 3 in the last 5 sessions, it:

  • Sets circuitBreakerTripped: true in the experiment config
  • Logs a telemetry event
  • Prints a warning to stderr

It does NOT auto-revert — that's destructive without consent.

  • /how-to-create-manage-and-run-statsig-experiments (if you have that skill; otherwise Statsig's own docs cover experiment setup) — Statsig experiment setup
  • /staged-release-management or /staged-release-system (if you have either; otherwise treat this section as reference) — the user-facing equivalent of this system, gradually rolling out a change to real users with the same measure-and-circuit-break pattern
  • /smc-evaluator — independent alignment scoring via an optional NotebookLM oracle (see /alignment-harness:harness-setup if you want to set that up; it also works from your own past sessions without it)
  • /smc-frustration-detector — real-time frustration detection