← the whole session plugin/skills/agent-debug/SKILL.md

Print a diagnostic overview showing what harness hooks fired, what got injected into context, what's silent, and what's shaping agent behavior — before every response. Use when the person wants to see what's actually happening behind the scenes.

Agent Debug — Diagnostic Visibility Mode

When this skill is loaded, you are in diagnostic mode for the rest of the session.

Before EVERY response you give (not just the first one — every single one), print the diagnostic block below. This is the first thing the person sees before your actual reply. It shows them what's happening behind the scenes so they can evaluate what's useful and what's noise.

This piece exists because most of what the alignment harness does happens through text injected into your context that the person never sees — a coherence-check instruction, a governer score, a memory-search result — arriving as tagged blocks alongside the person's own message, indistinguishable from it unless something points them out. A broken safeguard that looks like it's working is worse than an absent one, because the person keeps trusting it. This piece answers "is it set up, and is it actually working right now?" by reading the real files and the real command output, live, every turn — never by describing what the harness is supposed to do.

Before you start: find the real paths

Run once per session (not every turn):

alignment-harness paths

This prints the data folder the rest of these checks read from. If alignment-harness isn't on the PATH, use the full path the session-start message gave you (node ".../plugin/bin/alignment-harness" paths). Everything below calls it as alignment-harness.

What to Print

Print this block before every response:

🔬 DIAGNOSTIC — This Turn

Then for each of these, show what ACTUALLY happened this turn — not what the code says should happen, what actually happened. If something produced no output, say so. If something is broken, say so. If you can see the literal text that was injected, show a short excerpt (first 2-3 lines, not the whole thing).

1. UserPromptSubmit hook

Check the local telemetry log for the most recent entries (this is a plain JSONL file on this machine — no server involved):

DATA_DIR=$(alignment-harness paths | sed -n 's/^data: //p')
tail -10 "$DATA_DIR/telemetry.jsonl" 2>/dev/null

Show whether a UserPromptSkipped (trivial acknowledgement) or the substantive path (governer banner + coherence gate armed + memory search) fired this turn. Example output:

UserPromptSubmit: skipped (trivial acknowledgement, e.g. "ok") — no enrichment reached the agent this turn

or

UserPromptSubmit: real message — CoherenceGateArmed, governer score banner injected, MemorySearch ran (resultLen=0 → not configured or no hits)

If the telemetry file doesn't exist yet, say so plainly — it means no session has started with the plugin enabled, or telemetry is turned off in the switches file.

2. Coherence gate

Run:

alignment-harness status

Read the coherence gate: line (its strength, or off). For whether it's armed right now, look at the most recent CoherenceGateArmed / CoherenceGateAttested / CoherenceGateBlocked / CoherenceGateDowngraded line in the telemetry tail from check 1 — whichever is most recent tells you the live state.

Coherence gate: block strength, ARMED this turn (tools that change things wait until COHERENCE_CHECK + attest)

or

Coherence gate: off (turned off in the switches file)

3. Governer state

Run:

alignment-harness status

Show the score: line verbatim (it already says whether it's scored, by whom, and whether it's calibrated to the person's own risk table or using general defaults), plus the verification-contract line if one is printed.

Governer: 65/100 (scored by the agent, uncalibrated) — verification contract: 1/3 verified

or

Governer: none — unscored (no gate has a number to work from yet)

Never read this from a raw state file path directly — the per-session file's exact location can change between versions; alignment-harness status is the stable way to ask.

4. System reminders

Count how many system-reminder blocks you received this turn. You can see these — they're the <system-reminder> tags in your context. Count them and note what each one contains (skill list, explanatory style reminder, task reminder, etc). Show a 1-line summary of each.

System reminders this turn: 3
  - Skill list (available skills)
  - "Explanatory output style is active"
  - "Task tools haven't been used recently"

4.5. Ghost injections (CRITICAL — the person cannot see these)

These are blocks of text injected by hooks into your context that appear alongside the person's message but were NOT typed by them. The person CANNOT see them in the UI — they are invisible to everyone except you. This is a primary source of agent drift: injections that look like human instructions but aren't.

For EACH distinct injected block you received this turn, print:

GHOST {n}: <{tag_name}> — {one sentence describing what this injection tells you to do}
  Intent: {why this injection exists — what alignment problem it was designed to solve}
  First 3 lines: {literal first 3 lines of the injected text, verbatim}
  Source: {which hook produces this}

The tags this plugin's own hooks actually produce (verify this list still matches the installed hooks — see the note below):

  • <hook-injected-context source="alignment-harness" event="UserPromptSubmit"> — the wrapper around everything below, from hooks/user-prompt.js
  • <human-message priority="primary"> — the person's own stripped message, re-presented so you weight it over stale context — hooks/user-prompt.js
  • <calibrate-injection priority="high"> — fires only when the person's words match a confusion pattern ("I don't understand what you're saying") — hooks/user-prompt.js, tells you to run /alignment-harness:calibrate
  • <institutional-context source="memory-search"> — only appears if a memory-search command is configured (/alignment-harness:harness-setup) and found something — hooks/user-prompt.js
  • [coherence_check start] ... [coherence_check end] — the per-message understanding ritual, only when the coherence gate is on — hooks/user-prompt.js
  • <governer-alignment-gate score="..." scoredBy="..." confidenceFloor="..."> — the "state your understanding first" prompt (or a plain <governer-score> line if that's turned off) — hooks/user-prompt.js
  • <task-tracking-gate> — reminds you to create tasks for each stated requirement — hooks/user-prompt.js
  • <verification-contract-gate score="..."> — only injected once the task is scored 20+ — hooks/user-prompt.js
  • <evidence-gate status="BLOCKED" reason="..."> — appears in a tool result, not context, when the Stop hook sends you back — hooks/stop.js
  • A plain BLOCKED — ... message (no XML tag) from the coherence or governer gate on a tool call — hooks/pre-tool.js

If you received NO ghost injections, say so explicitly — that itself is diagnostic (it usually means the harness hooks aren't installed or enabled for this session, or the message was a trivial acknowledgement that skips enrichment).

Ghost injections this turn: 3
  GHOST 1: <hook-injected-context> — wrapper for everything the UserPromptSubmit hook adds
    Intent: mark harness-injected text as distinct from the person's own words
    First 3 lines: '<hook-injected-context source="alignment-harness" event="UserPromptSubmit">'
    Source: hooks/user-prompt.js
  GHOST 2: [coherence_check start] — forces a plain-language understanding check before tools that change things
    Intent: catch a misreading before it becomes work that has to be undone
    First 3 lines: "When the person speaks, print COHERENCE_CHECK on its own line first..."
    Source: hooks/user-prompt.js
  ...

A note on staleness: this tag list was accurate against the hooks shipped with this version of the plugin (hooks/user-prompt.js, hooks/pre-tool.js, hooks/stop.js). Hook text can change between releases. If you notice a tagged block that isn't in this list, or one listed here that never appears, say so under "Anything unexpected" (section 7) rather than silently trusting this list forever — it is a snapshot, not a live contract.

5. Skill files — what's loaded and how

Break this into three parts:

Auto-loaded this turn — any skill files that were loaded without the person explicitly asking. This includes skills triggered by hooks, plugins, or the harness telling you to invoke something. If nothing was auto-loaded, say "none."

Auto-loaded this turn: none

Manually loaded this turn — skill files the person explicitly invoked with /commands or that you loaded via the Skill tool because they asked.

Manually loaded this turn: agent-debug (person ran /agent-debug)

All skills loaded this session (running total) — every skill that's been loaded at any point during this session, in order. This accumulates across turns.

Session skill history: reflect → governer → agent-debug

6. Static context

One line each for the big static files actually present in this project. Don't print their content — just confirm they're loaded and roughly how much space they take:

Static context:
  - CLAUDE.md: loaded (~800 lines)
  - project memory file: not present

If a file this section usually mentions doesn't exist in this project, say "not present" rather than omitting the line — an absent file is itself diagnostic.

7. Anything unexpected

If you notice anything in your context that doesn't fit the categories above — an injected tag you don't recognize, a system message that seems new, a plugin doing something you haven't seen before — flag it here. If nothing unexpected, say so.

Unexpected: nothing this turn

or

Unexpected: received a tag not in section 4.5's list — the installed hooks may be a newer version than this skill's snapshot; worth re-checking against hooks/*.js.

Format

Keep the whole diagnostic block tight. The person wants signal. Each line should be one thing, one status. No paragraphs. No analysis. No opinions about what should be different. Just what IS.

After the diagnostic block, print a horizontal rule (---) and then give your actual response to the person's message.

Rules

  • Print this EVERY turn. Not just the first one. Every single response starts with the diagnostic.
  • Run the bash checks (telemetry, alignment-harness status) in parallel to keep it fast.
  • Do NOT add interpretation. "Hook skipped" is diagnostic. "Hook is broken and this means alignment is degraded" is interpretation — don't add it.
  • Do NOT skip the diagnostic because the message seems trivial. Even "hey" gets the diagnostic. That's the point — you see what fires on every kind of message.
  • If running the bash checks would slow things down noticeably, you can read from your most recent check and note "(cached from last check)" instead of re-running.