← the whole session plugin/skills/deploy-alignment-guardian/SKILL.md

Deploy a background agent that keeps a long, multi-agent project's work checked against the person's confirmed intent for its entire duration, not just at dispatch. Use when a task is scored high-risk/high-scope (governer >= 50), or when the person requests persistent alignment monitoring.

Deploy Alignment Guardian

Purpose: on a big project — many tasks, several agents, hours of work — the agents doing the work are handed the intent once, at dispatch, and slowly lose sight of it as they go. This stands up a separate background agent whose only job is to keep re-reading the confirmed intent and checking recent work against it, so drift gets caught while it's happening instead of after it's been built on.


Step 1 — Make sure a project seed exists

A "seed" here means a short written record of the person's confirmed intent for this specific project — what it should make possible, what success looks like. The guardian has nothing to monitor against without one.

Look for it in this order:

  1. The most recently modified file in this project's own seed folder (e.g. docs/intent/seeds/, or wherever /align writes them for you).
  2. A path recorded in the current session's governer state, if your setup stores one.
  3. Explicitly provided by the person or team lead.

If no seed exists yet, stop and run /align first — that's the skill that works out and confirms what the person actually wants, and writes the seed once they've confirmed it. Do not deploy the guardian without a seed.

Step 2 — Gather dispatch context

Collect:

  • SEED_PATH — absolute path to the confirmed seed document.
  • PROJECT_NAME — the project or team identifier.
  • RISK_SCORE — the score that triggered this deployment, if one exists (alignment-harness status); otherwise note "manual deployment, no score."

Step 3 — Dispatch the guardian

If your setup has a dedicated persistent agent type for this (a custom Opus agent registered as alignment-guardian is one example), use it. Most people won't have that registered — in that case, dispatch a general-purpose background agent and put the whole role in the prompt itself, so nothing depends on a pre-registered agent type:

You are the Alignment Guardian for this project. Your only job is watching, never building.

Seed path: {SEED_PATH}
Project: {PROJECT_NAME}
Risk score at dispatch: {RISK_SCORE}

Your job:
1. Read the seed document now and hold onto its core intent statements verbatim — don't paraphrase
   them into your own words, since a paraphrase drifts too.
2. Periodically (aim for roughly every 5 minutes of session time, or whenever you're re-invoked —
   see "how the cadence actually runs" below), re-read the seed and compare it against any task
   completions or agent outputs that have appeared since your last check.
3. Watch specifically for:
   - binary reductionism — an aspiration being squeezed into "it IS this, NOT that," losing the
     nuance the person actually stated
   - work marked done because the code runs, when what the person was meant to experience isn't
     actually there yet
   - new scope that doesn't trace back to anything in the seed
4. When you flag drift, quote the agent output, quote the seed line it breaks, and say what the
   aligned version would look like. Never just raise an alarm without a proposed fix.
5. When intent is confirmed to have grown (a new /align run, or explicit confirmation from the
   person), draft an ADDITION to the seed — never overwrite or delete existing lines — and hand
   it to the team lead for approval before anything gets written.
6. You do not talk to the person directly. Everything goes through whoever dispatched you.
7. You do not block work. You flag and propose. Never gate a commit or a task yourself.

Begin by reading the seed and sending one ALIGNMENT CHECK message to confirm you're active and
the seed is readable.

On the model: this job is long-context sense-making across everything that's happened in the project so far, which benefits from your strongest available reasoning model. Use an Opus-class model if you have access to one; otherwise the default model still does the job, just with less headroom on very long sessions — don't block deployment on having a specific model available.

On the cadence: a background agent can't reliably keep its own wall-clock timer running unmonitored. In practice this works one of two ways: (a) the lead agent pings the guardian with a short "anything to report?" message every so often as part of its own loop, or (b) the guardian re-checks whenever it's given a batch of new task completions to look at. Pick whichever fits how you're orchestrating the rest of the project, and say in the deployment confirmation which one you're using.

Step 4 — Confirm deployment

Print to chat:

🛡️ Alignment Guardian deployed — monitoring {SEED_PATH}
   Project: {PROJECT_NAME}
   Risk score: {RISK_SCORE}
   Cadence: {how it's being triggered — see above}

Step 5 — Keep a record it's actually working

Nothing about a background agent tells you on its own whether it's alive. Append each check it reports to a durable log — alignment-harness records alignment-guardian if you don't have anywhere better — with a timestamp and whether drift was found. If you haven't seen a check message in the interval you expect, treat that as broken, not as "nothing to report" — redeploy per below.

When to redeploy

Redeploy (restart the guardian) if:

  • The seed document path changes.
  • The project scope expands significantly (a new, higher risk score).
  • The guardian has gone silent for noticeably longer than its expected check interval with no check message.

Config worth setting per project

Nothing here is hardcoded — treat these as dials:

  • on/off for this project
  • check interval / trigger style (see Step 3's cadence note)
  • the risk-score threshold that makes you consider deploying this automatically (a reasonable starting default is 50, same as the governer's own "deeper reflection" band)

What the guardian does NOT do

  • It does not write to the seed without the team lead's approval — it only ever proposes additions, never overwrites.
  • It does not communicate directly with the person — only with whoever dispatched it.
  • It does not score or evaluate code quality — only intent alignment.
  • It does not block work — it flags and proposes, never gates.