← the whole session plugin/skills/governer2/SKILL.md
Opt-in upgrade to /governer: same scoring and routing core, plus a notebook built from your own past sessions that can be asked which steps this kind of task has needed before, forced workflow stages the agent can't quietly skip, and full observability (every oracle query and answer printed, every skipped stage logged). Off by default. 'governer2' is worth keeping as a distinct opt-in name rather than folding into /governer until its oracle has actually been exercised enough to judge.
Governer2 — Oracle-Backed, Forced-Stage Variant of /governer
Status: Opt-in, off by default. Master kill switch:
GOVERNER2_ENABLED=false(env var) — when false, this skill REFUSES to run and tells the agent to use/governerinstead. Don't trust this file's own claim about whether it's on or off — check for real (see "Is this actually on?" below); a config file can drift from what a skill file says, and only a live check tells you the truth.Difference from /governer: same scoring + routing core, plus four additive layers:
- Variant announce — every invocation prints in plain language which mode is active and WHEN this stage runs
- Oracle enrichment — when the person has a notebook built from their own history (see "Setting up the oracle" below), query it for what this kind of task has needed before; without one, say so and skip straight to the static rules
- Forced workflow stages — every required skill the oracle (or the static fallback) returns becomes a TaskCreate item the agent cannot mark the task complete without invoking
- Full observability —
ORACLE_INPUT,ORACLE_OUTPUT(full transcript, no summarization),DECOMP_DELTA(changes vs static rules),VIOLATIONS(any forced stage skipped) — all printed to session, and this file says plainly that these are self-reported by the agent, not independently verified by a hook — treat them as an honest transcript, not an audited oneSource of truth for
/governerbaseline:/governer's own skill file. This clone diverges from it ONLY in Phase 0, Phase 1.7, and the Forced-Stage section — everything else runs identically.
Before your output, print ## GOVERNER2_SCORE on its own line. This enables automated extraction. After your governer2 output is complete, print ## END_GOVERNER2_SCORE on its own line.
Is this actually on? (check before trusting this file's status text)
Read, in this order:
- The
GOVERNER2_ENABLEDenvironment variable, if set. - Otherwise, the config file at
$(alignment-harness records governer2)/config.json(create the folder withalignment-harness records governer2if it doesn't exist yet) — look for"enabled": true|false. - If neither exists: off. This is a real opt-in upgrade, not a default — the person turns it on deliberately, the same way they'd opt into any other advanced setup step.
Report the true answer, not what this file's banner says, every time you check.
Phase 0 — Variant Announce (mandatory first action)
Before scoring, print this exact block in plain language. The block tells the person which mode they're seeing and at what stage of the pipeline this fires.
Mode assignment
Read these env vars (or the config file from "Is this actually on?" above):
GOVERNER2_ENABLED— master switch. Iffalseor unset → halt with:🛑 governer2 is DISABLED. Use /governer instead.and stop. Defaultfalse.GOVERNER2_ORACLE_ENRICH— whether to query the person's own oracle.true→ oracle mode (query fires, if a notebook is configured)falseor unset → baseline mode (oracle skipped, /governer's static rules only)random— only meaningful if the person has told you they're deliberately running an A/B comparison of their own (see "Running your own A/B comparison" below); never the default for an ordinary run.
GOVERNER2_FORCED_STAGES— forced-stage enforcement flag.true→ required skills become TaskCreate items + completion gatefalseor unset → required skills are advisory only (legacy/governerbehavior)
Required announce block (verbatim format)
═══════════════════════════════════════════════════════════════════════════════
GOVERNER2 — MODE ANNOUNCE
═══════════════════════════════════════════════════════════════════════════════
This run uses: {Oracle mode — querying your own history notebook | Baseline mode — no oracle configured or oracle enrichment off}
WHEN ACTIVE: GOVERNER2_ORACLE_ENRICH={true|false}
GOVERNER2_FORCED_STAGES={true|false}
Selection method: {env var | config file}
Session ID: {short id}
WHEN IT RUNS: This fires at Phase 1.7 (between scoring complete and skill prediction gate).
Oracle query happens BEFORE the static skill-prediction tables are consulted.
With forced stages on, oracle (or static-fallback) output is REQUIRED before the task can be claimed done.
Master kill switch: GOVERNER2_ENABLED — currently {true|false}. Set to false to revert to /governer.
═══════════════════════════════════════════════════════════════════════════════
After printing, log a telemetry event to a plain local file (no server, this is just an append-only record for your own later review):
echo "{\"ts\":\"$(date -u +%Y-%m-%dT%H:%M:%SZ)\",\"event\":\"Governer2VariantAnnounced\",\"oracleEnrich\":{true|false},\"forcedStages\":{true|false},\"sessionId\":\"{id}\"}" >> .claude/logs/agent-telemetry.jsonl
Do not skip the announce. If you skip it, the person loses observability of which mode they got — the entire point of shipping this alongside /governer is that nothing about how the agent was steered is hidden.
Phase 1 — Score (identical to /governer Phase 1)
Run /governer's own Phase 1 exactly as it's written there: category and floor selection, the mandatory touchesUserState gate, and the local alignment-harness score command that computes the number — no server, no key. Print the same values /governer prints.
Do NOT diverge from /governer Phase 1 logic in this clone. The clone differs ONLY in Phase 0, Phase 1.7, and the Forced-Stage section.
Phase 1.5 — Adaptive Meta-Assessment (identical to /governer)
Run the same meta-assessment from /governer Phase 1.5 — ac, hp, ew self-scoring; critical-mass rule; unified score.
Memory persistence labels — same exact descriptive labels:
degreeToWhichThisTaskAffectsWhatUsersExperiencedegreeToWhichAgentMistakesWouldCompoundAcrossFutureSessionsdegreeToWhichGettingThisWrongWouldSpreadToOtherSystemsdegreeToWhichThisCouldBreakSomethingAlreadyWorkingoverallGovernerScore
Telemetry log (additionally tag via:"governer2"):
echo "{\"ts\":\"$(date -u +%Y-%m-%dT%H:%M:%SZ)\",\"event\":\"GovernerScored\",\"via\":\"governer2\",\"score\":{N},\"feature\":\"{name}\",\"ac\":{ac},\"hp\":{hp},\"ew\":{ew}}" >> .claude/logs/agent-telemetry.jsonl
Phase 1.6: UX-Breakage Reasoning Gate (MANDATORY — runs AFTER Phase 1.5, BEFORE Phase 1.7)
Primary gate. The static category scores are a backstop. This reasoning task is the actual gate on the autonomous-commit lane. It must be performed every single time — no exceptions, no pre-computed answers, no inherited conclusions from a previous invocation.
Before proceeding to Phase 1.7, the agent MUST write prose answers (not yes/no) to the following four questions and print them under the heading ## UX-Breakage Reasoning Gate:
Q1 — Would breakage change or break any user's experience?
"If this code as staged were broken or had a hallucinated bug introduced, would that change or break the experience of any of our users?"
Answer in prose. Consider every execution path that the modified code participates in — not just the obvious "happy path." Think about what calls this code, what depends on its return values, what fails silently if it throws.
Q2 — What UX outcomes could break?
"If yes, in what way? Describe every user-facing UX outcome that could break."
Answer in prose. Be specific about what the user would see, fail to see, or experience incorrectly. Use the user's perspective, not the code's perspective. Example: "avatar upload could fail silently — user submits the form, appears to succeed, but their image is not stored or displayed."
Q3 — What is the risk profile per item?
"What is the risk profile per item — frequency × severity?"
For each item from Q2, rate it. Frequency: how often does this path execute (every session, on first login, only on avatar upload)? Severity: if it breaks, how bad is the UX consequence (user blocked from core feature vs cosmetic glitch)?
Q4 — The comprehensive broken-UX list
"What is the comprehensive broken-UX list — every concrete user-facing thing that could go wrong?"
This is the definitive list used by the binary rule below. It must be numbered. Include items from Q2 plus any additional items discovered while answering Q1-Q3.
Required output format
Print the answers under this exact heading before proceeding:
## UX-Breakage Reasoning Gate
**Q1 — Would breakage affect users?**
{prose answer}
**Q2 — What UX outcomes could break?**
{prose answer}
**Q3 — Risk profile per item**
{prose answer}
**Q4 — Comprehensive broken-UX list**
1. {item}
2. {item}
...
(or: "None — this change has no execution path that reaches any user-facing behavior.")
Binary rule
Count the items in Q4.
- Zero items → change is eligible for the autonomous-commit lane (subject to all existing score-based rules). Continue to Phase 1.7.
- One or more items → change is NOT eligible for autonomous commit. The following are ALL required before any commit:
- Run
/sanity-checkon this change. - Collect browser-based evidence (screenshots or screen recordings of a real browser session exercising the actual UX path) for EVERY item on the Q4 list.
- "Browser-based evidence" explicitly excludes programmatic test passes. A green test suite does NOT satisfy this gate. The evidence must show a real browser exercising the actual user-facing UX — what the user would see, do, and experience.
- Obtain explicit human authorization before any commit.
- Run
Print at the end of this phase:
🚪 UX-Breakage Reasoning Gate: {N items} → {AUTONOMOUS-ELIGIBLE | HUMAN-REQUIRED}
If HUMAN-REQUIRED: list the specific browser-evidence items that must be collected before proceeding.
Phase 1.7 — Oracle Enrichment (NEW)
This phase fires AFTER Phase 1.5 and BEFORE the Skill Prediction Gate.
Baseline path (no oracle, or oracle enrichment off)
If GOVERNER2_ORACLE_ENRICH is not true:
- Print:
Baseline mode — skipping oracle enrichment (not configured, or turned off for this run). - Skip to the existing Skill Prediction Gate (Phase 1.5 → Skill Prediction Gate as in
/governer). - Predicted skills come from
/predict-required-skills(if you have it configured) OR the static fallback table — same as/governer.
Oracle path (enrichment active)
If GOVERNER2_ORACLE_ENRICH=true:
Step 1 — Resolve the oracle's notebook
Read the config file:
cat "$(alignment-harness records governer2)/config.json" 2>/dev/null | node -e "const j=JSON.parse(require('fs').readFileSync('/dev/stdin','utf8')); console.log(j.governer2_training_corpus_notebook_id || 'NOT_SET')"
If NOT_SET:
- Print:
⚠️ No history notebook configured yet. Falling back to baseline mode for this run. See "Setting up the oracle" below, or run /alignment-harness:harness-setup, to build one from your own history. - Continue as baseline mode.
If set:
- Use it as
{notebook_id}below.
Step 2 — Build the oracle query
Compose the input — this MUST be printed verbatim before the query fires:
═══════════════════════════════════════════════════════════════════════════════
ORACLE_INPUT (your own history notebook, {notebook_id})
═══════════════════════════════════════════════════════════════════════════════
SCOPE_SIGNATURE:
task_category: {category}
governer_score_band: {<20|20-59|60-79|>=80}
keywords: {top 5 keywords from first user message}
implicated_subsystem: {repos or feature areas touched}
UNIFIED_SCORE: {N}/100
META: ac={ac} hp={hp} ew={ew}
QUERY:
"Given this scope_signature, based on my own past sessions, return:
(1) Required skills in execution order with timing (early/mid/late)
(2) Correction patterns historically observed for similar scopes
(3) Forced-stage candidates: skills that, if skipped, caused multiple corrections in past sessions
(4) Auto-inject candidates: patterns that recurred 3+ times across sessions
Return as YAML. Cite session IDs for every claim."
═══════════════════════════════════════════════════════════════════════════════
Step 3 — Fire the oracle query
Use whatever oracle-query command you configured during setup (a NotebookLM CLI, an MCP tool, or your own equivalent). For example, if using nlm:
nlm notebook query {notebook_id} \
"{full query string from Step 2}" \
--profile <your profile name>
Or via MCP if available:
mcp__notebooklm__query_notebook({
notebook_id: "{notebook_id}",
question: "{full query string from Step 2}"
})
Step 4 — Print full transcript (NEVER summarize)
═══════════════════════════════════════════════════════════════════════════════
ORACLE_OUTPUT (verbatim, no summarization)
═══════════════════════════════════════════════════════════════════════════════
{paste the full oracle response here, verbatim, including citations}
═══════════════════════════════════════════════════════════════════════════════
If the oracle response was empty, errored, or timed out:
- Print:
⚠️ Oracle returned {error|empty|timeout}. Falling back to /predict-required-skills (if configured) or the static fallback table for this run. - Run the static-fallback path as in baseline mode.
- Tag this run as
oracle_fallback: truein the telemetry log so any comparison you run later can exclude these runs.
Step 5 — Compute decomposition delta
Compare the oracle's required-skill list to the static fallback table from /governer:
═══════════════════════════════════════════════════════════════════════════════
DECOMP_DELTA (changes from static /governer skill list to oracle-enriched list)
═══════════════════════════════════════════════════════════════════════════════
ADDED by oracle (not in static): [/skill-1 — reason from oracle, /skill-2 — reason]
REMOVED by oracle (in static, dropped): [/skill-3 — reason]
REORDERED (timing changes): [/skill-4 moved from late→early]
UNCHANGED: [count: N skills agreed]
ORACLE_CONFIDENCE: {high|medium|low based on oracle citations}
═══════════════════════════════════════════════════════════════════════════════
If DECOMP_DELTA is non-empty (oracle changed something), this is the actionable signal — the oracle is doing real work.
If DECOMP_DELTA is empty (oracle agreed entirely with static), still print it — that's also a signal (the oracle adds no value for this scope class, useful for future calibration).
Skill Prediction Gate (extended for forced stages)
This gate fires after Phase 1.7. Use the same threshold rules as /governer (FIRE if blast_radius>=2 OR user_facing OR complexity>=5 OR score>=40; SKIP otherwise).
When FIRE
Combine sources in order of authority:
- Oracle output (if Phase 1.7 produced one) — highest authority for domain-specific skills
/predict-required-skillsstatic rules, if you have that skill — authority for ALWAYS skills (governer, declare-scope, etc.)/governerfallback table — authority for category-default skills if neither above produced output
The merge rule:
- Oracle output WINS for domain-specific (overrides static for those keys)
- Static ALWAYS rules WIN for foundational (governer, declare-scope cannot be overridden)
- If conflict not resolvable: print as
CONFLICT:line in the Required Skill Sequence section, do not silently pick one
Required Skill Sequence (oracle-grounded, with FORCED-STAGE markers)
Print this format:
## Required Skill Sequence (governer2 — oracle-grounded)
The following skills are REQUIRED for this scope based on historical patterns from your own past sessions.
Each is a forced workflow stage when GOVERNER2_FORCED_STAGES=true.
1. /[skill] — [why, from oracle citation] [EARLY] [FORCED]
2. /[skill] — [why] [MID] [FORCED]
3. /[skill] — [why] [LATE] [FORCED]
N. /consume — admin-consumable output [LATE] [FORCED]
N+1. /commit [LATE] [ADVISORY]
Skipped-skill risk (oracle-flagged): [skills the oracle said are commonly skipped + corrections that resulted]
Skill source: oracle | static | fallback (per skill)
[FORCED] markers mean: see the Forced Stage Enforcement block below.
[ADVISORY] markers mean: included for completeness but not blocking — agent can skip without violation.
Forced Stage Enforcement (only fires when GOVERNER2_FORCED_STAGES=true)
This is the whole reason this clone exists rather than just being /governer with extra logging. Each [FORCED] skill becomes a TaskCreate item the agent CANNOT mark the parent task complete without invoking.
Step 1 — Create TaskList items
After printing the Required Skill Sequence, immediately call TaskCreate for each [FORCED] skill not already in the task list:
TaskCreate({
subject: "FORCED: Run /[skill-name] — [oracle reason]",
description: "Required pipeline stage per governer2's oracle (or static fallback). Must be invoked and marked completed before parent task can be claimed done. Why: [oracle reasoning + citation, or static-fallback reason]. Timing: [early|mid|late]. Skipping logs a VIOLATIONS event."
})
Print after creation:
🔒 Forced stages registered: {N} TaskCreate items added. Each must be marked complete before this scope is claimed done.
Step 2 — Pre-completion verification (agent must run this before claiming done)
Before any "task complete" / "I'm done" / "verification passing" claim, the agent MUST run this self-check:
VIOLATIONS_CHECK:
required_forced_stages: [list of /[skill] from Required Skill Sequence]
invoked_in_session: [list of /[skill] actually invoked in transcript]
missing: [forced stages NOT yet invoked]
status: {PASS — all forced stages invoked | FAIL — N missing}
Note on this check's honesty: this self-check is written by the same agent it's checking — it's a transcript, not an independent audit. Treat a PASS here as "I didn't notice anything missing," not as proof nothing was skipped. If you want a harder guarantee, have a separate agent re-run this check from the transcript alone.
If status: FAIL:
═══════════════════════════════════════════════════════════════════════════════
VIOLATIONS — forced stages skipped
═══════════════════════════════════════════════════════════════════════════════
Skipped:
- /[skill-1] — [oracle reason] — [why this matters]
- /[skill-2] — [oracle reason]
This run is NOT eligible for completion claim. Either:
(a) Invoke the missing stages now and re-run VIOLATIONS_CHECK, OR
(b) Write an entry to $(alignment-harness records governer2)/greenlight-requests.md
explaining why this scope justifies skipping the forced stage(s), and wait for
the person's explicit greenlight before proceeding.
Logging violation event...
═══════════════════════════════════════════════════════════════════════════════
Then log:
echo "{\"ts\":\"$(date -u +%Y-%m-%dT%H:%M:%SZ)\",\"event\":\"Governer2ForcedStageViolation\",\"sessionId\":\"{id}\",\"scope\":\"{scope-summary}\",\"missing\":[{skills}],\"score\":{N}}" >> .claude/logs/agent-telemetry.jsonl
The evidence gate (stop-evidence-gate.sh, part of this harness) already blocks unverified completion claims — the VIOLATIONS print above gives the agent the specific stages to address. The agent can then either invoke them, or file the greenlight-request entry.
Step 3 — Successful completion path
If status: PASS:
✅ All forced stages invoked. Scope eligible for completion claim.
Then proceed to /governer Phase 3 Summary as normal.
Phase 2 — Execute Steps (identical to /governer)
Same as /governer Phase 2: iterate the steps array, run inline actions (confidence-check, ground-check, evidence-grounded-analysis, phantom-scan, completion-reflection), trace each step.
The only difference: when oracle mode is active and the oracle returned forced stages, those forced stages take priority over any conflicting step in the static steps array.
Phase 3 — Summary (extended)
Print the same summary table as /governer, plus a row for this clone's own additions:
| Feature: {name} | Score: {N}/100 |
| degreeToWhichThisTaskAffectsWhatUsersExperience: {N} |
| degreeToWhichAgentMistakesWouldCompoundAcrossFutureSessions: {N} |
| degreeToWhichGettingThisWrongWouldSpreadToOtherSystems: {N} |
| degreeToWhichThisCouldBreakSomethingAlreadyWorking: {N} |
| overallGovernerScore: {N} |
| GOVERNER2 MODE: {oracle|baseline} | Forced stages: {true|false} |
| Forced stages required: {N} | Invoked: {N} | Skipped: {N} | Violations logged: {N} |
| Step | Certainty | Status | Notes |
| ... | ... | ... | ... |
Autonomous Commit Rule (extended for forced stages)
Same as /governer, plus:
| Condition | Commit? |
|---|---|
| UX-Breakage Reasoning Gate produced ANY items (Q4 non-empty) | NO — Sanity-check required + browser evidence per item + explicit human authorization |
| Score < 40 AND all tests pass AND zero forced-stage violations | YES — commit autonomously |
| Score 40–69 AND certainty >= 90% AND zero violations | YES — commit + trace |
| Score >= 70 | Queue brief, continue. Human batch-approves. |
| ANY forced-stage violation logged | NO — fix violations first, or write a greenlight-request entry and wait |
| Any test fails | NO — fix first or queue brief |
| Payment/auth/email code touched | Queue brief regardless of score |
Kill switch + revert
To turn governer2 off entirely:
export GOVERNER2_ENABLED=false
# OR remove the config:
rm "$(alignment-harness records governer2)/config.json"
To revert to plain /governer behavior fully:
/governer2is invoked only via explicit/governer2command./governerremains the default.- Nothing routes to
/governer2automatically unless the person has deliberately wired it in during setup.
To run a single one-off observation without enabling persistently:
GOVERNER2_ENABLED=true GOVERNER2_ORACLE_ENRICH=true GOVERNER2_FORCED_STAGES=true /governer2 {task description}
Setting up the oracle
The one thing that makes this "governer2" rather than "governer with extra logging" is a notebook built from your own past sessions. To build one:
- Run
/alignment-harness:harness-setup(the oracle checkpoint) or do it directly: read~/.claude/projects/*/*.jsonlfor this project, pull out which skills your own past tasks used, where you corrected an agent, and what got skipped. - Show yourself a summary and a few real examples before uploading anything.
- Only after confirming, upload to NotebookLM (or your own equivalent) and record the notebook id in
$(alignment-harness records governer2)/config.json:
{
"governer2_training_corpus_notebook_id": "<your notebook id>",
"default_variant": "oracle",
"enabled": true
}
Without this, governer2 still runs in baseline mode — it scores the task, runs the UX-breakage reasoning gate, and enforces whatever the static fallback table says the task needs. It just says, every run, that the oracle step was skipped and why. It never claims the oracle answered when it didn't.
Running your own A/B comparison
If you specifically want to measure whether the oracle helps — not just use it — set GOVERNER2_ORACLE_ENRICH=random (or alternate it deliberately across sessions) and compare correctionsPerSession or similar between the two arms over time. This is optional measurement scaffolding for someone deliberately running an experiment; it is not what a person gets by default just for turning oracle mode on.
Configuration file format
Path: $(alignment-harness records governer2)/config.json
{
"governer2_training_corpus_notebook_id": "<your notebook id, or omit>",
"default_variant": "oracle|baseline",
"ab_split_method": "always_a|always_b|random|env",
"ab_seed_field": "sessionId",
"violations_log_path": ".claude/logs/agent-telemetry.jsonl",
"enabled": false
}
enabled: false is the master kill switch in config form (env var takes precedence).
Why this clone exists (for future agents reading this)
This exists to test whether a notebook built from your own history actually improves agent behavior, without disturbing the live /governer baseline everyone gets by default. It lets you run side-by-side comparisons — oracle mode vs baseline — and decide, from real evidence, whether the oracle's suggestions should ever become load-bearing rather than advisory.
This is the operationalization of two governing rules from this skill's design: each required skill that comes back becomes an actual forced workflow stage, not something agents can skip; and observability happens — what got sent to the oracle PRINTED, what got returned FULL TRANSCRIPT PRINTED, when printed, any actual changes in decomposition printed, any agent failing to abide by it caught and logged and printed to session automatically.
End of /governer2 SKILL.md. To run: export GOVERNER2_ENABLED=true GOVERNER2_ORACLE_ENRICH=true GOVERNER2_FORCED_STAGES=true then invoke /governer2 {task}.