← the whole session plugin/skills/pipeline-compliance-audit/SKILL.md

Post-pipeline compliance gate for the `/process-actionable` pipeline. Audits every run's output files for every item in a batch — verifying that Notice ran with institutional-memory search and separated observed-from-imagined, that Score captured a full governer breakdown, that Verify showed real command output at the right depth, that the oracle gate (if configured) was queried, and that the final proposal is zero-jargon with a stated scope-stop. Outputs a compliance report with per-stage PASS/FAIL. BLOCKS completion if any stage has FAIL. Designed to be invoked every few minutes by a compliance agent during a live pipeline run, or once at the end of a batch.

Pipeline Compliance Audit

This audits the output of the /process-actionable pipeline. It has nothing to check until that pipeline has produced at least one run — if no run directory exists yet, say so plainly and stop, rather than auditing an empty batch as if it passed.

Why This Exists

In an earlier version of this kind of pipeline, an oracle quality gate meant to run on every item before it reached the person for review actually ran on only 5 of 31 items — silently, and nothing caught it. The pipeline produced output that looked complete but wasn't. The person reviewing it couldn't tell which items had actually gone through which gates and which had slipped through.

This skill exists so that skipping a stage is not just wrong — it is impossible to miss. Every item, every stage that actually exists in the pipeline being audited, every requirement. Before anything is marked "ready for review," this audit must pass.

This checklist is written against /process-actionable's real, current stage list (Stage 0 through Stage 5.5, plus a post-pipeline completion step) — not an idealized or invented one. If your project's pipeline skill uses different stage names or file names, adapt the checklist below to match what it actually writes; auditing for files a pipeline never produces isn't auditing anything.


What This Audits

Every item in the run must be checked against every stage its own run actually has files for. The audit is not complete until every item has a row in the output and every row has a verdict.

Stage 0 & 0.5 — Source Verification & Known-Decisions Check (inside 01-notice.md)

Requirements:

  • source trace present — an artifact id, file path, or explicit "agent observation — unverified" note. Not silently absent.
  • low-certainty ceiling applied when unverified — if the source couldn't be confirmed, the notice states an explicitly lowered certainty ceiling and says why, rather than treating an unconfirmed source as solid.
  • known-decisions check, if configured — if this project has a registry of prior decisions on recurring signals, it was checked, and a prior undismissed rejection stopped the pipeline rather than re-running it. If no such registry is configured for this project, mark this row N/A — not configured rather than FAIL.

FAIL conditions:

  • No source trace at all, and no explicit "unverified" note either
  • An unverified source treated with the same certainty as a confirmed one
  • A known-decisions registry is configured but wasn't checked

Stage 1 — Notice (01-notice.md)

Requirements:

  • certainty dots present — every claim is followed by a marker and a percentage. No bare assertions without one.
  • imagined vs observed separated — the output explicitly labels which statements are observed vs inferred/predicted. These are never blended.
  • institutional-memory search ran at least once — if this project has memory/semantic search configured (agent_find or an equivalent), the output shows a visible result block or explicitly states "no results for [query]." If nothing is configured, the notice says so rather than silently skipping the step.
  • source traced — a source path, artifact ID, or document link is present that an agent in Stage 2 can actually open and read.
  • key uncertainty identified — the one question that would most reduce uncertainty is named explicitly.

FAIL conditions:

  • Any claim written as a bare fact (no marker, no %)
  • Observed and imagined sections merged or absent
  • Memory search configured but no evidence it was called
  • Source is described but not linked or pathed

Stage 1.5 — Declare Scope (in the Stage 2 handoff inside 01-notice.md)

Requirements:

  • in-scope stated — the specific claim being investigated, in plain terms of what a person experiences
  • out-of-scope stated — adjacent issues noticed but not being pursued in this run
  • assumptions stated — what must be true for this claim to matter

FAIL conditions:

  • Scope section absent from the handoff
  • Adjacent issues mentioned in the notice but not carried into an explicit out-of-scope line

Stage 2 — Score (02-score.md)

Requirements:

  • touchesUserState gate ran — the output explicitly shows whether this affects what a user sees/receives/experiences, and which impact tier was assigned. If skipped, the score is invalid.
  • full scoring dimensions shown — every scoring dimension this project's governer actually uses appears with a score and brief reasoning, not just a final number.
  • adaptive meta-assessment scored — every adaptive dimension this project's governer defines appears with a score.
  • depth routing assigned, matching the score — the output states which depth was assigned (skip / quick / moderate / full, or this project's equivalent labels) and it matches the unified score against the project's own thresholds.
  • unified score stated and traceable — a single score is present and its inputs are visible, not just asserted.
  • "impact IF TRUE" framing — the scoring explicitly treats the claim as unverified at this point; it does not treat it as confirmed.
  • found-something-worse check — the output either flags a higher-priority concern discovered during scoring, or explicitly states none was found.

FAIL conditions:

  • touchesUserState absent or not explicitly stated
  • Scoring dimensions present as a bare number with no reasoning, or fewer dimensions than this project's governer defines
  • Depth assigned doesn't match the score against this project's own routing table
  • Score stated without the inputs that produced it

Stage 2.5 — Speak-Human Translation (appended to 02-score.md)

Requirements:

  • the score's mechanics are translated into plain language — not a restatement of the governer's own jargon, an actual translation: what this score means for the decision at hand, in a sentence a non-technical reader could act on.

FAIL conditions:

  • The scoring output has no plain-language translation section at all
  • The "translation" just repeats dimension names and numbers back

Stage 3 — Verify (03-verify.md)

Requirements:

  • actual commands with actual output — every verification step shows the exact command run AND the exact output received. Summaries of what a command "would show" don't count.
  • observation / inference / prediction separated — labeled distinctly.
  • negative check done for score ≥ 60 — a "what would I expect to see if this claim were completely false?" check, run for real, with the result reported.
  • depth honored — matches what Stage 2 assigned.
  • browser/UX evidence gate honored, when touchesUserState = true — an attempt at real browser/product verification is documented; if it couldn't reproduce the issue, certainty is explicitly capped and the attempt is described, rather than substituting a simulated state for real evidence.
  • emergent findings captured — anything discovered during verification that wasn't in the original claim is flagged explicitly. "Nothing new found" is acceptable — silence is not.
  • correct environment named — verification states which environment/data source it actually queried (for example: real/production data vs. a local stub or test fixture), so a claim about live behavior isn't quietly verified against a stub.

FAIL conditions:

  • Commands listed without output
  • "Verified by reading the code" without showing the code actually read
  • Negative check absent for score ≥ 60
  • touchesUserState = true and no browser/product-level verification attempted at all
  • No statement of which environment/data source was used

Stage 3.5 — Sanity Check (03.5-sanity-check.md)

Requirements:

  • checked against the original notice — does the verified claim still match what Stage 1 said? Divergence noted if not.
  • checked for a bigger emergent finding — is there something from verification that matters more than the original claim? Flagged if so.
  • checked whether verification disproved a scoring assumption — and whether the score should be revised.
  • cross-referenced against other runs — has this claim, or something adjacent, come up in another run directory? Named if so.

FAIL conditions:

  • File missing entirely for an item that reached Stage 4
  • Present but doesn't address any of the four checks above — a placeholder, not a real sanity check

Stage 4 — Propose (04-propose.md or 04-closed.md)

Requirements:

  • reflects the verified claim, not the original — if they diverged, the proposal states the divergence explicitly.
  • certainty doesn't exceed Stage 3's verified certainty
  • every affected user journey walked through — what the person does, what happens now, what the gap is, what happens after the fix — in the language of experience, not code.
  • zero jargon in the human-facing section — no variable names, code references, or unexplained acronyms. Test: could a non-coder read this and know exactly what's happening to whom?
  • risk stated in UX terms — what a real person would experience if the fix fails, not a technical risk statement.
  • scope-stop explicit — what this does NOT fix, never omitted.
  • admin questions pre-answered — what this fixes, who it affects, what could go wrong, how you'd know it worked, what's not covered.
  • sources linked — the proposal's frontmatter or references point back to the notice, score, and verify files, plus the original source.
  • quality-gate skill run on the draft, if this project has one (for example /align) — before finalizing.

FAIL conditions:

  • Proposal uses the original claim's numbers after Stage 3 changed them, with no divergence noted
  • Jargon present in the human-facing section
  • Scope-stop absent
  • No links back to the stage files that produced the claims

Stage 4.5 — Oracle Gate (04.5-jonathan-check.md, if this project has one configured)

If this project has a "predict how the reviewer will react" oracle configured (one common setup uses /jonathan-check2; generalize the name to whatever this project's own reviewer-prediction tool is) — check:

  • oracle queried — the actual call is visible, not just asserted.
  • question logged verbatim — not a paraphrase.
  • answer logged at auditable length — "it said it looks good" doesn't count.
  • predicted challenges listed and addressed — each one shows either a fix, evidence that already covers it, or an honest "not applicable" with reasoning.

If no such oracle is configured for this project, mark this stage N/A — oracle not set up rather than FAIL — see /alignment-harness:harness-setup for how to configure one.

FAIL conditions (only when an oracle IS configured):

  • Not run at all — this was the original failure this skill exists to prevent
  • Queried but the question or answer isn't logged
  • A predicted challenge is listed but never addressed

Stage 5 — Heal (05-heal-stage{N}-{date}.md, on-demand — only check if this run received reviewer feedback)

Requirements, when this file exists:

  • feedback quoted verbatim
  • root cause identified in the stage's own prompt/instructions, not just the symptom
  • updated instructions shown as a diff
  • regression-checked against at least one or two prior examples, not just the one that triggered the fix
  • quality-gate skill run on the healer's own output, if this project has one

FAIL conditions: present but missing the root-cause diff, or no regression check at all.


Stage 5.5 — Cross-Reference Update (05.5-cross-reference-update.md, on-demand)

Only produced when a cross-reference actually turned something up. When it exists:

  • names what was found in the other run
  • states the impact on this proposal
  • states whether the score or the proposal needs revision

FAIL conditions: file exists but doesn't answer one of the three points above.


Post — Completion Record

  • a completion record was filed — through whatever this project's own end-of-task reporting skill is (for example /complete-agentic-task), or, if none is configured, as a plain file under alignment-harness records pipeline-completions — so a finished run doesn't just end silently.

FAIL conditions: run reached Stage 4 (proposed or closed) with no completion record anywhere.


Per-Item Depth & Evidence Contract Enforcement

This section checks that the specific verification depth /process-actionable's own scoring stage prescribed for THIS item was actually honored — not a separate invented gate, but the pipeline's own rules, read back and confirmed.

How to read what was required

# The depth and score are in the Stage 2 output:
grep -i "depth\|unified score\|touchesUserState" $RUNDIR/02-score.md

Checks

This item's Stage 2 result Required in Stage 3 (03-verify.md) Status Gap (if FAIL)
Score < 20 (skip) A one-line note that scoring was too low for investigation — nothing more ✅/❌
Score 20–59 (quick) At least one targeted command and its real output ✅/❌
Score 60–79 (moderate) The main claim plus at least 2 stated assumptions checked, each with a command and output ✅/❌
Score 80+ (full) Every claim checked, source code read where referenced, negative cases run ✅/❌
2 of 3 adaptive-meta dimensions ≥ 8 (critical mass) Unified score treated as ≥ 80 regardless of the raw number, and Stage 3 ran at "full" depth ✅/❌
touchesUserState = true A real browser/product-level verification attempt is documented in Stage 3, with certainty explicitly capped if it couldn't reproduce the issue ✅/❌

Mark a row N/A only when its condition genuinely doesn't apply to this item (for example, the critical-mass row is N/A when no adaptive dimension reached 8). Never mark a row N/A to avoid grading it.

FAIL conditions:

  • Stage 3's actual depth is shallower than what Stage 2's own score required
  • The critical-mass rule applied in Stage 2 but Stage 3 still ran at a shallower depth
  • touchesUserState = true with no browser/product verification attempt anywhere in Stage 3

How to report this in the compliance output

Add a block per item after its stage checklist table:

### Depth & Evidence Contract — {item title}

**Score:** {N} (from 02-score.md)  **Depth assigned:** {skip/quick/moderate/full}  **touchesUserState:** {true/false}

| Requirement | Status | Gap (if FAIL) |
|---|---|---|
| Stage 3 depth matches Stage 2's assignment | ✅ PASS / ❌ FAIL | {gap} |
| Critical-mass rule honored (if triggered) | N/A / ✅ PASS / ❌ FAIL | {gap} |
| Browser/product verification attempted (if touchesUserState) | N/A / ✅ PASS / ❌ FAIL | {gap} |

**Contract result: PASS / BLOCK**

A FAIL here blocks the item exactly like a stage-checklist FAIL — include it in the item's FAIL count.


How to Run This Audit

Step 1: Locate the run directory

# Runs live wherever this project's /process-actionable writes them — check that
# skill's own "Run directory" section for the exact path. Typically something like:
ls <project-repo>/docs/intent/experiments/process-actionables/runs/

RUNDIR=<project-repo>/docs/intent/experiments/process-actionables/runs/{run-id}

If that directory doesn't exist yet, there's nothing to audit — say so and stop, rather than reporting an empty batch as passing.

Step 2: Identify all items in the batch

ls $RUNDIR/                # files for one item
ls <...>/runs/             # all run directories, for a batch

Step 3: Check each item against each stage's checklist

For each item and each stage that item's run actually has a file for:

  1. Read the stage file.
  2. Check every requirement in the checklist for that stage.
  3. Mark PASS or FAIL with a specific note for each FAIL (what's missing, where).

Step 4: Write the compliance report

Save it inside the run directory (for one item) or as a batch-level file next to the run directories (for a batch) — name it clearly, e.g. compliance-{date}.md.

Step 5: Block or pass

  • If any stage has FAIL for any item: print the BLOCK message and don't mark anything "ready for review."
  • If all stages PASS (or are legitimately N/A) for all items: print the PASS message.

Output Format

# Pipeline Compliance Audit — {run-id or batch label}
**Date:** {YYYY-MM-DD HH:MM}
**Items audited:** {N}
**Overall result:** PASS / BLOCK

---

## Item: {item title}

| Stage | Requirement | Status | Gap (if FAIL) |
|-------|-------------|--------|---------------|
| 0/0.5: Source & Known-Decisions | Source trace present | ✅ PASS | — |
| 1: Notice | Certainty dots present | ✅ PASS | — |
| 1: Notice | Institutional-memory search ran | ❌ FAIL | No result block in 01-notice.md |
| 2: Score | touchesUserState gate ran | ✅ PASS | — |
| 3: Verify | Actual commands + actual output | ❌ FAIL | 03-verify.md lists commands but shows no output |
| 3.5: Sanity Check | Cross-referenced against other runs | ✅ PASS | — |
| 4: Propose | Scope-stop explicit | ❌ FAIL | "What this does NOT fix" section absent |
| 4.5: Oracle Gate | Oracle queried | N/A | No oracle configured for this project |

**Item result: BLOCK**
**FAIL count: 3**

---

## Batch Summary

| Item | Stages present | FAIL count | Result |
|------|---------------|-----------|--------|
| {item} | 0,1,2,3,3.5,4 | 3 | BLOCK |
| {item} | 0,1,2,3,3.5,4,4.5 | 0 | PASS |

**Items passing: {N}/{total}**
**Items blocked: {N}/{total}**

BLOCK Message

PIPELINE COMPLIANCE BLOCK

{N} item(s) have stage failures. Nothing in this batch can be marked "ready for review" until all FAILs are resolved.

Items blocked: {list}
Total FAILs: {N}

Most common failures:
1. {stage}: {what's missing} — {N} items affected
2. {stage}: {what's missing} — {N} items affected

To resolve:
- Return each FAIL to whichever pipeline stage produced it, with the specific gap listed
- Re-run only the failed requirement, not the whole item
- Re-run this audit after fixes are applied

Do NOT advance any item past its current stage until this audit shows PASS.

PASS Message

PIPELINE COMPLIANCE PASS

All {N} items have completed every stage their run required, with no failures.
Every configured oracle gate was queried where applicable. Every verification command showed actual output.

This batch is ready for review.

Running This on a Timer

This skill can be invoked on a repeating loop during an active pipeline run, for example (if this project has a /loop skill):

/loop 5m /pipeline-compliance-audit — check all items in the current run that have reached Stage 4 or later, verify all prior stages passed, update the compliance report

What the compliance agent does on each tick:

  1. Check which items have reached Stage 4 (Propose) or later.
  2. For each, run the full per-stage checklist.
  3. Note PASS items in the compliance log; don't re-check them on later ticks unless their files changed.
  4. For any FAIL, immediately output a BLOCK notification with specific gaps.
  5. Update the compliance report with the current tick's findings.

Without a /loop-style skill, just re-run this audit manually at natural checkpoints (after each item reaches Stage 4, or once at the end of a batch).


Critical Rules

  1. Every item, every stage that actually exists for it — the audit is not a sample. The original failure this skill was built to prevent was 5 of 31 items getting the oracle gate; this never happens under this skill.
  2. The oracle gate is mandatory for every item, once an oracle is configured — not high-score items only. Before that, it's N/A, never silently skipped as a pass.
  3. A FAIL blocks the whole item — a single FAIL in an early stage blocks that item from being marked ready, even if later stages look perfect, because later stages were built on the incomplete earlier one.
  4. "Looks complete" is not compliance — read the actual stage files and check specific fields. An agent saying "I ran the search" without the result block present is a FAIL.
  5. This audit is self-contained — the compliance agent reads files only; it doesn't share context with the pipeline agents that produced them. That's what makes it independent.

  • /process-actionable — the pipeline this skill audits. This file's checklist must be kept in sync with that skill's actual stage list; if that skill's stages change, update this checklist to match, don't leave it checking for files that no longer exist.
  • /align — the quality gate this project's pipeline invokes at Notice and again before finalizing a proposal or a heal. Never modify it from here.
  • /jonathan-check2 (or this project's equivalent) — the oracle invoked at the Oracle Gate stage, if configured.
  • /governer — invoked at the Score stage.

What This Skill Does NOT Do

  • It does not fix failures — it identifies them and blocks advancement. Whichever stage produced the failing requirement is where the fix belongs.
  • It does not judge the quality of stage outputs — it checks for presence and format. Whether the content itself is good is what the oracle gate (if configured) or a human reviewer answers.
  • It does not run pipeline agents — it reads their output files.