← the whole session plugin/skills/multi-tier-audit-pipeline/SKILL.md

Run a three-tier autonomous audit pipeline that discovers findings, validates them against real code, and turns confirmed issues into individually-approvable fix proposals.

Multi-Tier Autonomous Audit Pipeline

When to Use

  • When the person says "audit everything", "find all bugs", "sweep the codebase", "what's broken"
  • When preparing for a production launch and need to surface all critical issues
  • When running periodic codebase health checks
  • When you need to discover, validate, and propose fixes across your project's full surface area

Architecture

Three tiers, each with a distinct job. This escalation — don't spend expensive, careful reasoning on a finding until a cheaper pass has produced it and a middle pass has actually proven it — is the reusable mechanism here; the specific list of what to check (below) is something you build up for your own project over time.

Tier 1: CHEAP DISCOVERY (wide net, fast model, parallel)
  → One agent per source. Finds raw findings.
  → Writes its own record per source.

Tier 2: MID-TIER VALIDATION (evidence-based, reads real code)
  → One agent per critical/high finding.
  → Proves or disproves each finding by reading actual code.
  → Writes [VALIDATED] or [FALSE POSITIVE] records.

Tier 3: CAREFUL PROPOSALS (/alignment-harness:speak-human + /alignment-harness:reflect + diff)
  → One agent per validated issue.
  → Writes [FIX PROPOSAL] records for instant approval.
  → Format: who it hurts, what they see, code diff, approve/reject.

Use whatever models you have access to for the three tiers — a fast/cheap model for Tier 1, a stronger one for Tiers 2-3. The tiering is about cost-matching effort to how far a finding has already been proven, not about specific model names.

Critical Rules

  1. One agent per source — never share work between agents. Prevents file conflicts and enables full parallelism.
  2. Write to records, not shared scratch files — see "Where findings live" below. Never have multiple parallel agents write to the same flat file (a real first-run failure mode: two discovery agents overwrote the same /tmp file and lost data).
  3. Tier 2 must read actual code — never trust a Tier-1 finding without code-level verification. Fast/cheap models misread things like conditional ordering and metadata labels that look like control flags.
  4. Tier 3 proposals use /alignment-harness:speak-human — zero jargon in proposals. If a smart non-coder would furrow their brow, rewrite it.
  5. Each Tier-1 agent maintains its own record — tagged with its source for traceability.

Where findings live

This ships with no server or database dependency. Each phase writes a short markdown or JSON file to a local folder:

  • Discovery findings → alignment-harness records audit-discovery
  • Validated/false-positive findings → alignment-harness records audit-validated
  • Fix proposals → alignment-harness records audit-proposals

Each command prints the folder path — write one file per finding there (filename can be the finding's short slug), tagged in its frontmatter or JSON the same way the phases below describe. If your project already has its own findings/proposal tracker (an admin UI, a database), write there instead and keep the same phase structure.

Phase 1: Dispatch Discovery Sweepers

Dispatch one agent per source, in the background where your tooling supports it. Each agent:

  • Runs its specific checks
  • Writes a record titled [DISCOVERY] {source name} — X findings
  • Tags: ["discovery", "{source-name}"]

Default Source Registry — start here, then extend it

These are the sources that work with nothing beyond this plugin installed. Extend this list with your own project's scanners, security tooling, and critical paths (payments, sign-in, your core feature) the same way you'd extend the governer's risk table — this registry is meant to grow into your own project's version of "audit everything," not stay generic forever.

# Source Skill/Command What It Checks
1 Your own lint/test/scan scripts whatever npm test, npm run lint, or your own scanner scripts are Whatever your project already checks in CI
2 Intent DB alignment /alignment-harness:intent-alignment-audit Confirmed-intent coverage gaps
3 Failing tests alignment /alignment-harness:tests-failing-check Real bugs vs stale tests vs environment issues
4 UX testing learnings /alignment-harness:ux-testing-learnings Known losses, blockers, patterns you've logged
5 Browser smoke check /alignment-harness:agentic-chrome-testing or /alignment-harness:devtools-site-testing Real browser: pages load, dead screens, console errors
6 Validate against live data /alignment-harness:validate-against-live-data Code decisions vs real data, if you have any to check against
7 Agent telemetry patterns — Your own agent failure-streak or hook-error logs, if you keep any
8 Compaction/session hygiene /alignment-harness:compact-agentic-session Thin session summaries missing structured fields

If a source in this table isn't configured or its skill isn't installed, say so plainly in the discovery output ("source N skipped — not configured") rather than inventing a plausible-sounding finding for it.

Example: a much larger, fully-configured registry (opt-in — from a real product, illustrative of how far this grows over time, not something to copy literally): a mature setup can easily reach 40+ sources — security reviews (OWASP-style checks, static analysis for cross-file vulnerability patterns, insecure-default detection), backend quality scanners (missing return before a response, unstable useEffect dependency arrays, console.log in hot paths, production-gate logic), product-specific critical-path checks (payment webhooks, sign-in flows, the product's core feature), user-feedback intake, and admin-registry sync checks. None of these ship with this plugin — build or install the ones relevant to your own stack, and add a row here once you have.

Phase 2: Validation

After discovery agents complete, look through the discovery records for findings not yet validated.

Recurring check (if you have a way to schedule one)

/alignment-harness:loop 5m Validation: check the discovery records folder for findings not yet validated. For each critical/high finding, dispatch an agent to read actual code and prove/disprove.

Per-Finding Validation Protocol

  1. Read the ACTUAL source file at the cited line
  2. Trace the execution path — don't assume from function names
  3. Check any logs you have for production evidence of the issue
  4. Write a record:
    • [VALIDATED] {title} with evidence OR [FALSE POSITIVE] {title} with disproof
    • Tags: ["validation", "{source}", "confirmed-issue"|"false-positive"]
    • If validated, note: "Needs a fix proposal"

Validation Priority

  • Critical findings → immediate manual dispatch (don't wait for a scheduled pass)
  • High findings → next pass
  • Medium/low → batch together

Phase 3: Fix Proposals

After validation, look through the validated records for [VALIDATED] + confirmed-issue findings without a fix proposal yet.

Recurring check (if you have a way to schedule one)

/alignment-harness:loop 10m Proposals: check the validated-findings folder for issues without a fix proposal yet. Dispatch one agent per issue.

Per-Issue Protocol

  1. Invoke /alignment-harness:speak-human — translate to zero-jargon language
  2. Invoke /alignment-harness:reflect — catch assumption cascades
  3. Read actual source code — trace full UX impact
  4. Capture browser evidence if UI-visible (via a browser-testing skill you have)
  5. Write a record with this exact format:
## {FINDING-ID} {One-sentence plain-English problem}
**Who it hurts**: {user type} during {UX moment}
**What they see**: {exact broken experience}
**Screenshot evidence**: {path or "not UI-visible"}

### Proposed Fix
**File**: `path/to/file.js:line`
**Change**: {2-3 sentence human description}
**Risk**: {low/medium/high} — {why}
**Reversible**: {yes/no}

### Code Diff
\```diff
- old line
+ new line
\```

**Approve?** Yes / No / Revise

Tags: ["fix-proposal", "awaiting-approval", "{repo}"]

Phase 4: Person's Review

All proposals land in the local folder alignment-harness records audit-proposals prints (or wherever your own project's tracker lives, if you wired one in).

  • Look for records tagged awaiting-approval
  • Each proposal is a quick approve/reject decision
  • Approved → agent implements and commits
  • Rejected → dropped or revised

Dispatch Pattern

// Phase 1: Dispatch all discovery agents in parallel
const sources = [
  { name: "lint-test-scan", prompt: "..." },
  { name: "intent-alignment", prompt: "..." },
  // ... one per source from your registry
];

// Dispatched in the background where supported
// Orchestrator stays available to the person

// Phase 2: As discovery agents complete, dispatch validation for critical findings
// Phase 3: As validation confirms issues, dispatch fix proposals
// Never block — keep dispatch in the background where you can

Learnings from a Real First Run

These are the general lessons — not tied to any one project — worth carrying forward:

  1. Mid-tier validation caught real false positives — a fast/cheap discovery model misread a conditional's ordering and a debug-mode metadata label as if it were a control flag.
  2. Shared scratch files cause race conditions — multiple discovery agents overwriting the same file lost data. Use the per-finding records described above instead.
  3. Manual dispatch is faster than waiting for a scheduled pass, for criticals — push critical findings through immediately.
  4. A first sweep tends to surface a lot at once — expect dozens of findings the first time you run this against an established codebase, most low-severity.
  • /alignment-harness:orchestrator — top-level coordination protocol
  • /alignment-harness:compact-agentic-session — record/compaction creation protocol
  • /alignment-harness:reflect, /alignment-harness:speak-human — used in Phase 3
  • /alignment-harness:tests-failing-check, /alignment-harness:ux-testing-learnings, /alignment-harness:validate-against-live-data, /alignment-harness:intent-alignment-audit — default Tier-1 sources