← the whole session plugin/skills/agentic-prelaunch-audit/SKILL.md

Run a multi-phase pre-launch audit: find every change since the last known-good state, make sure each is documented and tested, triage real bugs from noise, then give a risk-ranked go/no-go recommendation. Use when preparing any branch for production.

Pre-Launch Audit

Philosophy

A period of agent-assisted coding creates a surface area too large to test exhaustively by hand. This skill applies systematic coverage verification: every code change gets documented, every documented intent gets a testable UX statement, every UX statement gets a test, every test gets run, every failure gets triaged as a real bug or stale test. The output is a ranked list of actionables with confidence scores — not a wall of green checkmarks and not a wall of unexplained red ones.

You are the orchestrator. You dispatch sub-agents, verify their output (spot-check at least 1 in 5 by reading the files they claim to have changed), aggregate findings, and present a prioritized, evidence-backed launch decision.

Before you start: what does "safe to ship" mean for THIS product?

This audit only works once you know what to check. Before Phase 0, if you don't already have this recorded somewhere, ask (or infer from the repo and confirm):

  1. Which repos are in scope for this release? (could be one repo, could be several)
  2. What is the base branch this is measured against? (usually main)
  3. Which 2-4 flows in this product must never break? — this is the "critical path" list Phases 6 and 8 test. Look for hints first: payment/auth-looking code, existing CI test suites, deploy scripts, a CONTRIBUTING.md. Propose a draft list from what you find, and have the person confirm or correct it rather than silently guessing.

Record the answer somewhere durable (a file under alignment-harness records prelaunch-audit, or wherever this project keeps its own decisions) so the next audit doesn't have to ask again.

Example from a reference setup (opt-in — replace with your own): three repos (api, frontend, landing), critical paths = payment, login, and the core coaching flow, systemic checks include cookie-domain and CORS verification because those are the product's actual failure history.

If you have a records/compaction API set up for durable audit trails (see /alignment-harness:harness-setup), use it in the steps below wherever this file says "record it." Otherwise, write markdown files to the folder alignment-harness records prelaunch-audit prints, and say so — never invent an API call to an endpoint that doesn't exist for this project.

Prerequisites

Before starting, read whatever skills your own setup has for: session compactions/record-keeping (agentic-session-compactions if shipped), checking intent alignment (intent-alignment-audit), fixing failing tests, and triaging error logs. If a named skill below isn't installed, do the step's underlying job with plain reasoning instead of skipping it — say so ("no <skill> installed — doing this with direct code reading and test runs instead").

Run git worktree list in each in-scope repo and clean up prior stale worktrees first.


Phase 0: Deviation Discovery

Goal: Find every code change between the current branch and the base branch, across every repo in scope.

cd <repo>
git log <base-branch>..HEAD --oneline --no-merges

If the repo IS on the base branch, check for uncommitted/staged/untracked work instead:

git diff --name-only
git diff --staged --name-only
git status --porcelain | grep '??'

Group commits into logical clusters (by commit message, timestamp, and file grouping) — each cluster is roughly one past session's worth of work.

For each cluster, check if a record already exists (a compaction, a written summary, anything durable). If you have agentic-session-compactions or similar set up, search it; otherwise git log -p the cluster and grep your own past session transcripts (see the agentic-find skill's local fallback) for a matching description.

For any cluster with no record, dispatch a sub-agent per cluster (parallel):

You are documenting work that was done but never recorded, before a pre-launch audit.

Read every changed file and every commit diff in this cluster.
Write a short record covering:
- what changed and why (your best reconstruction from the diff)
- uxStoriesBroughtToReality: statements in the form "When {user does X under Y constraints}, they {experience Z}" —
  these are the testable UX statements Phase 2 will verify.

If a records/compaction system is set up, write it there. Otherwise write a markdown file to the
folder `alignment-harness records prelaunch-audit` prints, named for this cluster.

Gate: Do not proceed to Phase 1 until every cluster has a record.


Phase 1: Alignment Audit

Goal: Verify every record's stated UX intent actually matches what the project has decided it wants (its intent database, past decisions, or stated priorities — whatever that means for this project).

For each record, dispatch a sub-agent (batches of 5-10 records per agent):

You are auditing whether recent work matches previously stated intent.

If an intent-tracking system exists (see the `intent-db` skill), invoke it and search for each
UX story below. Otherwise, search institutional memory (the `agentic-find` skill — set up or its
local grep fallback) for anything that already decided this should work this way.

For each UX story:
1. If a matching prior intent/decision exists: note "ALIGNED — matches <what you found>"
2. If it contradicts a prior decision: note "MISALIGNED — <why>" and flag for review
3. If it's new (no prior decision either way): note "NEW — no conflicting record found" and,
   if you have an intent-tracking system, propose it there

UX stories to audit:
<list of records, titles, and UX stories>

Output per record:
- id: <record id>
- alignment: ALIGNED | MISALIGNED | NEW
- reasoning: <why>
- flagForReview: true/false (true if MISALIGNED or uncertain)

Collect flagged items and turn them into actionable tasks (see Phase 7).


Phase 2: UX Statement Coverage

Goal: Every record has complete, testable UX statements covering every change it describes.

For each record, verify:

  1. Every changed file has at least one UX statement describing its user-facing impact.
  2. Statements follow the form: "When {constraints in UX terms}, we {testable outcome}."
  3. Non-UX changes (refactors, internal tooling) still get a statement in the same shape from the agent's point of view: "When an agent searches for X, it finds it because Y."

For records with gaps, dispatch a sub-agent per record:

You are adding missing UX statements to a work record.

Read the changed files listed below. For each user-facing change, write:
"When {user type} {does action} {under constraints}, they {experience outcome}"

Update the record with the new statements (in your records system if set up, otherwise edit the
markdown file directly) and note what you added.

Gate: Do not proceed until every record has UX statements covering every changed file.


Phase 3: TDD Coverage

Goal: Every UX statement has a test that verifies it. Every edge case is covered.

For each UX statement:

  1. Search for existing coverage: grep -r "<keywords from the statement>" --include="*.test.*" --include="*.spec.*" <repo>
  2. Classify: COVERED (a test directly verifies this outcome), PARTIAL (a test exists but misses edge cases), or UNCOVERED (no test).

For UNCOVERED and PARTIAL statements, dispatch sub-agents (one per test file, in parallel):

You are writing tests for UX statements that lack coverage.

If a test-driven-development skill is installed, invoke it first. Otherwise: write the test
before assuming the implementation is right.

For each statement:
1. Read the implementation.
2. Write a test that verifies the EXACT UX outcome described, plus edge cases: null inputs,
   error states, boundary conditions.
3. Give the test a name that states the intent, e.g. "When {condition} then {outcome}" — not
   "test 1" or "handles edge case."
4. Run it.
5. If it FAILS because the code is wrong: do NOT fix the test to match broken code. Flag it as a
   real bug: "BUG FOUND: '<statement>' is not implemented correctly — test expects X, code does Y."
6. If it PASSES: confirm coverage.

Output: list of { uxStatement, testFile, testName, status: PASSING | BUG_FOUND, bugDescription? }

Any BUG_FOUND result is a real bug discovered through the act of writing the test — add it to the actionables list with high priority (Phase 7).


Phase 4: Full Test Suite Health

Goal: Run everything. Separate real bugs from test rot.

cd <repo>
NODE_ENV=test npx jest --no-coverage --no-watchman --forceExit 2>&1 | tail -30

(substitute your project's actual test runner if it isn't Jest)

For each failure, triage with plain reasoning if you don't have a dedicated test-fixing skill installed: read the failure, check whether it reflects the intended UX (from Phase 0/1's records or the code itself), and classify as FIX_CODE, UPDATE_TEST, or REMOVE_TEST — never silently mock away a failure to make it green. Dispatch one sub-agent per failing file, max 5 concurrent:

You are triaging a failing test for pre-launch readiness.

Test file: <path>
Failure output: <error output>

1. Establish the intended UX (check records/intent system, then the code itself).
2. Classify: FIX_CODE | UPDATE_TEST | REMOVE_TEST (with a documented reason for removal).
3. Implement the fix.
4. Re-run the test and confirm it passes.

Output: disposition, uxImpact (if this was a real bug), confidence 1-100, files changed.

Phase 5: Error Log Triage

Goal: Surface concerning errors from wherever this project's runtime errors are logged, and trace each to its cause.

If you have a structured error-logging system (a database, an admin dashboard, a log aggregator), query the highest-severity, unresolved, non-dismissed entries — recent first, most severe first, a reasonable cap (e.g. top 30). If you don't have one, grep your application's actual log output or crash reports for the same window this audit covers.

For each error found, dispatch a sub-agent (or do it yourself for a small list):

1. Check whether this exact error is already known and accepted (a "known issues" list, if you keep one).
2. If not: read the source at the failure location.
3. Classify: REAL_BUG | ALREADY_FIXED | NON_ISSUE.
4. For REAL_BUG: propose a fix in the form WHAT HAPPENED / INTENDED UX / PROPOSED FIX.

Output per error: severity, classification, uxImpact, proposedFix, confidence 1-100.

If no error-logging system exists at all, say so plainly and skip this phase rather than fabricating findings.


Phase 6: Systemic Checks

Goal: Run whatever validators and structural checks this project actually has, plus the handful that matter for almost any web product.

Generic checks worth running on most projects, if applicable:

  • npm audit --production (or your package manager's equivalent) for known-vulnerable dependencies.
  • Every environment variable the app requires at runtime is actually set (build the list from .env.example or the config-loading code, not from memory).
  • If you have admin or privileged routes: verify each one sits behind auth middleware (admin-security-check skill if installed, otherwise grep every admin route file for the auth guard your framework uses and confirm it's actually applied, not just imported).
  • If the frontend and backend are separate deployables: verify the frontend's expected response shape matches what the backend actually returns for each endpoint touched in this release (api-alignment-check skill if installed).
  • If you run feature flags or experiments: confirm nothing on a critical path (see below) is at 100% allocation with no way to roll it back, and nothing is targeted at "all users" when it was meant to be a partial rollout.

Example from a reference setup (opt-in — replace with your own): verifying CORS allowlist and cookie domain configuration for a multi-subdomain product, checking a specific admin-route-guard component pattern, running product-specific diagnostic endpoints, verifying a payment webhook responds 400 (not 404) to an unsigned request, checking deploy tags exist per repo before shipping.

If none of this project's own systemic checks are defined yet, this is where you write them down for next time, using whatever you found in Phase 0-2 about what actually broke before.


Phase 7: Risk-Ranked Actionables

Goal: Aggregate every finding from Phases 1-6 into one prioritized list with confidence scoring. This logic is general — it applies to any project.

For every actionable item, score four dimensions (1-100):

Dimension Question
Root Confidence How confident are you that you found the actual root cause?
Intent Confidence How confident are you in the INTENDED behavior here?
Risk If Wrong How much damage if your proposed fix is incorrect? (100 = catastrophic)
UX Impact If Unfixed What's the actual user impact if this ships unfixed? (100 = app unusable)
actionScore = (rootConfidence + intentConfidence) / 2
riskScore = (riskIfWrong + uxImpactIfUnfixed) / 2

if actionScore >= 60:
    if riskScore >= 70:  action = ACT + FLAG_FOR_REVIEW   # confident, but stakes are high
    else:                action = ACT                      # confident and safe
else:
    if riskScore >= 50:  action = ESCALATE_TO_ADMIN         # not confident enough, stakes too high
    else:                action = ACT_WITH_CAUTION          # low confidence, low risk

Tag each finding with which critical path it touches (from your "before you start" list) and how bad it is if that path breaks — this is what lets a payment-flow bug outrank a cosmetic one at the same confidence score.

Confidence calibration:

  • 90-100: you read the code, ran the test, saw it pass/fail, and know the intent from a real record. Act.
  • 70-89: you read the code and the logic makes sense, but you couldn't run it, or the intent is deduced rather than confirmed. Act with a review flag.
  • 50-69: you read the code but intent is ambiguous or the system is complex. Escalate if risk > 50.
  • Below 50: you're guessing. Always escalate. Never act.

Execute ACT items: dispatch a sub-agent to read the real code, make the minimal fix, run the relevant tests, and report what changed. If the fix introduces new failures, revert and escalate instead of pushing forward.

ACT + FLAG_FOR_REVIEW items: implement in a clearly labeled commit, log it as a task for review, include the full scoring rationale.

ESCALATE_TO_ADMIN items: log as a task with the finding, your reasoning, your uncertainty, and the UX impact. Do not attempt the fix yourself.


Phase 8: Critical Flow Sanity Testing

Goal: Actually exercise the 2-4 critical flows you named before starting — hit the real endpoint or the real UI, don't just read the code and assume it works.

For each critical flow, write (once, then reuse) a short script of real checks: what request or action to make, and what response or behavior means PASS. The point is testing against reality, not against your reading of the code (see /validate-load-bearing-claims-against-reality if you have it).

Example from a reference setup (opt-in — replace with your own critical flows):

Login flow:
1. curl the "who am I" endpoint unauthenticated → expect 401, not a connection error (proves the server is up)
2. POST valid test credentials to the login endpoint → expect a token or a clear error, never a 500
3. Use the token against an authenticated endpoint → expect the right user data back

Payment flow:
1. Create a checkout session with a test plan → expect a session id/URL, not a 500
2. POST an unsigned request to the payment webhook endpoint → expect 400 (rejected signature),
   never 404 (route not registered — that's a P0 failure)
3. Read the code path that decides what a user gets on a payment-provider error: does it fail
   toward giving them what they're already paying for, or toward locking them out? A paying
   user should never lose access because of a transient error on your side — verify every
   catch block on this path defaults toward access, not away from it.

This "fail toward access, not away from it" check is worth doing on ANY project with paid tiers, regardless of what payment provider you use — it's one of the more common ways an unrelated bug turns into a support incident.

For each flow, report PASS/FAIL per step with the actual response, not just "looks fine."


Phase 9: Final Evidence & Handoff

Goal: Produce one durable record of the whole audit and hand it to the person for a go/no-go decision.

If a records/compaction system is set up, write the audit there. Otherwise, write it as a single markdown file to alignment-harness records prelaunch-audit, containing:

## Pre-Launch Audit: <branch> -> <base branch>
Date: <date>  ·  Repos: <list>

### Coverage
- Deviations from base: <N commits across M repos>
- Records: <N total, M created during this audit>
- UX statements: <N total, M added during this audit>
- Tests: <N covering UX statements, M new tests written>

### Critical Path Status
| Path | Status | Evidence |
|------|--------|----------|
| <flow 1> | PASS/FAIL | <what was actually tested> |
| <flow 2> | PASS/FAIL | <what was actually tested> |

### Findings Summary
- Bugs found and fixed: <N>
- Bugs found and escalated: <N>
- Test failures resolved: <N>
- Systemic/log issues: <N> (M fixed, K escalated)

### Risk-Ranked Actionables (Remaining)
| # | Finding | Action | Root Conf | Intent Conf | Risk | UX Impact | Path |
|---|---------|--------|-----------|-------------|------|-----------|------|

### Recommendation
GO / NO-GO / CONDITIONAL-GO, with specific conditions if conditional.

### What the person should manually verify before shipping
1. <specific thing to check by hand>
2. <specific thing to check by hand>

If you have a task-tracking system, file the summary there too, flagged for review. Otherwise this markdown file IS the handoff — tell the person exactly where it is.

Deployment Gate Checklist

A short, general checklist that applies to almost any project:

- [ ] All defined critical paths tested and PASS in Phase 8
- [ ] Full test suite: no NEW failures introduced by this branch
- [ ] No unresolved high-severity errors from Phase 5 (define "high" for your own project)
- [ ] A rollback point exists (a tag, a known-good deploy, a stable branch)
- [ ] Every required environment variable is set in the target environment

Example from a reference setup (opt-in — replace with your own): cookie domain configured correctly for a multi-subdomain product, CORS allowlist matches production domains, payment webhook responds 400 not 404, no experiment at 100% allocation on a critical path, one payment flow and one login flow manually clicked through in a browser before shipping.

Recommendation: GO only if all boxes for YOUR project are checked.


Orchestration Principles

  1. Always pass the relevant skill to sub-agents — they start cold. If a skill informed your understanding, say "invoke <skill-name> first" if it's installed, or hand them the relevant reasoning directly if it isn't.
  2. Parallel when independent — missing records (Phase 0), alignment checks (Phase 1), test writing (Phase 3), test triage (Phase 4), and each critical-flow check (Phase 8) can all dispatch in parallel.
  3. Sequential when dependent — Phase 1 needs Phase 0 complete. Phase 3 needs Phase 2. Phase 7 needs Phases 1-6.
  4. Max concurrency — around 5 parallel sub-agents, to avoid overwhelming the system you're auditing.
  5. Verify sub-agent output — never trust blindly. Spot-check at least 1 in 5 by reading the files they claim to have changed.

When to stop

Continue until: every defined critical path shows PASS in Phase 8, every ACT item is implemented and verified, every ESCALATE item is logged, the full test suite has no NEW failures, and error logs show nothing above your own severity threshold unresolved. Only then state your launch recommendation with evidence.

Anti-Patterns (never do these)

  • Skip reading code and propose fixes based on error messages alone.
  • Bulk-fix failing tests by mocking away the failures.
  • Dismiss error-log entries without tracing the root cause.
  • Claim "ready to ship" without actually running Phase 8's real checks.
  • Act on low-confidence findings that touch a critical path.
  • Write UX statements in code terms instead of what a person actually experiences.
  • Skip Phase 0 and assume every past session already has a record.
  • Run phases out of order — the dependency chain exists for a reason.