← the whole session plugin/skills/intent-validation-engine/SKILL.md

Validate proposed UX intents against code and, if you have them, end-to-end tests — confirming a recorded decision still matches what the code actually does.

Intent Validation Engine

Anyone can write down what they wanted ("when a free user hits the limit, show an upgrade prompt"). This skill checks whether that written statement still matches what the code actually does — an intent nobody re-checks is a guess wearing the clothes of a settled decision, and code keeps changing underneath it. This is /alignment-harness:intent-leverage-audit's sibling: that one asks "of everything that might not match, which matters most" and works the priority order; this one is the plain check itself, run over whatever set of intents you give it.

When to Use

  • After an agent proposes new intents — validate them before human review
  • After code changes — re-validate intents that reference changed files
  • Pre-launch — audit all intents for accuracy
  • When asked to "validate intents" or "check intent accuracy"

Prerequisites

This reads from the Intent DB (/alignment-harness:intent-db, a local file store — intent.js next to that skill, no database or server required). If your project runs its own database-backed intent store instead, point at that; the validation logic below is the same either way.

Quick Commands

# List everything currently unvalidated (recently un-touched), highest-level view first
node <path-to-intent-db-skill>/intent.js recent 100

# Get one intent's compact brief before checking it
node <path-to-intent-db-skill>/intent.js brief <slug>

# Mark an intent as re-confirmed once you've checked it
node <path-to-intent-db-skill>/intent.js touch <slug>

There is no separate database or model to connect to — everything above operates on the JSON files the intent-db skill already manages.

How It Works

Three-Pass Validation

  1. Direct test coverage (strongest): if your project has end-to-end tests (Playwright, Cypress, or anything else), grep them for the intent's slug or a citation tag. An intent whose slug appears in a test file is validated with that test file's name recorded as evidence. If you don't have end-to-end tests yet, skip this pass and say so — it isn't a blocker for the other two.

  2. Journey batch validation: group the remaining intents into named "journeys" — a journey is just a user flow that several intents belong to (for example: onboarding, sign-in, checkout, or whatever flows your own product actually has). Checking one representative intent per journey against the code, then applying the result to the whole batch, is much faster than checking every intent alone once you have enough of them to make batching worth it. Start this list empty — there is no universal set of journeys; define your own as your intent store grows, the same way you'd define your own risk-table categories.

  3. Uncategorized: intents that don't match any journey (or when you have no journeys defined yet) get checked one at a time — read the code they reference, compare it to the UX promise, and record the result. These are also the ones worth building a journey bucket for, once enough of them accumulate around the same flow.

Status Flow

proposed → checked (matches) → approved (by a human)
proposed → checked (mismatch) → needs human review

In the local intent-db store this maps onto the existing proposed/approved status field plus the touch timestamp as your "last checked" marker. If you want a richer record than a timestamp — which test file covered it, which journey it was batched into, whether it mismatched and how — keep that as a note alongside the entry (for example, appended to the entry's own notes, or as a short file in the folder alignment-harness records validation-runs prints), since the shipped local schema doesn't carry a dedicated validation sub-object out of the box.

Recording a mismatch

When an intent doesn't match the code:

  • touch it anyway — you did check it, and a stale "never checked" state is worse than a checked-and-flagged one.
  • File a finding the same way /alignment-harness:intent-leverage-audit does: your own project's proposal-tracking system if it has one, otherwise a markdown file in alignment-harness records proposals, naming the slug, the broken promise, the affected files, and whether it's a code bug (the code stopped matching a still-correct intent) or a stale intent (the intent describes something removed or changed on purpose).

Journey Buckets

If you build these out, keep each one to:

  • matchTags — which intent tags put something in this bucket
  • matchParentPrefix — parent-slug prefixes that belong here
  • testFile — the end-to-end test file that covers this journey as a whole, if one exists

Store them however suits your project (a small JSON or JS file next to your own scripts) — there's no fixed schema this skill requires beyond "a way to group related intents and remember what covers them."

Staleness — say the true state, don't present an old check as current

Before reporting results, say when this last ran and how many intents have accumulated since. A validation from months ago, presented without that context, reads as "everything's fine" when it may mean "nothing has been checked in months while the store kept growing." If nothing has ever been checked, say that plainly rather than reporting 0% coverage as if it were a score to be alarmed about — it's just the starting state.

If you don't have any intents in your store yet, or none are approved, there's nothing to validate yet — this becomes useful once /alignment-harness:intent-lifecycle (or your own process) has recorded some confirmed intents and you have code or tests to check them against.