← the whole session plugin/skills/sanity-check/SKILL.md
Confirms that work matches intent — before starting, after completing, before committing, or when reviewing a plan or diff.
Sanity Check
Before your output, print ## SANITY_CHECK on its own line. This exact heading is what lets an optional end-of-session capture hook (see the bottom of this skill) find and harvest this exchange later, if you've set one up. Nothing reads it automatically otherwise — the check's real-time value (catching a misread before it compounds) doesn't depend on that capture existing.
Load the person's voice first, if you have it calibrated
Before producing sanity-check output, invoke /how-to-talk-like-the-founder (or your own equivalent voice-calibration skill, if you've built one) and read whatever voice samples and rules it returns. The sanity-check output is read by the person at accept/reject moments — the moments where voice friction is most costly. Output that sounds like an assistant narrating ABOUT the work creates friction. Output that sounds like the person talking about their own work removes friction.
After loading the voice skill, produce the numbered UX statements and surrounding prose in that voice register — short, direct, named, no hedging adverbs, no summary openers ("Here's what I built…"), no "this section will" constructions. Statements should read like the person would say them if they were narrating their own work back to themselves. The numbered list format itself is preserved; only the voice register changes.
If the voice skill isn't set up yet, or has no calibration samples for this person: don't halt. Run the check anyway, using the plain-communication rules below (short, direct, no jargon, one claim per line), and say once, plainly, that the voice isn't tuned to this person yet — "these are plain-language statements; I haven't calibrated your voice yet, want to?" A sanity check in a slightly generic register still catches a misread; refusing to run one at all because tuning is pending would be the worse failure.
MANDATORY — Statement scope and self-containment
For each statement, if it is only true within a specific scope, ANY variable you use, you MUST include the scope referred to specifically.
For each statement, if the reader will need to ask questions to understand what it means or applies to or which UX is intended or which actual semantically framed outcome is intended from the statement, it is incomplete. That additional context must be included precisely — but synthesized with the statement so it reads in a format that will be instantly understood by the reader even if looking at this context with none of the supporting context.
A statement that triggers a clarifying question in the reader's head has failed its own contract. The completion is not an additional sentence below the statement; it is woven into the statement itself.
Example — WRONG
Any future variant of the email machine that someone builds and wires through queueEngagementEmailSend() inherits the safety reroute automatically. Nobody has to remember to gate it — the gate is inside the function they call.
This reads true but leaves the reader asking: which email machines? Variants of what specifically? What does "the safety reroute" do? "Automatically" in what conditions? The undefined scope is a hidden trapdoor.
Example — RIGHT
Any future variant of any of the automated email machine systems which create re-engagement emails for users without human oversight, that someone might build and wire through queueEngagementEmailSend() or similar functions, must by default inherit the safety reroute automatically. Nobody has to remember to gate it — the gate is inside the function they call so that there is always the option of observability before go-live without any risk of that breaking.
This stands alone. The reader doesn't need to ask "which" anything — every variable is named in-line. The scope is explicit. The behavior is explicit. The intent is explicit.
How to apply this rule
After drafting each statement, read it as if you had zero context about this session. If you would ask any clarifying question, the statement is incomplete. Revise. Repeat until every statement passes the cold-read test.
You help the human confirm that work product matches intent — before, during, or after execution.
When This Fires
| Situation | What you do |
|---|---|
| Task start | Confirm you understand the assignment before writing code |
| Task end | Confirm the work you did matches the intended UX outcome |
| Before commit | Review all uncommitted changes and surface what they mean in UX terms |
| Commit batch review | Review a large set of commits and summarize their UX impact |
| Plan or diff review | Read a plan or diff and confirm it makes sense |
How It Works
1. Gather the work product
Depending on the situation:
- Task start: Restate the assignment scope
- Task end: Summarize what you built/changed
- Uncommitted files: Run
git diffandgit statusto see everything that changed - Commit batch: Run
git logover the range - Plan/diff: Read the document or diff provided
2. Translate to UX intent statements
For ALL the work product, produce a numbered list of UX statements — plain human language, no jargon. Each statement describes what will be true for the user as a result of this work.
Use /speak-human to translate. Every statement should be:
- Conditional: "When [user condition], [what happens]"
- Testable: Someone could verify it by using the app
- Specific: No vague claims like "improves performance"
Example output for uncommitted file review:
If all changes are intentional and correct, these UX statements will be true:
1. When a new user signs up, they see the onboarding wizard before the dashboard
2. When a paying user opens settings, the billing section shows their current plan name
3. When any user clicks "Export", the CSV downloads within 3 seconds
4. When an admin simulates a free user, the upgrade modal appears on the gated feature
3. Human confirms
Present the numbered list. The human reviews and confirms, edits, or rejects items.
4. Capture the confirmed statements
Where the confirmed UX statements go depends on the situation:
| Situation | Where statements are saved |
|---|---|
| Before commit | Include in the git commit message body |
| Plan review | Include in the plan document |
| Every situation | Nothing else is automatic unless you've set up the harness's optional end-of-session capture: a hook that scans the transcript for this exact heading and copies the block you print, plus the person's reply, verbatim into a harvest folder, later mined into per-domain "instinct ledgers" by /instinct-harvest if you have that skill. If you haven't set this up, the check still does its real-time job — catching a misread before it compounds — it just doesn't compound into a ledger yet. That is why the ## SANITY_CHECK heading must be exact and why one assumption per numbered line matters: the person's "3 is false" is what gets harvested. |
5. Scope confirmation (task start)
At task start, your sanity check is a scope declaration:
"My understanding: you want me to [X] so that [Y]."
UX statements that will be true when done:
- When [condition], [outcome]
- When [condition], [outcome]
Wait for confirmation before proceeding.
5b. Instincts I'm weighing (task start, additive 2026-09-08)
If the domain you're working in has an instinct ledger — a per-domain notes file recording past corrections, one working setup keeps these under docs/intent/<domain>/INSTINCTS/README.md — read the instincts for the situation you are about to enter and include, after the UX statements:
Instincts I'm weighing: [instinct headline] — applies because […]; [instinct headline] — I think this does not apply here because […].
If no ledger exists for this domain yet, skip this step; it costs nothing to skip and there's nothing to weigh yet. The person sees whether the instincts fired right. Their reply is harvested like any other sanity-check reply, so this is how the ledger compounds without turning instincts into rules.
6. Completion confirmation (task end)
At task end, your sanity check verifies intent:
"Here's what I built and the UX it delivers:"
- When [condition], [outcome] — verified / needs verification
- When [condition], [outcome] — verified / needs verification
Mark each statement as verified (you tested it or read the code path) or needs verification (human should check).
UX Truth Lens
When sanity-checking any work product, list all the statements that will be true in UX semantic meaning — not variables — if this change ships. Right, nuanced, one nuanced context-bound sentence each.
MANDATORY — Surface the DECISIONS, translated to semantic meaning (added 2026-06-07, additive)
This is additive to everything above, not a replacement. The UX Truth Lens above covers what the code does. This section covers the decisions that produced it — and those are a distinct thing the person needs and that agents routinely drop.
A diff is not just behavior; it is a record of choices. Every guard, threshold, ordering, default, and fallback is a fork where someone chose option A over option B. The behavior is the result of the decision. If you report only the behavior, you have left the decision implicit in the code — which is exactly the "using code as a substitution for effective communication" failure. The person cannot evaluate whether they agree with a decision they can't see.
The rule: for each non-trivial change in a batch, you must surface the decision as a decision and translate it to its semantic meaning — what was chosen, what the alternatives were, and why this one, in terms a person understands. Never leave the decision sitting inside the code for the reader to reverse-engineer.
The distinction, concretely
- Behavior (necessary but not sufficient): "When the kill-switch holds any value other than
false,"false","0", or0, the email routes to admin instead of the user." - Decision, translated (what was missing): "We decided the only thing that counts as 'safe to send to a real person' is an explicit, unambiguous off — and that every other possible value, including a half-set or garbled one, resolves toward holding the email, never sending it. The alternative was to treat 'not exactly on' as 'send,' which would mean a typo or a schema cast could leak a live email to a customer. We chose the direction where the failure is always 'too cautious,' because the cost of the other direction is a real person getting an email we never reviewed."
The second one is the decision. It names the fork, the rejected alternative, and the value judgment that settled it — all in human terms. That is what must not be left out.
How to do it
For each batch, after the behavior statements, walk the decisions the diff embodies. A decision is present wherever the code could reasonably have done otherwise:
- A chosen default or fallback (what happens when nothing is set / something fails) — and why that direction.
- A chosen threshold or set of accepted values (why exactly these, why not looser/tighter).
- A chosen ordering (why this step before that one, and what the ordering costs or buys).
- A chosen scope (why this set of cases is handled and that set is deliberately not).
- A chosen tradeoff (what was accepted as a cost in exchange for what benefit).
For each, write one synthesized statement: what was chosen, what it was chosen over, and why — in semantic meaning, with the code as the anchor, never as the substitute. Same voice rules as the rest of the skill: intent leads, code rides inside as the precision anchor.
A decision you report as only its outcome has failed this section. The reader should finish a batch knowing not just what it does, but every judgment call that went into it and whether he'd have made the same call.
Rules
- Never skip the numbered list. The list IS the sanity check. Without it you're just narrating.
- Each line reconstructs the PERSON'S INTENT — never validates the agent's steps. A line like "your approval covered X" or "I was authorized to Y" is the wrong register: it asks the person to confirm the agent's permission trail, which they cannot validate and which tells them nothing about whether the agent understood WHY. The correct register: "if this work was aligned, your intent must precisely have been: when [reader/user condition], [intended outcome] so that [why it matters]" — a statement the person can instantly recognize as their own intent or reject. This step is not for determining whether the steps taken were correct — it's for determining whether the agent's read on the person's intent actually matches that intent. The teaching example (Attempt 3, below) was always in this register; action-validation lines are the drift to guard against.
- ONE assumption per numbered line — never bundle. A numbered item containing two or more separately-falsifiable claims cannot be validated by number: the person can't say "3 is false" when 3 holds four assumptions. If a drafted item joins two falsifiable claims (watch for "and," "so," "plus," "which means"), split it into separate numbered lines. The whole purpose of this list is to articulate assumptions clearly enough that the person can identify exactly which ones are false, so bulky paragraphs bundling multiple assumptions defeat the point. Paragraphs that allude to assumptions while narrating work are the failure mode; the atomic list is the contract.
- Never reference an internal label without carrying its meaning. "Issue 6," "the source document," "the gate" mean nothing to a cold reader. If a line needs a label, the line also states what the label refers to in plain terms, inline — a label with no attached meaning tells the reader nothing about what actually happened.
- Never use technical jargon in the statements. Translate everything through
/speak-human. - Never assume correctness. If a diff looks odd or a file seems unrelated, flag it — "This file changed but I'm not sure why: [filename]".
- Always wait for human confirmation before acting on the list (committing, proceeding, etc.).
- If reviewing uncommitted files, check for accidental changes — files that shouldn't be in the diff, debug code left in, env files staged.
Teaching example — code-rephrase vs intent contract
Example from a real project (opt-in reference — product specifics genericized, method-teaching content kept close to verbatim)
This section captures a live correction cycle from a real sanity-check session on a real product (2026-05-11 / 2026-05-12, three-repo staging→main review) — a legacy-trial eligibility bug, with the product's own name and exact figures replaced by generic equivalents below; the correction dialogue itself, which teaches the general method, is kept close to how it actually happened. It walks through three agent attempts at translating one assertion, the feedback after each attempt, the comment about how to apply that feedback, and ends with the corrected version that was accepted. Every agent invoking /sanity-check should read this before producing output and aspire to land on the third version on the first try.
Attempt 1 — WRONG (code-rephrase as intent)
- The admin intends that
trial.used = truemeans the trial has been STARTED, not "the trial has been consumed/expired." (Restores the semantics from a prior period when trial conversion was much healthier.)
Why this was wrong
This is not a correct sanity check. All of the context is required, but the actual programmatic things referenced here need to translate to an intended UX statement.
A UX statement is a set of conditions that, if true, would result in an intended UX outcome, as managed by a specific manager, function, or controller — in the shape:
When X then Y should Z so that A
- X is the actual exact conditions (not code — actual human-step conditions)
- Y is the actual system in charge of this, the manager of it
- Z is the actual programmatic switch, control, or function name, in parens, outside of the semantic meaning or translation of that
- so that A is the actual intended UX experience in the context of that "when" statement
When completely translated, it still has all the details Attempt 1 has, but as reference — it primarily documents the intent, and the programmatic things are left as the connection between the code and the intended UX and conditions.
How to apply this feedback to arrive at the next attempt
The intent must lead the statement. The code symbol (trial.used) must NOT be the subject of the sentence — it should appear only as a precision anchor inside parens after the intent has already been stated in plain human conditions. The structure to use: When [exact human-step conditions] then [the actual system/manager in charge of this] should [the actual programmatic control or function name in parens, after the semantic meaning] so that [the intended UX experience]. The translated statement must still carry every detail from Attempt 1 — no detail dropped — but the primary thing it documents is the intent, not the code.
Attempt 2 — also wrong (invented placeholders, missed the conversational shape)
- When a brand-new user in a specific onboarding cohort has begun their free trial within the last 7 days, the app's access-tier system (the manager that decides what each user can see and do, controlled by
isUserOnLegacyTrial()inbackend/resolvers/legacy_trial/isUserOnLegacyTrial.jsandfrontend/src/utilities/isUserOnTrial.js) should treat them as actively on their legacy free trial (the function returnstrueonly when thetrial[]entry's.usedfield istrueAND today's date is inside the 7-day window from.startDate), so that they get full access for the duration of that 7-day window — which restores the conversion behavior from a prior period, before an inverted predicate quietly cut it down.
Why this was wrong
Closer to this shape — using all caps for variable placeholders:
When any user that does not get bucketed into a LIMITED_ACCESS or FULL_ACCESS state (not paying or on a cc trial) is accessing the app, we are going to check if they are qualified for a legacy 7 day trial.
The intent is that they would be given one unless they are disqualified, which would be determined by them actually having already had one. So effectively, the intent is: if you have not had a trial, and you're not already getting access by some other means, then we give you a legacy trial.
We use VARIABLE set to VALUE to measure and check this in the SYSTEM/FUNCTION NAME.
There are a couple of conditions to this that are intended. For example, this was designed for only brand new users, but it's also the fall back bucket for anyone who does not get caught by another bucket. Now this will effectively label them as VALUE in the SYSTEM IN CHARGE OF THIS —
and if they don't qualify, the intent is that they land in the next bucket, called FREE_ACCESS, so that there is some bucket defined for everyone.
How to apply this feedback to arrive at the next attempt
Two distinct corrections:
First — the shape is conversational, not declarative. A sanity-check intent statement is a walkthrough of the bucketing/qualification logic as a whole, not a single dense sentence. It names the upstream buckets (LIMITED_ACCESS, FULL_ACCESS), names the qualification check, names the exceptions and fallbacks, and names what happens to users who don't qualify (FREE_ACCESS). The reader follows the logic top-to-bottom in plain conversation. Variable names and function references appear inline where they connect, not as a parenthetical aside about implementation.
Second — when the person writes ALL CAPS in a template like "we use VARIABLE set to VALUE in SYSTEM/FUNCTION NAME," those are placeholders the agent must resolve to the actual real values from the code before writing the final output. The agent must NOT keep the ALL CAPS in the output as some pseudo-variable like THE_USER'S_HISTORY_OF_TRIALS. The agent looks up the actual field name (.used), the actual value (true), the actual function name (isUserOnLegacyTrial()), and writes those real strings into the prose. ALL CAPS survives in the final output ONLY when it refers to a literal constant name in the code (like LIMITED_ACCESS, FULL_ACCESS, FREE_ACCESS — which are real tier-name constants in accessTiers.js).
Attempt 3 — CORRECT (accepted as correct)
When any user who does not get bucketed into LIMITED_ACCESS or FULL_ACCESS (so, not a paying member, and not on a credit-card-on-file trial) opens the app, we are going to check whether they qualify for a legacy 7-day free trial.
Our intent is that they would get one — unless they are disqualified by already having had one. So effectively, the intent is: if you have not had a trial, and you are not already getting access by some other means, then we give you a legacy trial.
We use
.usedset totrueto measure and check this insideisUserOnLegacyTrial(), the single resolver that answers this question for both backend and frontend (lives atbackend/resolvers/legacy_trial/isUserOnLegacyTrial.js, mirrored infrontend/src/utilities/isUserOnTrial.js).There are a couple of conditions to this that are intended. For example, we designed this for a specific new-user cohort, but it's also the fallback bucket for anyone who does not get caught by another tier. Now this will effectively label them as
isOnLegacyTrial: truein the access-tier system (the manager in charge of deciding what each user can see and do, controlled inconfig/accessTiers.json the backend, surfaced throughuseSimulatedUserStateon the frontend).And if they don't qualify, our intent is that they land in the next bucket, called FREE_ACCESS, so that there is some bucket defined for everyone.
The reason this assertion matters right now: today on
main, the predicate is inverted.if (trial.used) continue;skips trials where.usedistrue, which is exactly backwards —.used = truemeans the trial has been STARTED, not that it's been exhausted. So every brand new user in that cohort who actually starts a trial instantly drops out of the legacy-trial bucket the moment they start it. This staging branch flips the predicate toif (!entry?.used) continue;which is the correct meaning. Promoting this restores the conversion behavior from a prior period, before this bug quietly regressed it.
Instructions for agents invoking this skill
When you are about to use /sanity-check, you should review the three attempts above and understand what perfect looks like and why, and aspire to do this perfectly based on Attempt 3 as your model.
Notice that in Attempt 3, the things the user would be confused about when trying to read the sanity check are predicted and unpacked so that the actual logic of the admin that must be true if the code is correct is clearly legible in a way the user would actually read. The walkthrough names the upstream buckets the user is NOT in, the qualification logic, the exception (brand-new users vs fallback bucket), the downstream bucket (FREE_ACCESS), AND the historical context (the bug on main, the inverted predicate, the conversion regression) — all in conversation, all with real values resolved, all with code anchors only where they earn their place.
A reader who has never seen this code should be able to read Attempt 3 cold, understand exactly what the admin must intend, and be able to grep the codebase for the named functions/fields/files to verify. That is the target every sanity-check statement aspires to.