← the whole session plugin/skills/show-proof-to-human/SKILL.md
The final 30-second sanity check before going live. Produces a tiny guide the person you're working with walks through to confirm the work actually does what it was supposed to do. Use after ALL agent testing is complete — that person is NOT the tester, they are the final stamp of approval. Also use when they say 'prove it', 'show me', 'is this actually working'.
Show Proof to Human
Why This Skill Exists
You can have 1,000 new files of code that intend to make it so that this thing you did actually happens for the user. But the sanity check is the actual user going to the actual pages that have been affected and seeing if they load and seeing if anything breaks when they click on things.
This skill is the guide for that final step. It's very simple. It should take 30 seconds to complete. It doesn't involve any code at all.
It's about confirming whether the intent was realized — you're reminding the human what the UX-level intent was, including under what conditions it's true, and then you're giving them a link that produces those exact conditions.
So if this change is only for a logged-out user, and it's only at one specific URL, and it's only when they have a specific account state — you're producing a link that puts the person in that exact context. Without hallucination.
This is not simple on the agent's end, because you actually have to derive and deduce what the intent of the work was. This is why the alignment work at the beginning matters — because ideally you clarified the intent before coding, so you know what to test. And ideally you've tested all this programmatically based on the severity (determined by the governer score). So then you're giving the person the steps to quickly click around and see that it actually does the thing. See that it's fixed.
You're taking something that would take five minutes and performing a lot of work on the backend so that it can happen in 30 seconds.
The person you're working with is NOT your testing person. They are the final stamp of approval that says this is safe to go live.
How to Consume
When you invoke this skill, you MUST have already:
Validated the work yourself. Run the tests. Hit the endpoints. Open the browser. You are the tester. You checked your own work before telling the person to go check it. If you haven't done this, you are not ready to invoke this skill.
Understood the original intent. What was the human trying to accomplish? What UX outcome were they after? Under what conditions does it apply? If you can't state this clearly, go back to the scope/intent docs.
Prepared the exact conditions. If the person needs to be logged in, be on a specific page, have a specific user state — you've figured out the URL or simulator params that produce those conditions.
Only THEN do you produce the proof guide.
The Output
Print SHOW_PROOF_TO_HUMAN mode activated on its own line first. Then:
SHOW_PROOF_TO_HUMAN mode activated
## What we changed (in human terms)
{1-2 sentences — what a person using the product will experience differently. No code, no file names, no variable names. Just: "When you create a new AI coach and upload an avatar, the avatar now actually saves and displays instead of silently disappearing."}
## The conditions where this is true
{Under what circumstances does this change apply? Logged in? Specific page? Specific user state? Be precise — "when a coach creates a new custom AI and uploads a profile image during creation" not "avatar uploads."}
## What could this have broken? (agent reasoning — write this down)
{The agent MUST write down its reasoning here. Think through: what else touches the same code paths I changed? What adjacent features share the same functions, data models, or UI components? What's the worst thing that could break?
Example: "I changed the upload utility — three controllers use it: bug reports, custom AI avatars, and UX assignments. I also changed how the frontend merges form data after creation. The explore page displays avatars using getProfilePhotoUrl() which reads linodeURL — if that function broke, every coach avatar on the explore page would disappear. Payments are not affected — nothing I touched is in the payment path."
Then for each risk: include a quick check in the "Go see it" section below. The reasoning drives the checks — don't check things you can't justify as being at risk.}
## Go see it (30 seconds)
1. Open: {exact clickable URL that puts you in the right conditions}
See: {what you should see that confirms it works — specific and visual}
2. Open: {next URL if needed}
See: {what confirms it}
{That's it. 2-3 links max. If it takes more than 30 seconds, you're doing it wrong.}
Rules
30 seconds or less. If your proof guide has more than 3 steps, you're making the person do your job. Compress.
Links must produce the exact conditions. If your project has an admin user-simulator or dev auto-login route (for example:
/admin2/user-simulator?subscription=<plan-state>&redirect=/some-page, and a/dev/auto-loginroute for auth), use its URL params to encode the exact state — subscription tier, page, account condition — rather than describing steps to get there manually. If you don't have anything like that, give the plainest real path to the same state (e.g. "log in as a test user, then visit /some-page"). The person should click ONE link and be in the right context immediately whenever that's possible.Zero code, zero jargon, zero variable names. This has nothing to do with programmatic tests. It has nothing to do with code. It's about the person clicking a link and seeing the thing work with their eyes.
Remind the intent first. Before the links, state what was supposed to change and under what conditions. The person needs to know what they're looking for.
Don't hallucinate conditions. If you're not sure what URL or state produces the right conditions, say so. "I couldn't determine the exact URL to reproduce this — here's what I know and what I'd need to figure out." Partial honesty beats confident bullshit.
You are NOT the person's QA team — they are YOUR approval gate. You've already tested everything. You've already run the tests, hit the endpoints, opened the browser. This guide is their 30-second confirmation that you got it right. If they find something broken, that's YOUR failure to test, not their job to discover.
What This Has Nothing to Do With
- Running tests (you already did that —
/verification-gate, or your own pre-completion verification skill) - Producing before/after evidence packages (that's
/anthropic-proof-its-fixed) - Writing programmatic checks (that's your job during implementation)
- Debugging (if it's broken, fix it before invoking this skill)
When to Use
- After ALL agent testing is complete — this is the last thing before declaring done
- When the human says "prove it" or "is this working" — produce this immediately
- Before
/commit— don't commit until the person has stamped approval
Composition
This skill fires AFTER the agent has already:
- Run
/verification-gate(or your own pre-completion verification skill) — verified own work - Run
/anthropic-proof-its-fixedif severity warranted it (before/after evidence) - Passed all relevant tests
And BEFORE:
/commit— don't commit without approval/complete-agentic-task— don't declare done without the stamp
Saving the Evidence
After producing the proof guide, save it somewhere persistent so the person can find it later and so it isn't lost when the session ends.
If you keep an Obsidian vault (or similar) for this project (for example: a vault holding intent/scope docs), save it there:
- Path (example):
<project>/docs/intent/maintenance/evidence-of-fix-{short-name}.md - Include two-way links: link TO the relevant scope/intent docs, and have those docs link BACK to this evidence
- Open it immediately if you have a way to (e.g. an
obsidian://URL for your own vault name)
If you don't, use the local fallback: run alignment-harness records evidence to get a folder, and write the proof guide there as evidence-of-fix-{short-name}.md. Tell the person the path.
This creates a persistent trust record — the person can always go back and see what was verified and when.
Honest Framing
This skill is v1. If the verification can't be reduced to 30 seconds (e.g., the change is purely backend infrastructure with no visible UX), say so and describe the simplest alternative the person can do. Never silently skip verification because it's hard to make visual.
Feedback Loop
If this skill's output was imperfect — the links didn't work, the conditions were wrong, the person found something broken — that's signal. Note what went wrong (in the evidence record, or wherever you track corrections) so the next agent, or a future edit to this skill file, gets it right.