← the whole session plugin/skills/agent-ux-verification/SKILL.md
Collect real screenshot evidence for a queued UX story by navigating to the state described, capturing it, and attaching an observation comment — never marking the item verified or failed yourself. Agents provide evidence; a human provides judgment. Works against a local, file-based queue by default.
Agent UX Verification
Purpose
Automate screenshot evidence collection for queued UX stories. Agents provide evidence, humans provide judgment. Agents never mark items as verified or failed — they only attach screenshots and comments. This is the same discipline behind the harness's evidence gate applied to visual/UX claims specifically: an agent saying "this works" is not enough; a human has to be able to look at what actually happened.
Where the queue lives
Default: a local, file-based queue, no server required. Run alignment-harness records ux-assignments for the folder. Each assignment is one JSON file: { "id": "...", "title": "...", "uxStory": "...", "steps": [...], "expectedResult": "...", "path": "...", "environment": "local", "category": "...", "howToReachState": "click-through | test-account | simulator", "simulatorPath": null, "claimedBy": null, "claimExpiresAt": null, "screenshotPath": null, "comments": [] }.
If the queue is empty, say so plainly: "no UX stories queued yet, and no way to reach any state to screenshot" — never report a clean queue that just happens to have nothing in it as if verification succeeded. Offer to draft a starter list from the person's own intent map (if they have one) or from reading their route/component list for anything that looks like a distinct user-facing flow, and ask them to confirm or edit each one before you run it.
If you've built (or a reference setup provides) a real admin-backed queue with its own HTTP API — fetch/claim/screenshot/comment endpoints, a swarm-dispatch hint, TTL-based claims — the same workflow below applies over HTTP instead of local files. That's optional, advanced infrastructure; if you have it, put a real key requirement on it or make it loopback-only, since an admin route that answers to anyone reachable on the port is a real gap worth closing before you rely on it.
Workflow Overview
┌─────────────────────────────────────────────────────────────────┐
│ 1. FETCH: │
│ Read the local queue folder for unclaimed items, │
│ or fetch from your own admin API if you have one │
└─────────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────────┐
│ 2. CLAIM (when running more than one agent at once): │
│ Write claimedBy + claimExpiresAt into the item's own file │
│ → Advisory lock (10 minutes), prevents two agents │
│ grabbing the same item │
└─────────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────────┐
│ 3. EACH AGENT: │
│ a) Can this state be reached with a few clicks as a real │
│ user would? Or with a test account? │
│ YES → Navigate + interact via a browser tool │
│ NO → Do you have (or can you build) an admin simulator │
│ for this state? Use it. If neither exists, say so │
│ as a blocker comment and stop — don't guess a URL. │
│ b) Navigate to the correct state │
│ c) Take a screenshot │
│ d) Attach it to the item (write the file path into the │
│ item's own record, or POST it if you're on an API) │
│ e) Add a comment explaining what was observed │
└─────────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────────┐
│ 4. RESULT: │
│ → Screenshot referenced from the item's own record │
│ → Comments form a running feed on that item │
│ → Agent does NOT change status │
│ → Human looks at the screenshot │
│ → Human marks verified/failed │
└─────────────────────────────────────────────────────────────────┘
Key Principles
- Agents provide evidence, not judgment — never mark items as verified/failed. This is the whole reason the skill exists; don't add a dial for it.
- Claim before working — prevents two agents duplicating the same item when running more than one at once.
- 10-minute claim TTL — if an agent fails or stalls, the item becomes available again.
- One step to attach evidence — screenshot + observation comment, written together.
- Human verification — only a human marks an item verified or failed.
Local queue operations (default)
Fetch
Read every *.json file in the queue folder; an item is available if claimedBy is null or its claimExpiresAt has passed.
Claim
Write claimedBy: "<your agent id>" and claimExpiresAt: "<now + 10 minutes, ISO>" into the item's file before starting work on it. If you can't complete it, clear the claim and add a blocker comment explaining why — so another agent (or a re-run) can pick it up.
Attach evidence
Save the screenshot file, set screenshotPath on the item to that file's path, and append a comment (see types below). This does not change the item's status — it only means the evidence is ready for a human to look at.
Comment types
| Type | When to use | Meaning |
|---|---|---|
observation |
What you saw/did | Plain narration of what happened |
blocker |
Can't proceed, needs a human | You hit something you can't resolve yourself |
question |
Unsure if behavior is correct | You have evidence but aren't sure it's right |
resolution |
Human closes the loop | The human's own note when they mark it verified/failed |
When you need more than a few clicks to reach a state
If the UX condition cannot be reached with a few clicks as a real user would, or with a test account:
- Check the item's
howToReachState/simulatorPathfield for a pointer to an existing simulator or seed script. - If one exists, use it.
- If not, and you have the ability to build one, create it — but this is your own project's own convention to establish, not something this skill prescribes. (One example project uses a
/admin2/{feature}-simulatorpattern; that's one example shape, not a requirement.) - If neither a real path nor a simulator exists, say so as a
blockercomment on the item and stop. Never guess a URL and report success without ever loading it.
Environments
Ask the person, once, what their local / staging / production URLs are for however many services their product has (could be just one) — during setup or the first time this skill runs — and use that instead of assuming any particular port or domain. Don't hardcode a fixed table of repo names and ports; every product's setup is different.
Self-Correction Protocol
If This Skill Fails
If you encounter a scenario not covered by this skill file:
- Document the gap: add a
blockercomment explaining what's missing. - Update Known Gaps: add a row to the table below.
- Propose a fix: describe how the skill file should be updated.
Known Gaps (Agent-Reported)
| Date | Agent | Gap Description | Proposed Fix | Status |
|---|
How to Update This Skill File
## Adding a Gap Entry
1. Read this skill file
2. Add a row to Known Gaps with:
- Date (ISO format)
- Your agent ID
- Clear description of what failed
- Proposed workflow addition
- Status: "reported"
3. If you can fix it, update the relevant section and change status to "fixed"
Example Agent Session
Human: verify ux assignments
Agent (Coordinator):
1. Reading the local queue...
→ Found 5 items, none claimed
2. Claiming and dispatching agents (one per item)...
Agent (Worker-1):
→ Item: "Active subscribers can enter the app from checkout"
→ howToReachState: use the checkout simulator
→ Navigating to the simulator path
→ Selecting "Active - Same Plan"
→ Clicking "Preview Modal"
→ Comment: "Modal appeared with two buttons, as expected"
→ Taking screenshot, attaching it to the item
→ Done ✓
Agent (Worker-2):
→ Item: "Error boundary shows a friendly message"
→ howToReachState: none set, no simulator exists
→ Comment (blocker): "No way to reach this state — need a test path or a simulator before I can verify this"
→ Stops, does not guess
3. Summary:
✓ 4/5 items have screenshots attached
⚠ 1/5 blocked, needs a way to reach its state
→ Ready for human review
Related Skills
how-to-add-modals-to-admin-simulator— creating modal simulators, if you use that patternhow-to-add-pages-to-admin-simulator— page-level simulators, if you use that pattern- Whatever generates your own UX stories/assignments in the first place — this skill only verifies a queue, it doesn't create one
devtools-site-testing— browser-driving patterns