← the whole session plugin/skills/agent-ux-verification/SKILL.md

Collect real screenshot evidence for a queued UX story by navigating to the state described, capturing it, and attaching an observation comment — never marking the item verified or failed yourself. Agents provide evidence; a human provides judgment. Works against a local, file-based queue by default.

Agent UX Verification

Purpose

Automate screenshot evidence collection for queued UX stories. Agents provide evidence, humans provide judgment. Agents never mark items as verified or failed — they only attach screenshots and comments. This is the same discipline behind the harness's evidence gate applied to visual/UX claims specifically: an agent saying "this works" is not enough; a human has to be able to look at what actually happened.

Where the queue lives

Default: a local, file-based queue, no server required. Run alignment-harness records ux-assignments for the folder. Each assignment is one JSON file: { "id": "...", "title": "...", "uxStory": "...", "steps": [...], "expectedResult": "...", "path": "...", "environment": "local", "category": "...", "howToReachState": "click-through | test-account | simulator", "simulatorPath": null, "claimedBy": null, "claimExpiresAt": null, "screenshotPath": null, "comments": [] }.

If the queue is empty, say so plainly: "no UX stories queued yet, and no way to reach any state to screenshot" — never report a clean queue that just happens to have nothing in it as if verification succeeded. Offer to draft a starter list from the person's own intent map (if they have one) or from reading their route/component list for anything that looks like a distinct user-facing flow, and ask them to confirm or edit each one before you run it.

If you've built (or a reference setup provides) a real admin-backed queue with its own HTTP API — fetch/claim/screenshot/comment endpoints, a swarm-dispatch hint, TTL-based claims — the same workflow below applies over HTTP instead of local files. That's optional, advanced infrastructure; if you have it, put a real key requirement on it or make it loopback-only, since an admin route that answers to anyone reachable on the port is a real gap worth closing before you rely on it.

Workflow Overview

┌─────────────────────────────────────────────────────────────────┐
│  1. FETCH:                                                       │
│     Read the local queue folder for unclaimed items,             │
│     or fetch from your own admin API if you have one             │
└─────────────────────────────────────────────────────────────────┘
                              ↓
┌─────────────────────────────────────────────────────────────────┐
│  2. CLAIM (when running more than one agent at once):            │
│     Write claimedBy + claimExpiresAt into the item's own file    │
│     → Advisory lock (10 minutes), prevents two agents            │
│       grabbing the same item                                     │
└─────────────────────────────────────────────────────────────────┘
                              ↓
┌─────────────────────────────────────────────────────────────────┐
│  3. EACH AGENT:                                                 │
│     a) Can this state be reached with a few clicks as a real     │
│        user would? Or with a test account?                       │
│        YES → Navigate + interact via a browser tool               │
│        NO  → Do you have (or can you build) an admin simulator   │
│              for this state? Use it. If neither exists, say so   │
│              as a blocker comment and stop — don't guess a URL.  │
│     b) Navigate to the correct state                             │
│     c) Take a screenshot                                         │
│     d) Attach it to the item (write the file path into the       │
│        item's own record, or POST it if you're on an API)        │
│     e) Add a comment explaining what was observed                │
└─────────────────────────────────────────────────────────────────┘
                              ↓
┌─────────────────────────────────────────────────────────────────┐
│  4. RESULT:                                                     │
│     → Screenshot referenced from the item's own record           │
│     → Comments form a running feed on that item                  │
│     → Agent does NOT change status                                │
│     → Human looks at the screenshot                               │
│     → Human marks verified/failed                                 │
└─────────────────────────────────────────────────────────────────┘

Key Principles

  1. Agents provide evidence, not judgment — never mark items as verified/failed. This is the whole reason the skill exists; don't add a dial for it.
  2. Claim before working — prevents two agents duplicating the same item when running more than one at once.
  3. 10-minute claim TTL — if an agent fails or stalls, the item becomes available again.
  4. One step to attach evidence — screenshot + observation comment, written together.
  5. Human verification — only a human marks an item verified or failed.

Local queue operations (default)

Fetch

Read every *.json file in the queue folder; an item is available if claimedBy is null or its claimExpiresAt has passed.

Claim

Write claimedBy: "<your agent id>" and claimExpiresAt: "<now + 10 minutes, ISO>" into the item's file before starting work on it. If you can't complete it, clear the claim and add a blocker comment explaining why — so another agent (or a re-run) can pick it up.

Attach evidence

Save the screenshot file, set screenshotPath on the item to that file's path, and append a comment (see types below). This does not change the item's status — it only means the evidence is ready for a human to look at.

Comment types

Type When to use Meaning
observation What you saw/did Plain narration of what happened
blocker Can't proceed, needs a human You hit something you can't resolve yourself
question Unsure if behavior is correct You have evidence but aren't sure it's right
resolution Human closes the loop The human's own note when they mark it verified/failed

When you need more than a few clicks to reach a state

If the UX condition cannot be reached with a few clicks as a real user would, or with a test account:

  1. Check the item's howToReachState / simulatorPath field for a pointer to an existing simulator or seed script.
  2. If one exists, use it.
  3. If not, and you have the ability to build one, create it — but this is your own project's own convention to establish, not something this skill prescribes. (One example project uses a /admin2/{feature}-simulator pattern; that's one example shape, not a requirement.)
  4. If neither a real path nor a simulator exists, say so as a blocker comment on the item and stop. Never guess a URL and report success without ever loading it.

Environments

Ask the person, once, what their local / staging / production URLs are for however many services their product has (could be just one) — during setup or the first time this skill runs — and use that instead of assuming any particular port or domain. Don't hardcode a fixed table of repo names and ports; every product's setup is different.


Self-Correction Protocol

If This Skill Fails

If you encounter a scenario not covered by this skill file:

  1. Document the gap: add a blocker comment explaining what's missing.
  2. Update Known Gaps: add a row to the table below.
  3. Propose a fix: describe how the skill file should be updated.

Known Gaps (Agent-Reported)

Date Agent Gap Description Proposed Fix Status

How to Update This Skill File

## Adding a Gap Entry

1. Read this skill file
2. Add a row to Known Gaps with:
   - Date (ISO format)
   - Your agent ID
   - Clear description of what failed
   - Proposed workflow addition
   - Status: "reported"
3. If you can fix it, update the relevant section and change status to "fixed"

Example Agent Session

Human: verify ux assignments

Agent (Coordinator):
1. Reading the local queue...
   → Found 5 items, none claimed

2. Claiming and dispatching agents (one per item)...

Agent (Worker-1):
   → Item: "Active subscribers can enter the app from checkout"
   → howToReachState: use the checkout simulator
   → Navigating to the simulator path
   → Selecting "Active - Same Plan"
   → Clicking "Preview Modal"
   → Comment: "Modal appeared with two buttons, as expected"
   → Taking screenshot, attaching it to the item
   → Done ✓

Agent (Worker-2):
   → Item: "Error boundary shows a friendly message"
   → howToReachState: none set, no simulator exists
   → Comment (blocker): "No way to reach this state — need a test path or a simulator before I can verify this"
   → Stops, does not guess

3. Summary:
   ✓ 4/5 items have screenshots attached
   ⚠ 1/5 blocked, needs a way to reach its state
   → Ready for human review

  • how-to-add-modals-to-admin-simulator — creating modal simulators, if you use that pattern
  • how-to-add-pages-to-admin-simulator — page-level simulators, if you use that pattern
  • Whatever generates your own UX stories/assignments in the first place — this skill only verifies a queue, it doesn't create one
  • devtools-site-testing — browser-driving patterns