← the whole session plugin/skills/agentic-ux-verification/SKILL.md
Workflow for agents to capture UX screenshot evidence without marking verification status.
Agent UX Verification Skill
Purpose
Automate screenshot evidence collection for UX verification assignments. Agents provide evidence, humans provide judgment. Agents never mark items as verified - they only attach screenshots and comments.
Two ways to run this
This whole workflow needs somewhere to keep the list of "things that need a screenshot" and their status. The author built that as a small admin API (/api/.../ux-assignments/... below) — that's optional, private infrastructure most people won't have. If you haven't set up an equivalent, use the local fallback described after each step instead. Either way, the workflow and the principles are identical; only where the list and the evidence live changes.
Local fallback, in one line: keep each assignment as its own file under alignment-harness records ux-assignments/ (run that command to get the folder), with a status field (needs-screenshot / claimed / done) and a comments array in the file itself; screenshots go in the same folder as image files the record links to.
Workflow Overview
┌─────────────────────────────────────────────────────────────────┐
│ 1. FETCH: │
│ If configured: GET /needs-screenshots (returns a batch, │
│ triggers swarm) or /needs-screenshots/next (auto-claims) │
│ Local fallback: list files under the local folder whose │
│ status is "needs-screenshot" │
└─────────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────────┐
│ 2. CLAIM (swarm only): │
│ If configured: POST .../items/:itemId/claim │
│ Local fallback: set status: "claimed", claimedBy, and a │
│ claimExpiresAt field directly in the item's file │
└─────────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────────┐
│ 3. EACH AGENT: │
│ a) Can verification be done with a few clicks? │
│ YES → Navigate + interact via your browser/devtools tool │
│ NO → Create simulator function first (see below) │
│ b) Navigate to the correct state │
│ c) Take screenshot │
│ d) Attach it — one API call, or save the file locally and │
│ update the record's path field │
│ e) Add comments explaining what was observed │
└─────────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────────┐
│ 4. RESULT: │
│ → Screenshot stored (with a TTL if your store has one) │
│ → Comments visible wherever the person reviews these │
│ → Agent does NOT change status to "verified"/"failed" │
│ → Human sees the screenshot and marks it verified/failed │
└─────────────────────────────────────────────────────────────────┘
Key Principles
- Agents provide evidence, not judgment - Never mark items as verified/failed
- Claim before working - Prevents swarm collision
- 10-minute claim TTL - If agent fails, item becomes available again
- One function to attach - Screenshot + auto-comment in single call (or, locally, one file write)
- Human verification - Only humans mark verified/failed status
If you've configured a UX-assignments API
The shape below is one reference setup, for reference — build the equivalent against whatever store you use, or skip this whole section and use the local fallback.
Fetch Endpoints
1. GET .../ux-assignments/needs-screenshots
Purpose: Fetch batch of items for swarm dispatch
When to use: Coordinator agent fetching work for parallel processing
Query params: limit (default: 20), includeClaimed (default: false)
Returns:
{
"success": true,
"data": [
{
"assignmentId": "...",
"itemId": "...",
"title": "Active subscribers can enter the app from checkout",
"uxStory": "When a user with an active subscription checks out...",
"steps": ["Go to /checkout", "Enter a subscriber's email", "..."],
"expectedResult": "Modal shows two buttons...",
"path": "/checkout",
"environment": "local",
"category": "payment",
"combinedScore": 81,
"simulatorPath": "/admin/subscription-simulator",
"claimedBy": null,
"claimExpiresAt": null
}
],
"count": 20,
"total": 47,
"hint": "Trigger a swarm: dispatch one agent per item for parallel processing"
}
Agent action: If count > 1, trigger swarm dispatch (one agent per item)
2. GET .../ux-assignments/needs-screenshots/next
Purpose: Fetch single highest-priority item and AUTO-CLAIM it
When to use: Single agent working sequentially
Query params: agentId (required)
Returns: single item, claimed: true, claimExpiresAt. data: null when nothing needs a screenshot.
Claim Endpoints
3. POST .../items/:itemId/claim
Body: { "agentId": "...", "durationMinutes": 10 }. 409 if already claimed by another agent, with claimedBy and claimExpiresAt.
4. DELETE .../items/:itemId/claim
Body: { "agentId": "...", "reason": "..." }. Clears the claim and adds a blocker comment.
Evidence Endpoints
5. POST .../items/:itemId/screenshot
Body: { "screenshot": "<base64>", "agentId": "...", "capturedUrl": "...", "notes": "..." }. Stores the screenshot, sets capture metadata, adds an observation comment, clears the claim.
6. POST .../items/:itemId/comments
Body: { "text": "...", "agentId": "...", "type": "observation" }.
| Type | When to use | UI color |
|---|---|---|
observation |
What you saw/did | Gray |
blocker |
Can't proceed, needs human | Red |
question |
Unsure if behavior is correct | Yellow |
resolution |
Human closes the loop | Green |
If you haven't configured one: work the local files directly
Each item is one file under alignment-harness records ux-assignments/, e.g. ux-item-042.json:
{
"itemId": "ux-item-042",
"title": "Active subscribers can enter the app from checkout",
"uxStory": "When a user with an active subscription checks out...",
"steps": ["Go to /checkout", "Enter a subscriber's email", "..."],
"expectedResult": "Modal shows two buttons...",
"environment": "local",
"simulatorPath": null,
"status": "needs-screenshot",
"claimedBy": null,
"claimExpiresAt": null,
"screenshotPath": null,
"comments": []
}
- Fetch: list files with
status: "needs-screenshot". - Claim: set
status: "claimed",claimedBy,claimExpiresAt(now + 10 minutes). - Attach evidence: save the screenshot as a file in the same folder, set
screenshotPathto it, append acommentsentry withtype: "observation", and setstatusback toneeds-screenshot— never toverified/failed; only the human does that when they review the folder. - Release without finishing: set
statusback toneeds-screenshot,claimedBy: null, and append ablockercomment explaining why.
Tell the person where the folder is so they know where to review; there's no admin UI to point them to unless they've built one.
Swarm Dispatch Pattern
When a coordinator receives multiple items (from the API or by listing the local folder):
// Coordinator fetches batch (API or local-file listing — same shape either way)
const items = /* fetch or list-local-files */;
if (items.length > 1) {
// Dispatch swarm - one agent per item
for (const item of items) {
// 1. Claim item (API call, or write status:"claimed" to its file)
// 2. Spawn agent with item context
spawnAgent({
task: 'capture-ux-screenshot',
item: item,
agentId: `worker-${i}`
});
}
}
When to Create Simulator Functions
If the UX condition cannot be reached with a few clicks, check for an existing simulator:
- Check the item's
simulatorPathfield - If it exists, navigate there directly
- If not, create a simulator page in whatever admin/dev surface your project has (the author's convention:
/admin/{feature}-simulator)
Environment URLs
Fill this in for your own project — local dev server, staging, and production, for each app/service you're verifying against. There's nothing generic to put here; every project's ports and domains differ.
Self-Correction Protocol
If This Skill Fails
If you encounter a scenario not covered by this skill file:
- Document the gap: Add a
blockercomment explaining what's missing - Update Known Gaps: Add a row to the table below
- Propose fix: Describe how the skill file should be updated
Known Gaps (Agent-Reported)
| Date | Agent | Gap Description | Proposed Fix | Status |
|---|---|---|---|---|
How to Update This Skill File
## Adding a Gap Entry
1. Read this skill file
2. Add row to Known Gaps table with:
- Date (ISO format)
- Your agent ID
- Clear description of what failed
- Proposed workflow addition
- Status: "reported"
3. If you can fix it, update the relevant section and change status to "fixed"
Example Agent Session
Human: verify ux assignments
Agent (Coordinator):
1. Fetching items needing screenshots (API if configured, otherwise the local folder)...
→ Found 5 items, hint says "trigger swarm"
2. Claiming and dispatching agents...
→ item-1/claim → worker-1
→ item-2/claim → worker-2
... (one per item)
Agent (Worker-1):
→ Item: "Active subscribers can enter the app from checkout"
→ simulatorPath: /admin/subscription-simulator
→ Navigating there, selecting "Active - Same Plan", clicking "Preview Modal"
→ Comment: "Modal appeared with two buttons"
→ Taking screenshot, attaching it (API call or local file write)
→ Done ✓
Agent (Worker-2):
→ Item: "Error boundary shows a friendly message"
→ simulatorPath: null (need to create)
→ Creating a simulator page for it...
→ Navigating and triggering the error
→ Attaching screenshot
→ Done ✓
3. Summary:
✓ 5/5 items have screenshots attached
→ Ready for human verification — point them at the API's admin UI if configured,
otherwise at the local records folder
Related Skills
how-to-add-modals-to-admin-simulator- Creating modal simulatorshow-to-add-pages-to-admin-simulator- Page-level simulatorsux-assignment-generator- Creates the assignments this skill verifiesdevtools-site-testing- Chrome DevTools MCP patterns
Triggers
This skill should be invoked when user says:
- "verify ux assignments"
- "capture ux screenshots"
- "run verification swarm"
- "check what needs screenshots"