← the whole session plugin/skills/agentic-ux-verification/SKILL.md

Workflow for agents to capture UX screenshot evidence without marking verification status.

Agent UX Verification Skill

Purpose

Automate screenshot evidence collection for UX verification assignments. Agents provide evidence, humans provide judgment. Agents never mark items as verified - they only attach screenshots and comments.

Two ways to run this

This whole workflow needs somewhere to keep the list of "things that need a screenshot" and their status. The author built that as a small admin API (/api/.../ux-assignments/... below) — that's optional, private infrastructure most people won't have. If you haven't set up an equivalent, use the local fallback described after each step instead. Either way, the workflow and the principles are identical; only where the list and the evidence live changes.

Local fallback, in one line: keep each assignment as its own file under alignment-harness records ux-assignments/ (run that command to get the folder), with a status field (needs-screenshot / claimed / done) and a comments array in the file itself; screenshots go in the same folder as image files the record links to.

Workflow Overview

┌─────────────────────────────────────────────────────────────────┐
│  1. FETCH:                                                       │
│     If configured: GET /needs-screenshots (returns a batch,     │
│       triggers swarm) or /needs-screenshots/next (auto-claims)  │
│     Local fallback: list files under the local folder whose     │
│       status is "needs-screenshot"                               │
└─────────────────────────────────────────────────────────────────┘
                              ↓
┌─────────────────────────────────────────────────────────────────┐
│  2. CLAIM (swarm only):                                         │
│     If configured: POST .../items/:itemId/claim                 │
│     Local fallback: set status: "claimed", claimedBy, and a     │
│       claimExpiresAt field directly in the item's file           │
└─────────────────────────────────────────────────────────────────┘
                              ↓
┌─────────────────────────────────────────────────────────────────┐
│  3. EACH AGENT:                                                 │
│     a) Can verification be done with a few clicks?              │
│        YES → Navigate + interact via your browser/devtools tool │
│        NO  → Create simulator function first (see below)        │
│     b) Navigate to the correct state                             │
│     c) Take screenshot                                          │
│     d) Attach it — one API call, or save the file locally and   │
│        update the record's path field                           │
│     e) Add comments explaining what was observed                │
└─────────────────────────────────────────────────────────────────┘
                              ↓
┌─────────────────────────────────────────────────────────────────┐
│  4. RESULT:                                                     │
│     → Screenshot stored (with a TTL if your store has one)      │
│     → Comments visible wherever the person reviews these         │
│     → Agent does NOT change status to "verified"/"failed"       │
│     → Human sees the screenshot and marks it verified/failed    │
└─────────────────────────────────────────────────────────────────┘

Key Principles

  1. Agents provide evidence, not judgment - Never mark items as verified/failed
  2. Claim before working - Prevents swarm collision
  3. 10-minute claim TTL - If agent fails, item becomes available again
  4. One function to attach - Screenshot + auto-comment in single call (or, locally, one file write)
  5. Human verification - Only humans mark verified/failed status

If you've configured a UX-assignments API

The shape below is one reference setup, for reference — build the equivalent against whatever store you use, or skip this whole section and use the local fallback.

Fetch Endpoints

1. GET .../ux-assignments/needs-screenshots

Purpose: Fetch batch of items for swarm dispatch When to use: Coordinator agent fetching work for parallel processing Query params: limit (default: 20), includeClaimed (default: false) Returns:

{
  "success": true,
  "data": [
    {
      "assignmentId": "...",
      "itemId": "...",
      "title": "Active subscribers can enter the app from checkout",
      "uxStory": "When a user with an active subscription checks out...",
      "steps": ["Go to /checkout", "Enter a subscriber's email", "..."],
      "expectedResult": "Modal shows two buttons...",
      "path": "/checkout",
      "environment": "local",
      "category": "payment",
      "combinedScore": 81,
      "simulatorPath": "/admin/subscription-simulator",
      "claimedBy": null,
      "claimExpiresAt": null
    }
  ],
  "count": 20,
  "total": 47,
  "hint": "Trigger a swarm: dispatch one agent per item for parallel processing"
}

Agent action: If count > 1, trigger swarm dispatch (one agent per item)

2. GET .../ux-assignments/needs-screenshots/next

Purpose: Fetch single highest-priority item and AUTO-CLAIM it When to use: Single agent working sequentially Query params: agentId (required) Returns: single item, claimed: true, claimExpiresAt. data: null when nothing needs a screenshot.

Claim Endpoints

3. POST .../items/:itemId/claim

Body: { "agentId": "...", "durationMinutes": 10 }. 409 if already claimed by another agent, with claimedBy and claimExpiresAt.

4. DELETE .../items/:itemId/claim

Body: { "agentId": "...", "reason": "..." }. Clears the claim and adds a blocker comment.

Evidence Endpoints

5. POST .../items/:itemId/screenshot

Body: { "screenshot": "<base64>", "agentId": "...", "capturedUrl": "...", "notes": "..." }. Stores the screenshot, sets capture metadata, adds an observation comment, clears the claim.

6. POST .../items/:itemId/comments

Body: { "text": "...", "agentId": "...", "type": "observation" }.

Type When to use UI color
observation What you saw/did Gray
blocker Can't proceed, needs human Red
question Unsure if behavior is correct Yellow
resolution Human closes the loop Green

If you haven't configured one: work the local files directly

Each item is one file under alignment-harness records ux-assignments/, e.g. ux-item-042.json:

{
  "itemId": "ux-item-042",
  "title": "Active subscribers can enter the app from checkout",
  "uxStory": "When a user with an active subscription checks out...",
  "steps": ["Go to /checkout", "Enter a subscriber's email", "..."],
  "expectedResult": "Modal shows two buttons...",
  "environment": "local",
  "simulatorPath": null,
  "status": "needs-screenshot",
  "claimedBy": null,
  "claimExpiresAt": null,
  "screenshotPath": null,
  "comments": []
}
  • Fetch: list files with status: "needs-screenshot".
  • Claim: set status: "claimed", claimedBy, claimExpiresAt (now + 10 minutes).
  • Attach evidence: save the screenshot as a file in the same folder, set screenshotPath to it, append a comments entry with type: "observation", and set status back to needs-screenshot — never to verified/failed; only the human does that when they review the folder.
  • Release without finishing: set status back to needs-screenshot, claimedBy: null, and append a blocker comment explaining why.

Tell the person where the folder is so they know where to review; there's no admin UI to point them to unless they've built one.


Swarm Dispatch Pattern

When a coordinator receives multiple items (from the API or by listing the local folder):

// Coordinator fetches batch (API or local-file listing — same shape either way)
const items = /* fetch or list-local-files */;

if (items.length > 1) {
  // Dispatch swarm - one agent per item
  for (const item of items) {
    // 1. Claim item (API call, or write status:"claimed" to its file)
    // 2. Spawn agent with item context
    spawnAgent({
      task: 'capture-ux-screenshot',
      item: item,
      agentId: `worker-${i}`
    });
  }
}

When to Create Simulator Functions

If the UX condition cannot be reached with a few clicks, check for an existing simulator:

  1. Check the item's simulatorPath field
  2. If it exists, navigate there directly
  3. If not, create a simulator page in whatever admin/dev surface your project has (the author's convention: /admin/{feature}-simulator)

Environment URLs

Fill this in for your own project — local dev server, staging, and production, for each app/service you're verifying against. There's nothing generic to put here; every project's ports and domains differ.


Self-Correction Protocol

If This Skill Fails

If you encounter a scenario not covered by this skill file:

  1. Document the gap: Add a blocker comment explaining what's missing
  2. Update Known Gaps: Add a row to the table below
  3. Propose fix: Describe how the skill file should be updated

Known Gaps (Agent-Reported)

Date Agent Gap Description Proposed Fix Status

How to Update This Skill File

## Adding a Gap Entry

1. Read this skill file
2. Add row to Known Gaps table with:
   - Date (ISO format)
   - Your agent ID
   - Clear description of what failed
   - Proposed workflow addition
   - Status: "reported"
3. If you can fix it, update the relevant section and change status to "fixed"

Example Agent Session

Human: verify ux assignments

Agent (Coordinator):
1. Fetching items needing screenshots (API if configured, otherwise the local folder)...
   → Found 5 items, hint says "trigger swarm"

2. Claiming and dispatching agents...
   → item-1/claim → worker-1
   → item-2/claim → worker-2
   ... (one per item)

Agent (Worker-1):
   → Item: "Active subscribers can enter the app from checkout"
   → simulatorPath: /admin/subscription-simulator
   → Navigating there, selecting "Active - Same Plan", clicking "Preview Modal"
   → Comment: "Modal appeared with two buttons"
   → Taking screenshot, attaching it (API call or local file write)
   → Done ✓

Agent (Worker-2):
   → Item: "Error boundary shows a friendly message"
   → simulatorPath: null (need to create)
   → Creating a simulator page for it...
   → Navigating and triggering the error
   → Attaching screenshot
   → Done ✓

3. Summary:
   ✓ 5/5 items have screenshots attached
   → Ready for human verification — point them at the API's admin UI if configured,
     otherwise at the local records folder

  • how-to-add-modals-to-admin-simulator - Creating modal simulators
  • how-to-add-pages-to-admin-simulator - Page-level simulators
  • ux-assignment-generator - Creates the assignments this skill verifies
  • devtools-site-testing - Chrome DevTools MCP patterns

Triggers

This skill should be invoked when user says:

  • "verify ux assignments"
  • "capture ux screenshots"
  • "run verification swarm"
  • "check what needs screenshots"