← the whole session plugin/skills/founder-watchable-canonical/SKILL.md

CANONICAL entry for person-watchable browser testing with screen recording. Pins the walkthrough doc, the recording command, the cohort definitions, and the append-only Instincts the person has taught across sessions. Use when the person wants to watch an agent test the UX live, or when learning a new instinct that should never be forgotten.

Person-Watchable Walkthrough — CANONICAL

(Named founder-watchable because it was written for a solo founder watching their own product; the method works the same for anyone reviewing a walkthrough live — a PM, a solo dev, a team lead. Read "the person" below as whoever is watching.)

Status: CANONICAL. This is the single entry point for ALL UX-validation walks. A handful of prior overlapping skills are referenced below — use them ONLY when their specialty applies; otherwise this canonical is the front door.

Lifecycle: Active. Last verified 2026-05-12.


📋 MENU — what's in this file (read first)

Section What you get When you need it
🎯 INTENT The 3 things this skill exists to do First read — orient on purpose
🧠 INSTINCTS Append-only behavioral rules the person has taught (currently 27+) Read BEFORE asking the person anything OR proposing a fix
📚 Source-of-truth doc Full 8-step recipe + overlay-injection JS When you need the exact mechanical sequence
🧪 Launch command Bash + MCP sequence to start a walk First-time launch of a new walk
🔧 Agent tools (a) Resume from checkpoint, (b) Side-panel overlay variant Resuming a mid-walk OR the person requests side-panel layout
🎭 Test accounts you+<number>@yourdomain.com plus-aliasing format + login institutional knowledge Any walk needing a fresh signup
🚫 Anti-patterns What NOT to do (auto-browser, Playwright --headed, sleep pacing, etc.) Before reaching for a tool you "remember" worked
🗺️ Supersession map Which prior skills are subsumed by this one + when to use them anyway When you almost reach for an old skill — check here first

🧩 HOW THIS SKILL IS MODULAR — read this if you find yourself reinventing

This canonical is a growing menu of reusable modules, not a static document. The pattern:

  1. You discover a new technique during a walk (a better recording flag, a clearer overlay layout, a new checkpoint format, a way to detect a specific failure class)
  2. You add it to this file — either as a new instinct (behavioral rule) or a new agent-tools entry (reusable mechanical capability)
  3. Future walks load this file and inherit your module — they don't reinvent the wheel, they stack on top

If you're about to write a one-off helper, ask first: "Does this belong in the canonical so the next walk has it?" If yes, add it here. If it's truly walk-specific, write it inline.

The same applies to the person's feedback: every correction they give → instinct here → next session inherits the lesson, no re-teaching.


🎯 INTENT — what this skill exists to do

  • autonomous testing with full visibility to admin
  • detection and proposed fixes of bugs for admin to approve with /sanity-check
  • (reserved — the person may add)

🧠 INSTINCTS — append-only, person-taught

Every entry below is a behavioral rule the person has stated. Append-only. Date-stamped. Newest at bottom.

Confirm before acting: When the person gives feedback, repeat it back verbatim (or close paraphrase) before changing behavior. This is the hallucination-detection surface — if I misquote, the person catches it before the wrong lesson lands in this file.

Read this section BEFORE asking the person anything. If the answer is already here, act.

  • 2026-05-12 — Prefer autonomy over making the user decide arbitrary things. Autonomous testing is visible to the admin. Take the most direct path to the stated goal. Do not offload arbitrary choices back to the person mid-run.

  • 2026-05-12 — Prefer testing things that actually had code changes to make sure those code changes worked, over random testing surfaces. The scope is defined by the diff, not by my urge to be thorough.

  • 2026-05-12 — Prefer high-impact critical-path tests over arbitrary tests. One end-to-end story over a buffet of unit-style cohort scans.

  • 2026-05-12 — User feedback → update intent & instincts in THIS skill file. Every correction the person gives me must update this file before I change behavior. Behavior-only correction = compounding loss.

  • 2026-05-12 — Prefer localhost as default unless otherwise specified. When testing code changes, the running local stack is the higher-fidelity target.

  • 2026-05-12 — Persona-driven coaching, not "this is a test" strings. When the test involves real coaching, simulate a high-LTV persona from ~/.claude/skills/critique-from-target-audience/personas/persona-{1..10}.md. Stacks integration testing with coaching-quality data on the same compute spend.

  • 2026-05-12 — One walk, one user, one watch. Founder-watchable = integration-as-story, not unit-test-of-cohorts. Cohort matrix belongs in the retrospective, never in the plan.

  • 2026-05-12 — Recording must survive shell exit. Plain & in a single Bash call dies when the call ends. Use run_in_background: true or nohup + detached subshell.

  • 2026-05-12 — Use mcp__chrome-devtools__navigate_page with generous timeout, NOT evaluate_script + window.location.href. The latter kills the MCP page target on CRA dev. The canonical doc's persistent-connection warning extends to dev-mode navigation.

  • 2026-05-12 — Existing institutional knowledge beats improvisation. Before generating ANY test artifact (account, email, password, payload), search for an existing canonical skill that defines it. For login/test-accounts: if the project has its own login/test-account skill, that's the source of truth. For personas: one example is critique-from-target-audience's personas/ folder, if installed. If you don't know if a skill exists, search before improvising.

  • 2026-05-12 — Test accounts MUST use the person's own email with plus-aliasing: <their-local-part>+<number>@<their-domain> format. Variable: the digits after +. Fixed: everything else. Reasoning: (1) Gmail (and most providers') plus-aliasing delivers to the person's real inbox so they can verify the actual emails the system sends. (2) Trivial cleanup via User.deleteMany({ email: /^<their-local-part>\+\d+@<their-domain>$/i }) (ask the person for their real local-part/domain once, then reuse it). Never invent a test domain. Never use an email the person doesn't own or control. Shape of the convention (fill in your own address): you+<number>@yourdomain.com, cleaned up with User.deleteMany({ email: /^you\+\d+@yourdomain\.com$/i }) — swap you and yourdomain.com for the person's real, owned address. Never write anyone's actual email address into this file, including your own — keep it in the person's private test notes instead.

  • 2026-05-12 — Never propose a fix without reading every named function in your causal chain. If your diagnosis names a function (handleLoginSuccess, issueRefreshToken, getIdToken), open and read each one before claiming you know the bug. A bug story built on read-some-call-the-rest is a probable story, not a certain one. The person's bar for "go fix it" is certainty grounded in code you have read, not in network logs you have inferred from.

  • 2026-05-12 — Reproduce a failure twice with clean state before believing it. One observation is data, two with isolation is signal. A bug that fires once may be stale cookie, prior session JWT, browser state, or race. Re-test with a fresh isolatedContext and zero prior auth state before naming the root cause.

  • 2026-05-12 — The person testing a different path successfully is evidence your diagnosis is incomplete. When the person says "I just tested X and didn't see the error you described" — that is not a vibe check, it is a real datapoint. Surface the contradiction explicitly: "Your test of [path] passing means the bug is specific to [different path] OR [state condition], because [reasoning]." Then either close the gap or back off the diagnosis. Never wave the contradiction away.

  • 2026-05-12 — Confidence-tag every fix proposal. State the confidence level explicitly before asking to proceed: "Confidence: 60% — I have read N of M named functions, reproduced 1 of 2 needed times, ruled out X but not Y." The person grants permission based on the confidence, not on the proposal sounding plausible.

  • 2026-05-12 — Programmatic fix ≠ UX fix. The test is "does the user reach the original intent?" A try/catch that lets the user SEE the error is NOT a fix. The user's experience must actually succeed. For any error-path fix, the minimum bar is: (1) state resets so retry is possible, (2) the system auto-retries with backoff before surfacing failure to the user, (3) during retry, a soft human-language status message communicates "this is taking longer than expected, but we're working on it" rather than a dead spinner, (4) the user eventually reaches the original intent without ever knowing the path was rocky — OR if all retries truly fail, they get a clear "try again" button that works the second time, never a permanent dead state. Anything less is destruction-of-intent disguised as a fix.

  • 2026-05-12 — Production data > one browser observation. Query the funnel before believing a single-walk bug claim. When a walk surfaces a bug that looks like "everyone is broken," dispatch a subagent to query the last 7-30 days of real production conversion data for the affected funnel step. If conversion is at normal levels, the bug is environment-specific (your browser, your state, your race condition) and should NOT be shipped as a production fix — investigate further first. If conversion is at zero or dramatically depressed, the bug is universal and the fix is urgent. Real data tells you which one.

  • 2026-05-12 — Every bug-fix proposal MUST include the optimal UX, not just the programmatic fix. Forgetting to think about what the user experiences is a CRITICAL error in agent thinking. The product IS the experience, not the code. Every backend failure is a frontend opportunity to curate: a 500 becomes "we're working on it — your data is safe," a timeout becomes "hang tight, this is taking a moment," a 401 becomes a silent re-auth the user never sees, a 403 becomes a clear "log in again" with one-click action. Before proposing ANY bug fix, ask: "What does the user see at each moment along the failure path, and how do we curate that into a graceful experience that still gets them to their original intent?" If you only describe what the code does, you have failed the proposal. If a fix proposal causes no harm in the happy path AND curates the failure path gracefully, it is safe to ship even without root-cause certainty — defensive UX is strictly additive value.

  • 2026-05-12 — Stale Chrome DevTools MCP profile cookies are a known false-positive bug source. The chrome-devtools-mcp profile at ~/.cache/chrome-devtools-mcp/chrome-profile/ persists cookies across walks even with isolatedContext. A failed walk that left a partial-auth state can corrupt the next walk's signup flow. Always confirm: did this bug reproduce in a fresh-profile run? If you only have one observation in the MCP browser and production data refutes universality, the MCP profile is the most likely culprit. Mitigation: kill the entire profile (not just the SingletonLock) between walks that involve auth state.

  • 2026-05-12 — The most important lesson of 2026-05-12: query production data BEFORE building the bug story. The cost of one subagent query (~30 seconds) is dramatically less than the cost of writing a 15-minute diagnosis based on one observation that turns out to be environment-specific. If a walk surfaces what looks like a "100% broken" funnel bug, the FIRST action is the production-data query, not the code-reading diagnosis. The diagnosis is for AFTER you've confirmed the bug is real.

  • 2026-05-12 — Know which gates in the product are INTENDED behavior vs a bug, before you walk one. Most products have more than one gate along a funnel (an anonymous-usage cap, a signup wall, a trial-usage cap, a paywall), and they can be nested — hitting one on the way to testing another is often correct, not a finding. Before starting a walk that needs to reach a specific gate, read the code (or ask the person) which gates exist between the start and the target, in what order, and confirm each intermediate one firing is expected rather than flagging it as a bug in the walk report. Example (opt-in illustration): anonymous users get 4 free exchanges, then a signup modal fires on message 5 — that's the intended conversion mechanism, not a bug. A separate, later gate around message 16 (in that codebase's usage-check middleware) is the actual trial-limit gate for signed-up users. A walk validating the trial-limit gate has to pass through the earlier signup gate first, by design.

  • 2026-05-12 — "Test reached the goal" means the original test target fired, not "everything except the goal worked." Stopping at "I validated the path but didn't drive to the actual gate the test was designed to verify" is a failure of completion disguised as strategic prudence. If the user asked "does the gate fire at msg 16," done = you saw the gate fire at msg 16. Until then, the test is in-progress, period. Compute-budget rationalizations for stopping early are a tell that you forgot the test target.

  • 2026-05-12 — Walks MUST support resume-from-checkpoint. Every walk writes a checkpoint file at /tmp/watch-checkpoint-<slug>.json after each substantive milestone (signup gate fired, registration succeeded, authed dashboard reached, msg N completed). A resume command (/founder-watchable-canonical resume <slug>) re-opens the browser to the dashboard session ID, re-injects the overlay, and continues from the last checkpoint message count — no re-driving from msg 1. The checkpoint file MUST contain: sessionId (URL path segment), test account email, last completed user-message count, last coach-message preview, recording path if active, walk goal string.

  • 2026-05-12 — Side-positioned overlay (when the person requests it) MUST be done via injected DOM only — zero risk to the app's own UI code. Strategy: create a wrapper <div> that becomes the new viewport for the app by transforming document.body width to calc(100vw - 460px) via inline style on body, with the overlay panel rendered at position: fixed; right: 0; width: 460px; height: 100vh;. ALL CSS lives in the overlay-injection JS — none in the app's own source. The app sees a narrower viewport and naturally responsive-shrinks. If anything breaks, the person dismisses the overlay and body style is removed — the app's own UI is untouched. Test the side-panel variant against the dashboard, coaching surface, and at least one modal before declaring it safe. If ANY UI breakage detected, fall back to the top-right overlay variant.

  • 2026-05-12 — Substantive findings → a written report, always, in a fixed place. When a walk produces substantive findings — bugs found, fixes shipped, instrumentation gaps, quality observations, UX moments worth fixing for users — write a markdown report at <project>/docs/qa/browser-testing/walk-reports/walk-YYYY-MM-DD-<slug>.md (create the folder if it doesn't exist). Write it in the person's own established voice if one is set up (/how-to-talk-like-the-founder, which explains how to calibrate to the person's own writing or falls back to an opt-in example), otherwise plain and direct is fine. If the person has an Obsidian vault or other note viewer configured for this project (see /alignment-harness:harness-setup), open it there; otherwise print the file path. The walk-report is the durable artifact; the .mov is supporting evidence. Without the report, the work evaporates after the session ends.

  • 2026-05-12 — Resume-from-checkpoint requires explicit auth-restoration when Chrome MCP profile cycles. The original resume protocol assumed cookies survive across new_page calls. They don't — killing the Chrome MCP profile (which we do to clear stale-cookie pollution) drops the auth cookie. Resume protocol must: (1) clear MCP lock only, NOT the whole profile if auth-continuity matters, OR (2) accept the profile kill AND add a login-form step before navigating to the session URL. Use the project's own login skill for the credentials, if one exists, or the persona's test account if the signup flow is being tested.

  • 2026-05-12 — When a backend rejection is correct but the frontend doesn't render its response (e.g. a modal that should fire doesn't), that IS the UX bug — not a "button broken" bug. If a button appears non-responsive at a specific user state (a gate, a limit, an error), first verify the backend response before assuming the request never fired — the bug is often "frontend doesn't handle this response shape," not "frontend can't fire the request." Example (opt-in illustration): during one walk's resume, the send button appeared to silently swallow clicks. The person manually verified send worked for non-gated messages; the usage-check middleware was correctly rejecting the request; the actual bug was the frontend handler not catching that rejection and opening the expected upgrade modal.

  • 2026-05-12 — Walk reports go in <project>/docs/qa/browser-testing/walk-reports/YYYY-MM-DD-<slug>.md (the same default path as the instinct above — keep them in one place). Frontmatter must include session-date, walks-covered (array), test-target, verdict, recording-files (array), related-canonical. Body in the person's own voice if calibrated (/how-to-talk-like-the-founder), otherwise plain and direct. Body sections: (1) what the session was for in plain language, (2) one section per walk, (3) findings numbered + severity-tagged, (4) durable artifacts produced, (5) what's still open as a table. Open it wherever the person reads notes (their configured vault, or just the file path) immediately after writing.

  • 2026-05-12 — Score 80+ critical-path UX bugs use /governer2 invocation, not direct fixes. When a UX bug surfaces that affects the conversion gateway, payment flow, or any code path where a wrong fix could regress every user, escalate to /governer2 {description}. The pipeline produces: variant announce, score, meta-assessment, UX-Breakage Reasoning Gate (Q1-Q4), required skill sequence, autonomous-commit verdict. Q4 item count > 0 means HUMAN-REQUIRED — sanity-check + browser-evidence-per-item + explicit authorization from the person before commit. Demonstrated 2026-05-12 on legacy-trial gate-modal handoff at score 82/100 with 10 Q4 items (one worked example — your own numbers will differ).

  • 2026-05-12 — Skill hygiene is recurring, not one-off. Run the biweekly supersession audit agent (skill-supersession-audit-agent.md in the browser-testing domain) every 14 days. It scans ~/.claude/skills/ for new overlap clusters, produces a proposal in Obsidian, waits for the person's approval. Proposes-never-executes. Without this loop, the cleanup compounds and the canonical front-doors get muddier over 90 days.

  • 2026-05-12 — Safety-critical paths fail open, not closed. "Give them the coaching and log it" is the default for ANY error in any gate-handling, paywall-handling, or trial-handling chain. A trial user blocked from coaching due to a bug at the conversion moment is the single most expensive failure mode in the entire product — company-ending bug class. Coaching delivered + bug logged to high-severity telemetry = recoverable. Coaching blocked = lost user + invisible bug. The asymmetry is so steep that fail-open MUST be the default. Architecture: (1) backend never JUST rejects on internal error — returns a response shape the frontend can handle that includes the coaching content if anything in the gate-check itself errors; (2) frontend gate-handling has its own try/catch that on failure calls back to the backend to generate the coaching response anyway; (3) every error path in the chain terminates at "user gets coaching" with high-severity admin-diagnostics log entry. Block-on-bug is acceptable ONLY when the cost of false-allow exceeds the cost of false-block — for trial coaching the cost of false-block is catastrophic. Applies to gates, paywalls, auth checks, every "can the user do this thing" decision in the conversion path.


The 30-second summary

Open a real Chromium window on the person's machine via Chrome DevTools MCP. Start screencapture -v recording the screen. Drive the funnel anonymously via isolatedContext. Inject a glass overlay top-right that narrates every thinking/action/result. The person watches live and scrolls the overlay log at their own pace. On completion (or visible break), kill recording — .mov lands on Desktop.

Source-of-truth doc

Everything you need to run a walk — the launch command, the overlay-injection JS, the anti-patterns, the checkpoint format — is in this file. One setup additionally keeps a longer step-by-step recipe doc in the project's own notes vault; if this project has an equivalent doc (check <project>/docs/qa/browser-testing/ or ask the person), read it for extra detail, but nothing here depends on one existing.

The launch command

# 1. Start recording BEFORE opening the browser
TEST_NAME="<short-slug-for-this-test>"
RECORDING_PATH="$HOME/Desktop/${TEST_NAME}-$(date +%Y%m%d-%H%M%S).mov"
screencapture -v -k -D 1 "$RECORDING_PATH" &
RECORDING_PID=$!
echo "Recording PID: $RECORDING_PID → $RECORDING_PATH"

Then via Chrome DevTools MCP:

  1. mcp__chrome-devtools__new_page({ url, isolatedContext: TEST_NAME })
  2. Inject overlay (snippet in source-of-truth doc, Step 4)
  3. Snapshot → log thinking → action → log result, looped per step
  4. On done: kill -INT $RECORDING_PID

Default target — INSTINCT: localhost unless told otherwise

For any test of code in middlewares/, controllers/, route handlers, or any backend logic, the default target is http://localhost:3000 talking to <your-api-server> (or whatever ports the running dev stack uses — verify with curl -s -o /dev/null -w "%{http_code}\n").

Test against production ONLY when explicitly asked, or when the test is about deployed-experience verification (e.g. "does the ad-to-landing flow work on the real site").

Test accounts and login — institutional knowledge

Existing login skill (source of truth), if the project has one: check for a project-specific login/test-account skill first (one example is a project login skill) before improvising credentials.

Pre-existing test account (logged-in walks):

  • Email/password: whatever standing test account the project already has (check the project's own dev docs, or ask the person once and note it in your own private test notes — never write a real credential into this file).
  • Local: many projects have a /dev/auto-login route or equivalent to skip the form — check for one before assuming there isn't.
  • Production: use the real form.

New test account (signup-flow walks ONLY):

  • Format: the person's own email with plus-aliasing — <their-local-part>+<unix-timestamp-or-random-digits>@<their-domain>.
  • Why: plus-aliasing → emails actually arrive in an inbox the person controls, so they can verify what the system sends. Trivially cleanable via a regex query on the email field once you know their real local-part/domain.
  • NEVER use synthetic test domains like *@test.invalid, *@example.com, etc — they create noise and the person can't observe the emails.
  • Password convention: use the project's standard test password, or a memorable one noted in the test run — never a value copied from production.

Cleanup query shape (for the person, after testing — fill in their real local-part/domain, never write a real one into this file):

db.users.deleteMany({ email: /^<their-local-part>\+\d+@<their-domain>$/i })

How a single walk is structured

One user. One path. One watch. If the project has a persona library for simulating realistic users (one example: critique-from-target-audience's personas/ folder), drive the walk with one of those; otherwise role-play a specific, named user type yourself and stay in character. The persona/role reacts authentically each turn — no canned test strings. Hit the code-change-relevant moment (e.g. a gate or threshold the change affects), observe what happens, continue past it (reject the prompt, continue the flow, fill a form, try to circumvent it). Cohort matrix only enters the retrospective report, not the plan.

Anti-patterns (from prior failures — see source doc)

  • Do NOT use auto-browser / noVNC
  • Do NOT use Playwright --headed for human watching
  • Do NOT chain sleep commands to pace the walkthrough — harness blocks long sleeps. Plow through. Overlay log accumulates. The person scrolls.
  • Do NOT auto-login or pre-seed the session when testing new-user flow
  • Do NOT skip isolatedContext — cookie bleed corrupts the test
  • Do NOT use navigate_page with persistent connections (Stripe/Pusher) open — use evaluate_script + window.location.href + wait_for

🔧 Agent tools — continue-test + side-panel overlay

Resume from checkpoint

Checkpoint file shape (/tmp/watch-checkpoint-<slug>.json):

{
  "slug": "<short-slug-for-this-test>",
  "isoTimestamp": "2026-05-12T15:55:00Z",
  "url": "http://localhost:3000/dashboard/<session-id>",
  "sessionId": "<session-id>",
  "testAccount": "person@example.com",
  "lastUserMessageCount": 6,
  "lastCoachReplyPreview": "<first ~80 chars of the last real reply, for sanity-checking on resume>",
  "recordingPath": "~/Desktop/<slug>-walk.mov",
  "goal": "<the specific outcome this walk is driving toward, in one sentence>",
  "personaName": "<persona or role name, if using one>",
  "overlayMode": "top-right" | "side-panel",
  "cohort": "<whatever segment/cohort label this project uses, if relevant>"
}

Resume protocol:

  1. Read checkpoint JSON
  2. Clear stale Chrome MCP profile lock (NOT the whole profile — we want the auth cookies to survive)
  3. mcp__chrome-devtools__new_page({ url: checkpoint.url, isolatedContext: checkpoint.slug }) — the URL navigates directly to the in-progress coaching session
  4. Take snapshot — confirm coaching surface loaded with prior messages visible
  5. Re-inject overlay (same mode as checkpoint.overlayMode), log "RESUMED from checkpoint at msg N"
  6. Start new recording with -resume- in filename so it doesn't overwrite the original
  7. Continue driving the persona/role from message N+1

Checkpoint write convention: After every reply renders successfully, write/overwrite the checkpoint file via evaluate_script calling a dev endpoint if the project has one, or via the agent's own Write tool. Simpler: agent writes after every turn that gets a real reply.

Side-panel overlay (alternative to top-right)

When the person requests "side-panel overlay so I can see the full UI", use this injection instead of the default top-right glass panel. It modifies ONLY the injected DOM and a single inline style on document.body — zero changes to the app's own source files.

() => {
  if (window.__agentOverlay) return 'already';

  // Compensate body width so the app's UI shrinks to fit
  const SIDEBAR_WIDTH_PX = 460;
  document.body.style.width = `calc(100vw - ${SIDEBAR_WIDTH_PX}px)`;
  document.body.style.marginRight = `${SIDEBAR_WIDTH_PX}px`;
  document.body.style.boxSizing = 'border-box';
  document.body.dataset.agentSidePanelActive = '1';  // marker for cleanup

  const panel = document.createElement('div');
  panel.id = '__agent-overlay';
  panel.style.cssText = [
    'position:fixed','top:0','right:0',
    `width:${SIDEBAR_WIDTH_PX}px`,'height:100vh',
    'overflow-y:auto','z-index:2147483647','padding:18px 20px',
    'background:rgba(15,17,22,0.95)','color:#e8eef7',
    'border-left:1px solid rgba(255,255,255,0.12)',
    'font:13px/1.5 -apple-system,BlinkMacSystemFont,"SF Pro Text",sans-serif',
    'box-shadow:-12px 0 40px rgba(0,0,0,0.45)','pointer-events:auto',
    'box-sizing:border-box',
  ].join(';');
  panel.innerHTML = '<div style="font-weight:600;margin-bottom:12px;font-size:14px;opacity:.85">Agent log — side panel</div><div id="__agent-log"></div><button id="__agent-overlay-dismiss" style="position:absolute;top:12px;right:12px;background:rgba(255,255,255,0.08);border:0;color:#e8eef7;padding:4px 8px;border-radius:6px;font-size:11px;cursor:pointer">close</button>';
  document.documentElement.appendChild(panel);

  // Dismiss removes BOTH the panel AND the body style — the app's UI returns to full width
  document.getElementById('__agent-overlay-dismiss').onclick = () => {
    document.body.style.width = '';
    document.body.style.marginRight = '';
    document.body.style.boxSizing = '';
    delete document.body.dataset.agentSidePanelActive;
    panel.remove();
    window.__agentOverlay = false;
    delete window.__agentLog;
  };

  const ICONS = { thinking: '🧠', acting: '👉', pass: '✅', fail: '❌', info: 'ℹ️', finding: '🔍', persona: '👤', target: '🎯', resume: '⏯️' };
  window.__agentLog = (text, status = 'info') => {
    const log = document.getElementById('__agent-log');
    if (!log) return;
    const row = document.createElement('div');
    const t = new Date().toLocaleTimeString([], { hour12: false });
    row.style.cssText = 'padding:10px 0;border-top:1px solid rgba(255,255,255,0.07)';
    row.innerHTML = '<div style="opacity:.5;font-size:11px;margin-bottom:3px">' + t + '</div><div>' + (ICONS[status] || 'ℹ️') + ' ' + text + '</div>';
    log.prepend(row);
  };
  window.__agentOverlay = true;
  window.__agentLog('Side-panel overlay active. App UI compressed to leave 460px sidebar. Click "close" top-right to restore full width.', 'info');
  return 'side-panel installed';
}

Validation before relying on side-panel: Take a snapshot after injection, confirm the dashboard buttons (Start New Session, Login, intent buttons, send button) are still clickable and the textbox is still focusable. If any of those return errors, dismiss the overlay and fall back to top-right mode.

🗺️ Overlapping-skill map — when this canonical is the front door vs when a specialty skill applies

Browser/UX-testing skill libraries tend to accumulate overlapping entries over time — several skills that each do a slice of "test the app in a browser." This canonical exists to be the one front door for a watchable, narrated, recorded walk; other skills earn their keep only for a genuinely different surface. If you're about to reach for one of the table below, check first whether this canonical already covers it.

Mostly superseded by this canonical — use only for the specific reason listed

Skill (if installed) What it does When to still use it instead of this canonical
agentic-chrome-testing Chrome via AppleScript or Playwright Almost never — only if explicitly asked for AppleScript-driven Chrome without MCP.
devtools-site-testing Chrome DevTools MCP for sites Never separately — this canonical IS that workflow with overlay + recording + checkpoint + persona on top.
webapp-testing, playwright-e2e Playwright for local web apps / generic e2e patterns Only for headless CI regression suites — never for a watchable walk.
playwright-validator, agent-ux-verification, agentic-ux-verification Screenshot evidence collection Only for unattended UX-assignment artifact collection where no one is watching live.
autonomous-app-testing Self-healing autonomous tests Only if no human is in the loop at all and full autonomy is genuinely wanted.
validate-in-browser Reusable browser validation orchestrator Never separately — its decompose → preflight → run → cross-verify shape is folded into this canonical.
feature-completion-ux-test End-to-end check after a ship Rarely — this canonical handles end-to-end with checkpoint resume, which that skill doesn't have.

If a skill in this table isn't installed in your plugin set, that's fine — it just means there's nothing to defer to for that edge case; this canonical still covers the main path.

Adjacent — different surface, genuinely complementary

Adjacent skill What's different When to use it alongside this one
validate-against-live-data Validates code against real production data (queries, not a browser) When the question is "does the data match the code" — no browser needed
A project-specific login/test-account skill, if one exists Source of truth for login credentials Always, if the project has one
critique-from-target-audience (personas), if installed Persona library for realistic user simulation When you want a persona rather than role-playing one yourself
intent-lifecycle, intent-db, intent-validation-engine UX intent contracts and validation When you're working WITH intents, not driving a browser walk
verify-ux-assignment-state Per-assignment state verification When you specifically have UX assignments to verify
ux-testing-learnings UX testing learnings log When logging or searching past UX learnings
ux-assignment-reasoning-protocol Required reasoning before UX assignment creation When creating UX assignments, not walking
agent-outcome-summary End-of-task outcome reports When summarizing after a walk completes

If your own skill library grows overlapping entries

One pattern in use: when a new skill's job is fully covered by an existing canonical, move its SKILL.md out of the auto-loaded skills folder into a plain reference folder in the project (or an Obsidian vault, if configured) rather than deleting it — it stops consuming context-window budget on every session start, but stays one Read away for the rare case its specific legacy reason applies. This is a housekeeping technique, not something this skill requires you to do; do it only if your own skill count grows large enough to be confusing.


  • Followup audit: /discoverability-audit-learnings-findable-by-agents, if installed (run after first successful walkthrough)
  • Related skills (route here, not the other way around, when installed): /validate-in-browser, /autonomous-app-testing, /agentic-chrome-testing, /devtools-site-testing, /feature-completion-ux-test