← the whole session plugin/skills/ux-pre-execution-reflection/SKILL.md
Structured pre-execution reasoning that agents run BEFORE building, triggered when ux-criticality-measurement returns a score — or run directly any time you want to think through what could break before you build.
UX Pre-Execution Self-Reflection
When to Use
This skill is triggered by the ux-criticality-measurement skill when the criticality score is 4 or higher (moderate to high). Do NOT skip this step when triggered.
Trigger conditions:
ux-criticality-measurementreturnedreflectionRequired: true- You are about to build something with a criticality score >= 4
- You suspect a delivery chain issue but haven't mapped it yet
It also works stand-alone. If ux-criticality-measurement isn't installed, hasn't run, or you just want to think something through before building, invoke this directly. Say plainly that you're running it without a calibrating score, and default to full depth (all 9 questions, thorough reasoning) rather than skipping the exercise for lack of a number.
Depth:
reflectionLevel: "light"(score 4-6): Answer all 9 questions, brief answers acceptablereflectionLevel: "full"(score 7-10, or no score available): Answer all 9 questions with thorough reasoning and specific examples
The 9 Reflection Questions
Answer ALL of these before writing any implementation code.
1. What is the target intent in plain UX terms?
Format: "When {user} does {action}, they experience {outcome}"
This forces you to articulate what success looks like from the user's perspective, not from the code's perspective.
Good: "When a user submits this form, they see a confirmation and their data actually persists" Bad: "Save the form using the v2 handler with the retry wrapper"
2. What is the full delivery chain from code to user eyeballs?
Map every single layer between your code and the moment the user experiences the result. Include third-party systems, classifiers, and environmental factors.
Example:
Code -> Email template -> Transactional email provider -> Recipient's mail server -> Inbox classifier -> Inbox tab -> User opens email
3. For each layer: does this layer SERVE or FIGHT the intent?
Go through your delivery chain and honestly assess each layer. If ANY layer fights the intent, that is a red flag that must be resolved before building.
Worked example (a real incident, kept because it's a clean illustration — not something you need to have experienced yourself):
| Layer | Serves/Fights | Note |
|---|---|---|
| Content pipeline | Serves | Produces high-quality, personal-sounding content |
| HTML marketing template | FIGHTS | Wrapping a personal note in marketing chrome triggers the recipient's Promotions-tab classifier |
| Email transport | Neutral | Adds bulk headers by default |
| Inbox classifier | FIGHTS | HTML + multiple links + bulk headers = Promotions tab, not Primary inbox |
4. What could break fulfilling this intent?
List failure modes, especially silent ones where the code "works" but the user never experiences the intended outcome.
Example failure modes:
- Email lands in a Promotions/spam-like tab (silent — code reports "sent successfully")
- Auth token expires mid-flow (user sees a generic error, not a helpful recovery path)
- Feature flag misconfigured (feature silently hidden from target users)
5. What are the exact requirements to hit the nail on the head?
Not code requirements — UX requirements. What must be true about the user's experience?
Example:
- Must land in the primary inbox (not promotions, not spam)
- Must look like a human wrote it (no marketing chrome, no footer)
- Maximum 1 link in the body
- From name must be a real person's name, not the company name
6. In what ways could fulfilling this intent cause more harm than good?
Think about second-order effects. Sometimes a well-intentioned feature actively damages trust or creates negative associations.
Example: "A marketing-looking message from someone the user thinks of as a real person is WORSE than no message at all — it trains the user to ignore future messages and erodes the relationship the product is built on."
7. What assumption am I making that is most likely to be wrong?
Name it explicitly. The most dangerous assumptions are the ones you don't realize you're making.
Example: "I assumed 'sent successfully' equals 'reached the user'. The delivery layer between send and receive was invisible to me."
8. What question should I be asking myself that I haven't asked?
This is the meta-question — force yourself to find the blind spot.
Example: "Does every layer of my implementation serve the stated intent, or does any layer fight it?"
9. What is the answer to that question?
Answer your own meta-question honestly.
Example: "No — the HTML template actively fights the intent by making a personal note look like marketing. The template must be removed or replaced with plain text."
Output Format
The reflection produces this JSON object:
{
"targetIntent": "When {user} does {action}, they experience {outcome}",
"deliveryChain": [
{ "layer": "Content pipeline", "servesOrFights": "serves", "note": "Produces high-quality content" },
{ "layer": "HTML marketing template", "servesOrFights": "fights", "note": "Triggers spam/promotions classification" },
{ "layer": "Email transport", "servesOrFights": "neutral", "note": "Adds bulk headers" },
{ "layer": "Inbox classifier", "servesOrFights": "fights", "note": "HTML + multiple links + bulk headers = filtered" }
],
"failureModes": [
"Lands in a filtered tab",
"User never sees it",
"Personal voice undermined by marketing chrome"
],
"uxRequirements": [
"Must land in the primary inbox",
"Must look like a human wrote it",
"Max 1 link"
],
"harmPotential": "A marketing-looking message from 'a real person' is WORSE than no message — it trains the user to ignore future messages",
"riskyAssumption": "I assumed 'sent successfully' = 'reached the user'. The delivery layer was invisible to me.",
"metaQuestion": "Does every layer of my implementation serve the stated intent?",
"metaAnswer": "No — the HTML template actively fights the intent by making a personal note look like marketing"
}
How This Would Have Caught the Worked Example's Failure
The email-deliverability incident above is the canonical example this skill was designed to prevent:
- Criticality measurement would have scored it high (silent failure, sits in the delivery chain, first impression, harm if wrong all rated significant)
- Question 2 (delivery chain mapping) would have surfaced the marketing-template layer
- Question 3 (serves or fights) would have flagged that the HTML template fights the intent of a personal note
- Question 4 (failure modes) would have identified "lands in a filtered tab" as a silent failure
- Question 7 (risky assumption) would have caught the assumption that "sent = received"
Without this reflection, the agent jumps straight to code, produces good content, wraps it in a template that fights the goal, and the message lands somewhere the user never sees — defeating the entire purpose.
Saving the Reflection
Default (works with nothing set up): write the JSON object above to a file in the folder printed by alignment-harness records compactions — for example {session-id}-pre-execution-reflection.json. This is a plain local file; no database or API required. Print the reflection in the session either way, so it's visible even before it's saved.
If the person has their own session-tracking store or API for this (see /alignment-harness:harness-setup, and /agentic-session-compactions or /compact-agentic-session if the plugin ships them), write it there instead, using the same preExecutionReflection field shape shown above.
In one real build, this fed a database field that an internal admin page rendered as a delivery-chain visualization with serves/fights indicators. You don't need any of that infrastructure for the skill to work — the reasoning and the local JSON file are the whole point; a nicer viewer is a later, optional convenience.
Cross-References
- Criticality Measurement:
ux-criticality-measurementskill (triggers this reflection, optional — this skill also runs stand-alone) - Session Compactions:
agentic-session-compactions/compact-agentic-sessionskills, if installed (can store this output; otherwise use the local file described above)