← the whole session plugin/skills/agent-outcome-summary/SKILL.md
Generate a trustworthy, human-legible outcome report after completing dispatched work. Each item translates code changes into UX-testable statements with verification evidence grounding. The person you're working with reads this in 30 seconds and knows what happened.
Agent Outcome Summary
A "seed" here means a written work assignment — a short doc that states intent and what success looks like, handed to an agent (often working in its own isolated copy of the repo, a "worktree") to implement. This skill is the report that agent hands back.
Purpose
Produce a summary that a non-technical reader can read in under 60 seconds and know:
- What changed in the user's experience
- What evidence supports each claim (not self-assessment)
- What the agent could not do and why
Anti-Self-Assessment Rule (MANDATORY)
An agent CANNOT mark its own work as VERIFIED. The status is determined mechanically:
| Status | Condition |
|---|---|
| VERIFIED | Transcript contains actual command output (curl response, test result, screenshot path, DB query result) that directly proves this specific claim. The evidence must be quoted verbatim. |
| UNVERIFIED | The change was made but no external verification command was run against it. The agent believes it works based on reading the code — but reading is not running. |
| BLOCKED | The agent could not complete this item. Must include the specific reason and what would unblock it. |
What counts as evidence for VERIFIED:
- curl output showing the actual HTTP response body
- Test runner output: "Tests: 5 passed, 0 failed" with test names
- Screenshot file path from playwright/devtools capture
- Database query result showing the stored/changed data
- Server console output showing request processing
- Browser DOM inspection output
What does NOT count:
- "I read the code and it looks correct"
- "The pattern matches the working endpoint"
- "It should work based on my changes"
- Self-eval scores (85/100) without running anything
- "I verified it" without showing output
Summary Format
## OUTCOME SUMMARY
Seed: {seed name} | Implements-Seed: {seed-id}
Agent: {model name} | Session: {session identifier}
Duration: {approximate time}
### Seed Success Criteria
{For each "What success looks like" item from the seed:}
- [VERIFIED|UNVERIFIED|BLOCKED] {Criterion restated as testable UX statement}
Evidence: {verbatim command output, or "No verification run" for UNVERIFIED, or "Blocked: {reason}" for BLOCKED}
### UX Changes Delivered
{For each discrete change the agent made:}
- [VERIFIED|UNVERIFIED|BLOCKED] When {situation}, {what the user now experiences} because {what changed}.
Evidence: {verbatim excerpt from actual command output}
### What Was Not Done
{Explicit list of anything in scope that was not completed, with reasons.
If everything was completed, state: "All seed scope items addressed."}
### Verification Commands for Human
{List the exact commands the person can run RIGHT NOW to verify each claim:}
- {command 1} — proves {which claim}
- {command 2} — proves {which claim}
### Files Changed
{List of files created or modified, one per line, with one-sentence description of what changed in each}
How to Generate the Summary
Step 1: Collect seed success criteria
If the agent was dispatched from a seed document, read the "What success looks like" section. Each sentence becomes a criterion line in the summary. If no seed exists, skip this section.
Step 2: Scan transcript for verification evidence
Search backward through the conversation for actual command outputs:
curlcommands with response bodiesnpm test/npx jestwith test resultsnodescripts with output- Screenshot captures with file paths
- Database queries with results
- Server responses logged to console
For each piece of evidence found, note:
- The command that was run
- The output it produced
- Which UX claim it supports
Step 3: Map code changes to UX statements
For every file changed (via Edit or Write tool), translate the change into a UX statement:
- BAD: "Updated SettingsCard.jsx to add isExpandable prop"
- GOOD: "When a user opens Settings, sections now start collapsed with only title and subtitle visible — clicking expands to show full controls"
Step 4: Assign status mechanically
For each UX statement:
- Check if there is a matching evidence artifact from Step 2
- If YES and the evidence directly tests this specific claim → VERIFIED
- If YES but the evidence tests something adjacent (e.g., curl /health when you changed /settings) → UNVERIFIED
- If NO evidence exists → UNVERIFIED
- If the work was attempted but could not complete → BLOCKED with reason
Step 5: Write verification commands for human
For each UNVERIFIED item, write the exact command the person would run to verify it. For VERIFIED items, include the command that was already run so the person can re-run it.
Step 6: Output the summary
Write the summary to TWO locations:
- File in worktree:
.claude/outcome-summary.md— human-readable markdown - Local compaction record: if the person has set up a private compaction API of their own (see
/alignment-harness:harness-setup), post it there the way that setup describes. Otherwise, runalignment-harness records compactionsto get the local folder the harness keeps these in, and write the same summary there as a JSON or markdown file named for the session — this is the working default and needs no setup.
Integration Points
With /communicate-what-you-finished
This skill REPLACES the format from /communicate-what-you-finished for seed-dispatched work. The old skill's "when x then y because z but d" pattern is preserved in the "UX Changes Delivered" section but upgraded with evidence grounding.
With stop-evidence-gate.sh
The evidence gate already checks for verification output before allowing completion claims. This skill consumes the SAME evidence the gate checks for — the two are complementary. The gate enforces "you must verify." This skill structures "here's what you verified and what you didn't."
With Seed Documents
When a seed file path is known (e.g., <project>/docs/seeds/human-legible-outcome-summary.md), read its "What success looks like" section and create one criterion line per sentence.
With a Worktree Manifest
If a worktree manifest exists at .claude/worktree-manifest.json (a JSON file some setups use to track which agent is working in which worktree), the summary file path should be recorded in it under outcomeSummary. This allows whatever is coordinating the dispatched agents (a person, or an orchestrating agent) to find the summary without searching.
Example Output
## OUTCOME SUMMARY
Seed: Human-Legible Outcome Summary | Implements-Seed: human-legible-outcome-summary
Agent: claude-opus-4-1 | Session: agent-a50264c2
Duration: ~25 minutes
### Seed Success Criteria
- [VERIFIED] A summary format exists with VERIFIED/UNVERIFIED/BLOCKED status per item
Evidence: cat .claude/outcome-summary.md shows the format with all three status types
- [VERIFIED] Each VERIFIED item has corresponding command output or evidence
Evidence: grep "Evidence:" .claude/outcome-summary.md shows evidence lines on every VERIFIED item
- [UNVERIFIED] The summary is readable by a non-technical reader in under 60 seconds
Evidence: No verification run — requires human judgment, not a command
### UX Changes Delivered
- [VERIFIED] When an agent finishes dispatched work, it now produces a structured summary where each line says what changed in the user's experience, with evidence status (VERIFIED/UNVERIFIED/BLOCKED) that is determined mechanically from transcript evidence — not self-assessment.
Evidence: cat <plugin>/skills/agent-outcome-summary/SKILL.md | head -20 shows the skill file exists with the documented format
- [UNVERIFIED] When the person reads 5 agent summaries, each takes under 60 seconds to scan because the format leads with status badges and UX statements before evidence details.
Evidence: No verification run — requires human reading the actual output
### What Was Not Done
All seed scope items addressed.
### Verification Commands for Human
- cat .claude/outcome-summary.md — read the generated summary to verify format
- grep -c "VERIFIED\|UNVERIFIED\|BLOCKED" .claude/outcome-summary.md — count status-tagged items
### Files Changed
- <plugin>/skills/agent-outcome-summary/SKILL.md — new skill defining the summary format, generation mechanism, and anti-self-assessment rules
Generating the Summary File
At the end of your work, create the file:
# Write the summary to the worktree
cat > .claude/outcome-summary.md << 'SUMMARY_EOF'
{your generated summary here}
SUMMARY_EOF
Then write the local record (no setup required):
RECORDS_DIR=$(alignment-harness records compactions)
cat > "$RECORDS_DIR/outcome-summary-$SESSION_ID.json" << 'SUMMARY_EOF'
{
"sessionId": "$SESSION_ID",
"compactedSession": "outcome-summary",
"scopeDeclaration": {
"intent": "Generate human-legible outcome summary for seed work",
"outcomeSummary": "...the summary text..."
}
}
SUMMARY_EOF
Tell the person the path you wrote to. If they've configured their own private compaction API instead (/alignment-harness:harness-setup), post there according to that setup and skip the local file.
Triggers
Invoke this skill when:
- An agent finishes work dispatched from a seed document
/agent-outcome-summaryor/outcome-summaryis called- The stop-evidence-gate passes and the agent is composing its final report
- Any dispatched agent is about to report completion to the orchestrator