← the whole session plugin/skills/incomplete-tasks-get-from-session/SKILL.md
Capture unfinished tasks from the current Claude session transcript and persist them to the compaction system.
Capture Incomplete Tasks from Session
Before your output, print ## INCOMPLETE_TASKS on its own line. This enables automated extraction, if you've wired one up, to whatever surface you review compactions on.
Extracts incomplete work from the current session's transcript and persists it to the compaction system (see agentic-session-compactions) so future agents can pick it up.
The core principle: the transcript JSONL file is the source of truth. Not agent memory, not the compacted summary, not what the agent thinks happened. The actual file on disk.
Intelligence Routing
Step Model Dispatch
──────────────────────────────────────────────────────────────────────
Extract transcript from disk None (Python) Inline
Extract TaskList state None (TaskList tool) Inline
Clean + separate utterances HAIKU (data work only) Background
Identify incomplete work SONNET (reasoning) Background
Map to incompleteItem shape SONNET (reasoning) Background
Persist via /compact-agentic-session None (delegated) Inline
Verify persistence None (read-back) Inline
Model selection rationale (one real project corrected this split after seeing Haiku produce shallow results on the reasoning step):
- Haiku = data cleaning ONLY: separating transcript into distinct utterances, stripping system tags, preserving full content without summarizing or truncating. NO reasoning, NO interpretation.
- Sonnet = reasoning: identifying what's incomplete, mapping intent to actionable items, determining priority, assessing risks. This requires actual sense-making — a smaller/cheaper model produces shallow results here.
Use whatever your setup calls its "fast/cheap" tier for the cleaning step and its "reasoning" tier for the interpretation step — the specific model names don't matter, the division of labor does.
Step 1: Extract Verbatim Transcript (PROGRAMMATIC — no AI)
CRITICAL: The agent MUST use the Bash tool to execute this script. DO NOT reason about what it would return.
This function reads the session's JSONL file and returns the actual user messages — verbatim, not summarized, not interpreted.
This skill ships two files alongside this SKILL.md: run_extract.sh and extract_transcript.py. Run the copy that lives in this skill's own directory. If you're not sure of the absolute path, locate it once and reuse it for the rest of the session:
SCRIPT="$(find ~/.claude -path '*incomplete-tasks-get-from-session/run_extract.sh' 2>/dev/null | head -1)"
bash "$SCRIPT"
Output shape:
{
"session_uuid": "c06eab6d-...",
"transcript_path": "~/.claude/projects/.../.jsonl",
"total_user_messages": 18,
"user_messages_verbatim": ["see if you can create...", "ok got it...", ...],
"tasks_created": 7,
"tasks_completed": 5,
"incomplete_tasks": [
{"subject": "Task subject verbatim", "description": "Full description", "status": "pending"}
],
"tool_summary": {"Edit": 22, "Write": 3, "Bash": 45, "Agent": 16, "Skill": 5, "other": 12}
}
Step 2: Dispatch Sonnet (or your project's reasoning-tier model) to Find ALL Unaddressed Intent
Model: your project's reasoning tier (MANDATORY — this is reasoning, not data work)
The core job: find every piece of unaddressed intent in the transcript. Intent is NOT just explicit requests. Intent includes:
- Explicit requests: "build X", "add Y", "fix Z"
- Complaints as intent: "that tool is misaligned" = implicit intent to fix its alignment
- Corrections as intent: "no not that", "the style is wrong" = implicit intent to change approach
- Objections as intent: "this doesn't actually do the thing" = implicit intent to make it do the thing
- Aspirational framing as intent: "this is about surfacing alignment" = implies a set of requirements for what the system must do
- Frustrations as intent: "I keep having to check your work" = implicit intent for the system to self-verify
- Observations that imply gaps: "I don't see any place where I would articulate intent" = implicit intent that such a place should exist
- Questions that imply requirements: "how do we confirm that deterministically?" = implicit intent for deterministic confirmation to be built
Every single statement the person makes that describes a difference between what IS and what SHOULD BE — that is intent. The reasoning-tier agent must find ALL of them.
HOW the reasoning agent reasons about intent (MANDATORY — include in the prompt)
The prompt must teach the agent a DECISION PROCESS, not just a list of categories:
Step 1: Does this message describe a GAP? A gap is any difference between what IS (current reality) and what SHOULD BE (desired reality). The gap can be stated explicitly ("build X") or implicitly ("that tool is misaligned" — implying it should be aligned). Ask: "Is the person describing something about reality that they want to be different?" If yes → this contains intent.
Step 2: What KIND of gap is it? Use this decision table:
| Signal in the person's words | What it means | How to derive the intent |
|---|---|---|
| "Build X" / "Add Y" / "Create Z" | Explicit request | The intent IS the request |
| "That's broken" / "This is misaligned" | Complaint about current state | Intent = the FIXED state |
| "No not that" / "Wrong approach" | Correction of agent behavior | Intent = the CORRECTED approach |
| "This is about surfacing alignment" | Aspirational framing | Aspiration IMPLIES requirements |
| "I keep having to..." / "Why do I have to..." | Frustration with repeated problem | Intent = system handles it automatically |
| "I don't see any way to..." | Observation of absence | Intent = that thing EXISTING |
| "How do we confirm...?" | Question implying requirement | Intent = the answer being built |
| "That button is unclear" | UX feedback | Intent = clear UX |
| "I want this for every..." | Scope declaration | Intent = full scope realized |
| "Remember, it should..." | Constraint reminder | Intent = constraint honored |
Step 3: Is this intent already captured as a task? Check the existing task list. "Specifically" means the task description would lead an agent to fulfill THIS exact gap — not just a task in the same general area.
Step 4: Is the task actually COMPLETE with evidence? "Code exists" is NOT evidence. The intent is fulfilled when the PERSON EXPERIENCES the intended outcome.
The prompt MUST include a reasoning field
Every output item must have: "reasoning": "How I determined this is intent: [the specific gap between IS and SHOULD BE]" — this is how the agent shows its work and how downstream consumers can verify the reasoning wasn't hallucinated.
Reasoning-tier agent prompt (dispatch with run_in_background: true):
You are an intent extraction agent. You receive the FULL verbatim transcript of a Claude Code session between a person and an AI agent.
Your job: find EVERY piece of unaddressed intent in the transcript. Not just explicit requests — ALL intent, including implicit.
INTENT IS: any statement that describes a difference between what exists and what the person wants to exist. This includes:
- Direct requests ("build X", "add Y")
- Complaints ("that's broken", "this is misaligned") — the fixed state is the intent
- Corrections ("no not that", "never do X") — the corrected approach is the intent
- Objections ("this doesn't actually work") — working correctly is the intent
- Aspirational statements ("this is about surfacing alignment") — the aspiration implies specific requirements
- Frustrations ("I keep having to check") — the system self-checking is the intent
- Observations of absence ("I don't see any way to do X") — X existing is the intent
- Questions that imply requirements ("how do we confirm this?") — confirmation existing is the intent
- UX feedback ("that button is unclear") — clear button is the intent
- Scope declarations ("I want this to work for every conversation type") — full scope is the intent
For EACH intent you find:
1. Quote the EXACT person's words (never paraphrase)
2. Derive the implied intent — what reality would exist if this was addressed?
3. Check: does a corresponding task exist in the task list? Is it marked complete?
4. If no task exists OR the task exists but isn't verified as complete → this is UNADDRESSED
DO NOT summarize or compress the person's words. The full nuance must survive.
DO NOT skip statements because they seem minor. A frustration about a button is still intent.
DO NOT group multiple intents into one item. Each distinct intent gets its own entry.
Return JSON array:
[
{
"adminQuote": "the FULL exact quote — never truncated, never summarized",
"intentType": "explicit_request | complaint | correction | objection | aspiration | frustration | absence | question | ux_feedback | scope_declaration",
"intentTitle": "one sentence — what the person wants to exist, stated as the desired end state",
"intent": "the precise, context-bound UX the person wants — who experiences what, when, how — with full nuance preserved",
"hasMatchingTask": true/false,
"taskStatus": "not_started | in_progress | completed | no_task",
"evidenceOfCompletion": "what evidence exists that this was actually done — or 'none'",
"isVerifiedComplete": false,
"priority": 60,
"whats_missing": "specific gap between intent and current state",
"doneMeans": "1-3 sentences, in the PERSON'S OWN VOICE and words, stating the SEMANTIC meaning of completion — the lived outcome that tells them it is done, NOT the code-level fix. The test: could they read this one line and instantly feel what 'done' is, without parsing any code? Lead with 'Done means:' then describe the experience a real person has when this is complete, and WHY it matters. Reuse their exact phrasing from the transcript wherever they gave it. Example shape: 'Done means: a search crawler hits a page on the site and it loads, instead of getting blocked with a 403. The proof is a stranger from Google can actually reach the page.' NEVER write this as 'the X function returns Y' or 'the file is updated' — that is code meaning, which is exactly what this field exists to replace."
}
]
doneMeans is mandatory and is the single most important field for the person reviewing this list. They read this list to decide what to act on; a code-meaning description ("populate the vector store") forces them to translate it themselves, while a lived-outcome description ("you type a feeling into search and get the right practices back, ranked") lets them decide in one glance. Every item MUST have a doneMeans written in their voice.
The caller pre-loads the user_messages_verbatim and incomplete_tasks arrays from Step 1 directly into the reasoning-tier prompt. That agent does NOT read files — it receives the extracted data.
Step 2.5: Caller Prints to Screen (MANDATORY — before any persistence)
The first time this list is shown to the person, lead with the lived-outcome version. They don't want a schema dump — they want to read each item and instantly feel, in their own language, what "done" means and whether it matters. Print the items SORTED BY PRIORITY (highest first), each as a [priority] tag, the short title, then the doneMeans sentence(s) in their voice. This is the default presentation — produce it without being asked:
[{priority}] {intentTitle}
{doneMeans}
[{priority}] {intentTitle}
{doneMeans}
...
Keep it conversational and scannable — one item flows to the next like a person walking them through the list, not a table. Do NOT lead with adminQuote/evidence/risks fields; those are the audit detail, available on request, not their first read.
The full audit record (for the compaction + for an agent that needs the detail) still carries every field — print this expanded block only if the person asks for the detail, or include it silently in the persisted compaction content:
INCOMPLETE TASK {i}/{total}:
They said: "{adminQuote}"
IntentTitle: {intentTitle}
Done means: {doneMeans}
Intent: {intent}
Evidence: {evidence}
Status: {status}
Priority: {priority}/100
What's missing: {whats_missing}
Risks: {risks}
RiskAbatement: {riskAbatement}
This prints BEFORE any persistence step. The person reads it in their terminal first.
Step 2.75: Create TaskCreate Items (MANDATORY — DO NOT SKIP)
After printing and BEFORE compacting, create a TaskCreate item for EACH incomplete item the reasoning step identified. This is NOT optional. Compacting without TaskCreate means the items exist in the compaction system but NOT in the session's active task list — which means the agent has no tracked work items to act on.
For each item:
TaskCreate({
subject: "{intentTitle}",
description: "They said: \"{adminQuote}\"\n\n{intent}\n\nWhat's missing: {whats_missing}",
activeForm: "{short present-tense description}"
})
Why this step exists: in one real project, an agent once compacted 16 incomplete items but forgot to create TaskCreate entries — the task list didn't grow, and the person had to catch it manually. Compaction is for cross-session persistence. TaskCreate is for in-session tracking. BOTH must happen. Neither replaces the other.
Step 3: Delegate to Canonical Compaction Skill
This skill does NOT write to any compaction store directly. Persistence is delegated to /compact-agentic-session (if installed — otherwise follow agentic-session-compactions' own local-file default directly) — the single source of truth for all agent compaction writes. Fragmenting compaction logic across multiple skills risks writing into the wrong place, and that is exactly what happened previously in one real project (see below).
Past mistake — never repeat
If your project has a separate store for end-user-facing conversation data (support chats, coaching sessions, anything a real customer sees) as well as a store for agent work-session compactions, these must stay in two different collections/files. In one real project, an earlier version of this skill was mistakenly pointed at the user-facing conversation store instead of the agent-work store, and polluted real customer-facing data for about three and a half weeks before anyone noticed. Whatever this skill writes is agent work-session data — never write it into a store meant for end-user conversations.
Hand off to /compact-agentic-session
For the items identified in Step 2, invoke /compact-agentic-session (or, if you don't have that skill, write directly using agentic-session-compactions' own format) with this payload shape:
{
"sessionId": "{CLAUDE_SESSION_UUID}",
"title": "[Incomplete Tasks] {first item's intentTitle}",
"compactedSession": "Tasks captured at session end from transcript analysis. {plain-language summary ≥ 200 chars describing what the agent worked on, what got completed, and what remains}",
"tags": ["incomplete-tasks", "session-end", "auto-captured"],
"incompleteItems": [
{
"title": "{intentTitle}",
"content": "They said: \"{adminQuote}\"\n\nIntent: {intent}\n\nEvidence: {evidence}\nWhat is missing: {whats_missing}\nRisks: {risks}\nRisk Abatement: {riskAbatement}",
"intentStatement": "When {actor} does {action}, they experience {outcome}",
"status": "pending",
"priority": {priority}
}
]
}
If you have /compact-agentic-session installed, it handles the actual write, body validation, and retry logic on failure, and prints confirmation so the person sees proof. None of that should be duplicated here — if you don't have it, do the equivalent yourself using the local file format agentic-session-compactions describes, and still print the same confirmation.
Field rules (so the handoff payload is correct)
sessionId: the real Claude Code session UUID — NOT a custom slug, NOT auserId.compactedSession: minimum 200 characters. If you have less, expand the summary before handoff.incompleteItems[].statusenum:pending | approved | completed | dismissedincompleteItems[].priority: integer 1–100 (not a string)incompleteItems[].intentStatement: every item needs a "When X, then Y" UX statement
Why delegation matters
- Single source of truth: when the compaction format changes, only one place needs updating
- No fragmentation: agent work goes to one store, not scattered across multiple
- No user-data pollution: keeping agent work and end-user data in separate stores prevents the mistake described above
- Failure handling: retry logic for connection errors belongs in one place, not duplicated here
Verify (read-only)
After persistence, confirm it actually landed:
Local store (default): re-read the file you just wrote (cat "$(alignment-harness records compactions)/{SESSION_UUID}.md") and check the incompleteItems you expect are present.
Hosted store, if you built one: fetch it back with whatever read endpoint your own integration exposes and confirm the same items are present. Either way, print how many items you verified and their statuses — never claim persistence without actually reading it back.
Rules
- The JSONL file on disk is the source of truth. Not agent memory, not compacted summaries.
- User messages are extracted programmatically — the Python script reads them, not an agent.
- The reasoning-tier agent receives pre-loaded data — it does NOT read files or hit APIs. The caller feeds it the extracted transcript.
- Never claim tasks were persisted without reading them back. If verification fails, say so.
- Never summarize user messages. The
adminQuotefield must be a direct quote. If you can't quote it, you don't have it. - Merge, don't overwrite. If a compaction already has incompleteItems, fetch them first and include them in the write.
After Persisting
Print, in whatever form fits your setup:
PERSISTED: {path to the local compaction file, or a link if you have an admin UI wired to one}
Review, approve, or dismiss each item there.
If persistence failed:
NOT PERSISTED — {what went wrong}. Items printed above but not saved.
- Incomplete items should be reviewable per-compaction, however your project surfaces that (a file you can open, or an admin page if you built one)
- The person approves/dismisses/prioritizes from there
- Approved items become consumable by future agents (via
/agentic-actionable-tasks, if installed)
Schema Note
The incompleteItems in the compaction carry richer fields inside the content string:
- The person's verbatim quote
- Full intent statement (not summarized)
- Evidence of what was done
- What's still missing
- Risks and risk abatement (if applicable)
If you build a UI over this data and want to render these as separate fields instead of a single content blob, that's a parsing change on your UI's side — the underlying record already supports it via the existing content field; no format change needed here.