← the whole session plugin/skills/extract-wip-from-past-sessions/SKILL.md

Sweep a range of past Claude Code sessions with a fleet of Sonnet agents (one per session) to surface implied-but-unfinished work the person wanted done, filter to a target initiative you name, and convert each finding into a priority-prefixed, self-contained proposal sorted into work-state triage folders. Use when the person wants to mine recent or older sessions for dropped actionables before a push, or says "extract WIP", "sweep the last sessions", "what did I leave unfinished", or "run the WIP sweep on the next N sessions".

Extract WIP From Past Sessions

Why this skill exists

Across many sessions, a person implies work they want done — sometimes explicitly, sometimes by frustration or "this should…" — and it gets dropped when the session ends or compacts. This skill recovers that dropped work systematically: one agent per session reads the transcript, decides if there's implied-unfinished work tied to the target initiative, and if so writes a proposal the person understands on first read and another agent can execute end-to-end. This is an on-demand tool for a deliberate push, not a background sweep — run it when you're about to make a push and want to know what's been dropped along the way, not as something that runs unattended on a schedule.

Validated by the author on a 90-session sweep targeting a real front-end launch: 88 reviewed, 31 relevant, 41 proposals — the early-pulse finding held all the way to completion. That was his own product's launch; the mechanism itself doesn't care what the target is.

Before you start: name the target initiative

"Relevant" has no meaning without a target. Ask the person, up front, what initiative, launch, or theme to filter for (a feature launch, a bug-fix push, a specific area of the product) — never silently sweep everything and call the result targeted. If they don't have an Obsidian vault or equivalent triage system already, see "Output location" below.

Fleet model

Sonnet, always (never Haiku — this is judgment work: deciding what's implied, what's relevant, and writing a proposal someone else can act on). Escalate a specific item to Opus only when root-cause reasoning or live proof is required — that's the evidence-folder's job, not the sweep's.

The five requirements for each finding (keep this precise)

Each per-session agent must:

  1. Determine if there is incomplete work the person strongly implied they wanted completed.
  2. If so, determine whether it is directly implicated in the target initiative you named above.
  3. If so, write proposals immediately actionable, with a priority 1-99 as the filename prefix (NN_specific-completed-ux-outcome.md), in separate files.
  4. Make each proposal clear enough that another agent can execute it end-to-end and know where to get the context (repo, likely files, runnable verify recipe, source session id).
  5. Make each proposal clear enough that the person understands it the first time they read it — lead with "what the person will experience when this is done," zero jargon.

Workflow (orchestrator = parent thread, never delegated)

Step 1 — Ground reality first (do NOT skip)

  • Resolve the session window: build a newest-first manifest of *.jsonl in ~/.claude/projects/<project-slug>/, ranked by mtime. "Last 48h" or "next 100" resolves to a concrete rank range.
    DIR=~/.claude/projects/<your-project-slug>
    find "$DIR" -maxdepth 1 -name "*.jsonl" | while read -r f; do node -e 'const s=require("fs").statSync(process.argv[1]);console.log(Math.floor(s.mtimeMs/1000)+"|"+s.size+"|"+process.argv[1])' "$f"; done \
      | sort -rn | awk -F'|' 'NR>=START && NR<=END {printf "%d|%d|%s\n", NR, $2, $3}'
    
    (Works on macOS and Linux alike: find -printf and stat flags differ between them, so Node reads the time and size.)
  • Confirm the output location exists (see below) / create the dated sweep folder.
  • If a memory-search tool is set up (/alignment-harness:harness-setup), run it once to ground the relevance filter — what "the target initiative" concretely means (which surfaces, bugs, work). Otherwise, ground the filter by reading the project's own docs/README and recent git log for the same keywords, and say plainly that no memory-search tool was configured.

Step 2 — Fan out, newest-first, with an in-agent relevance gate

  • One Sonnet agent per session, ordered newest→oldest.
  • Each agent's FIRST move is a cheap relevance check (read the transcript head/tail): is this session related to the target initiative? If not → return relevant:false with a one-line reason and STOP (cheap skip). This is what keeps a 100+ session sweep affordable — most agents skip in seconds.
  • Relevant agents extract implied-unfinished work and write proposals per the five requirements.
  • Use a Workflow with pipeline() (no barrier) so a relevant session flows straight into proposal-writing while others are still being reviewed. Pin model: 'sonnet' on the agent calls. Return a structured verdict per agent (sessionId, sessionTitle, relevant, reason, proposalsWritten).

Step 3 — Output location and triage folders

If the person has a vault (Obsidian or similar) configured for proposals, write there, sorted into these work-state subfolders (each carrying an agents.md contract naming the skill sequence and owning tier — Opus or Sonnet, never Haiku):

  • 1_1 next action - person's own — the person's own; don't auto-execute
  • 1_1_get me a proposal — needs a full proposal authored (/task-get-one-important → /align → speak-human translation → /jonathan-check2 or your own response-predictor → /how-to-submit-and-track-proposals)
  • 1_1_2 Get me evidence — suspected, not proven; reproduce live + attach a clickable repro/screenshot before fixing (Opus)
  • 1_2_approved for escalation — escalate, don't act
  • 1_3_approved for action — execute + commit (/init → /resolve-proposals → /governer → whatever UI-pattern convention you use, if any → real verification → commit with a Resolves-Proposal: tag)
  • 1_9_agent-completed — archive
  • 1_99_deferred for after launch — leave

If there's no vault configured, write to a plain, clearly dated folder — alignment-harness records wip-sweep prints where — using the same triage-folder shape inside it, and tell the person where it landed. The six-folder taxonomy is a reasonable default either way; the labels are yours to rename.

Step 4 — Tier rule

Owning agent tier per folder/item is Opus or Sonnet, chosen by gravity / risk-of-error. Haiku is never part of this work. Opus for judgment / root-cause / live proof / high blast-radius; Sonnet for standard review and execution.

Step 5 — Initiative flags

When the person names a related concern beyond the primary target, instruct every agent to explicitly flag any session touching it, so that WIP surfaces even if it's not the primary filter.

Skill-pattern grounding

For per-folder skill sequences, if you have an oracle set up for "which skills does this kind of work need" (see /alignment-harness:harness-setup), query it. Otherwise reason from the sequences above and from what's actually installed — check /predict-required-skills2 if you have it; it's fine to fall back to a plainer sequence if no oracle is configured.

Known gaps (carry forward, don't re-discover)

  • A stamping/attribution step at the end of each reviewed transcript did not reliably land on the author's first run — if you add a similar "mark this session reviewed" step, verify it explicitly after the run (grep transcripts) rather than assuming it fired.
  • Structured-output loss ~2% — a few agents finish without returning the verdict schema after nudges. They fail loud (named in workflow <failures>); re-run just those sessions if they matter.
  • Untracked output — a sweep folder outside your normal doc tree may be untracked by git; use mv rather than git mv if that's the case (reversible since nothing's in history), then git add it if you want the proposals version-controlled.

When to re-run

This is a deliberate tool, not a scheduled job: run it on any session-count range not yet extracted, before a push where you want dropped actionables surfaced ahead of time.