← the whole session plugin/skills/skill-extractor/SKILL.md

Scans markdown files across the codebase and converts documentation into searchable agent skills.

BEFORE YOU START

BEFORE starting the creation of this skill, NAME the file you will read, Read this file end to end with no exceptions, Abide by it, THEN create the skill FROM it as the Seed.

The file is the /skill-file-create-pre-check-self-healing-protocol skill — it ships alongside this one in the same plugin. Read its SKILL.md in full (invoke it via the Skill tool, or open the file directly if you know where the plugin is installed).

Skill Extractor - Mine Skills from Markdown Files

extract-skills

Automatically discover and extract skills from new or updated .md files across the codebase. Converts documentation into searchable, actionable agent knowledge.

When to Use

  • After adding new documentation files
  • During periodic knowledge base refresh
  • When onboarding a new repo or system
  • When asked "what new skills exist?"

Quick Start

# Find new .md files (last 7 days)
/extract-skills --since 7d

# Extract from specific directory
/extract-skills --path web-app/docs

# Full scan (all .md files)
/extract-skills --full

What It Extracts

For each .md file, the extractor looks for:

Pattern Becomes Example
# How to X Skill: "how-to-x" "How to Deploy to Vercel"
## When to Use Trigger patterns "Use when deploying..."
Code blocks with commands Executable steps npm run deploy
## Examples Usage examples Real-world patterns
YAML frontmatter Metadata name, description, triggers

Extraction Algorithm

#!/bin/bash
# skill-extractor.sh

SINCE="${1:-7d}"
# Where new skill files get written. Claude Code's own per-user skills folder
# is the portable default; use a project-local .claude/skills if you'd rather
# keep extracted skills inside the repo they were mined from.
SKILL_DIR="${SKILL_DIR:-$HOME/.claude/skills}"

echo "🔍 Scanning for new .md files (since $SINCE)..."

# Find recent .md files
git log --since="$SINCE" --name-only --pretty=format: -- "*.md" | \
  sort -u | grep -v '^$' | while read file; do

  if [ -f "$file" ]; then
    echo ""
    echo "📄 Analyzing: $file"

    # Extract title
    TITLE=$(head -20 "$file" | grep -m1 "^# " | sed 's/^# //')

    # Check for skill indicators
    HAS_HOWTO=$(grep -c "^## How to\|^# How to" "$file" 2>/dev/null || echo 0)
    HAS_WHEN=$(grep -c "^## When to Use" "$file" 2>/dev/null || echo 0)
    HAS_COMMANDS=$(grep -c '```bash\|```shell' "$file" 2>/dev/null || echo 0)

    # Score skill potential
    SCORE=$((HAS_HOWTO * 3 + HAS_WHEN * 2 + HAS_COMMANDS))

    if [ "$SCORE" -gt 3 ]; then
      echo "   ✅ High skill potential (score: $SCORE)"
      echo "   Title: $TITLE"
      echo "   Has 'When to Use': $([ $HAS_WHEN -gt 0 ] && echo 'Yes' || echo 'No')"
      echo "   Command blocks: $HAS_COMMANDS"
    else
      echo "   ⚪ Low skill potential (score: $SCORE)"
    fi
  fi
done

Output Format

Proposed Skills Report

🔍 Skill Extraction Report
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

📄 web-app/docs/DEPLOYMENT.md
   ✅ High skill potential (score: 8)

   Proposed Skill:
   - Name: deployment-guide
   - Triggers: "deploy", "vercel", "production"
   - Key Commands:
     - `vercel --prod`
     - `npm run build`
   - When to Use: "When deploying to production..."

   Action: Create skill file? [Y/n]

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

📄 api/docs/WEBHOOKS.md
   ✅ High skill potential (score: 6)

   Proposed Skill:
   - Name: webhook-handling
   - Triggers: "webhook", "stripe event", "callback"
   ...

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Summary: 3 potential skills found, 12 files scanned

Skill Template Generation

When a skill is confirmed, generate:

---
name: {extracted-name}
description: Scans markdown files across the codebase and converts documentation into searchable agent skills.
---

# {Title from H1}

{Content extracted from source}

## When to Use

{Extracted from "When to Use" section or inferred from context}

## Quick Start

{Extracted code blocks}

## Source

Originally extracted from: {source_file_path}
Extraction date: {date}

A skill file dropped into Claude Code's own skills folder becomes usable immediately (via the Skill tool, or a matching trigger phrase). If you also have an institutional-memory search set up (agent_find or your own equivalent — see /alignment-harness:harness-setup), it should be included in future queries once that search re-indexes:

  1. Skill file created in ~/.claude/skills/{name}/SKILL.md (or your project's .claude/skills/{name}/SKILL.md)
  2. Your institutional-memory search picks it up on its next re-index, if you have one configured
  3. If you also keep a separate tool registry (see /how-to-bulk-register-atomic-tools), register it there too so it surfaces in tool search as well as skill search

Directories to Scan

Priority order for extraction:

1. ~/.claude/skills/**/*.md         # Existing skills (check for updates)
2. */docs/**/*.md                   # Documentation folders
3. **/README.md                     # Project readmes
4. **/CONTRIBUTING.md               # Contribution guides
5. **/ARCHITECTURE.md               # System architecture
6. **/*.md (remaining)              # Everything else

(Add your own extra locations here if you keep documentation somewhere else — for example a separate internal-tools repo or wiki export.)


CLI Commands

# Recent files only
git log --since="7 days ago" --name-only --pretty=format: -- "*.md" | sort -u | grep -v '^$'

# Count skill indicators in a file
grep -c "^## When to Use\|^## How to\|^## Quick Start" path/to/file.md

# Extract frontmatter
head -20 path/to/file.md | sed -n '/^---$/,/^---$/p'

# Find files with high skill potential
find . -name "*.md" -exec grep -l "## When to Use" {} \;

Automation

Scheduled Extraction (Weekly)

# Add to crontab or launchd
0 9 * * 1 ~/.claude/skills/skill-extractor/extract.sh --since 7d >> /tmp/skill-extraction.log   # save the algorithm above as extract.sh first

Git Hook (On Commit)

# .git/hooks/post-commit
#!/bin/bash
CHANGED_MD=$(git diff --name-only HEAD~1 -- "*.md")
if [ -n "$CHANGED_MD" ]; then
  echo "New/updated .md files detected. Run /extract-skills to check for new skills."
fi

Best Practices

  1. Don't over-extract - Not every .md file is a skill
  2. Merge related docs - Combine small related files into one skill
  3. Preserve source - Always link back to original documentation
  4. Human review - Confirm before creating skill files
  5. Dedupe first - Check if a skill already exists (your institutional-memory search if you have one, otherwise ls ~/.claude/skills/ and a grep over skill descriptions)