← the whole session plugin/skills/workflow-factory/SKILL.md
Fully autonomous workflow production system. Takes any objective + confirmed seed, produces a verified self-healing workflow document. Uses institutional-memory search and, if configured, reaction-prediction oracles for autonomous alignment. Self-heals on failure — patches its own stage designs when auditors catch errors. The factory that builds workflows IS a self-healing workflow.
Workflow Factory — Autonomous Workflow Production System
What This Is
A fully autonomous system that takes ANY objective and produces a verified, self-healing workflow document. No human hand-holding required. The factory uses institutional knowledge (memory search, and an institutional-principles oracle if you've configured one), a reaction-prediction check (predicting how the person you're working with would react — /alignment-harness:jonathan-check2, generalized, if configured), and self-reflection (/alignment-harness:reflect) to maintain alignment without human intervention.
The factory IS a self-healing workflow. Every stage has an agent producing output and another agent auditing that output. Failures get caught, fixed, AND traced back to the factory's own design — which gets patched so the failure can't recur.
This is an advanced piece — reach for it once you have a confirmed seed (from /alignment-harness:align) and, ideally, institutional-memory search set up (/alignment-harness:harness-setup, checkpoint 4). It still runs without either, with honest fallbacks noted at each phase below — it just does less checking.
When to Use
- You have a confirmed sub-seed and need a workflow to execute it
- You need to produce workflows for multiple domains in parallel
- You want to throw an objective at a system and get a production-ready workflow back
- The domain is unfamiliar and you need the factory to LEARN what failure modes apply
When NOT to Use
- Single-agent tasks (use
/alignment-harness:governerdirectly) - The workflow already exists (use it, don't rebuild it)
- You need human alignment first (use
/alignment-harness:align, THEN feed the confirmed seed to this factory)
The Factory Pipeline
The factory has 6 phases. Each phase has a PRODUCER and a CHECKER. The checker audits the producer's output before the next phase begins. If the checker finds a problem, the producer revises AND the factory logs a self-healing patch.
Phase 1: FIND & UNDERSTAND
Producer: Seed Finder
Checker: Alignment Verifier
Phase 2: REFLECT & CLASSIFY
Producer: Work Classifier
Checker: Classification Auditor
Phase 3: LEARN FAILURE PATTERNS
Producer: Pattern Researcher
Checker: Pattern Verifier
Phase 4: DESIGN STAGES
Producer: Workflow Designer
Checker: Design Auditor
Phase 5: VALIDATE WITH A REACTION-PREDICTION CHECK
Producer: (Phase 4 output)
Checker: reaction predictor (if configured)
Phase 6: PACKAGE & SAVE
Producer: Document Writer
Checker: Scope Guardian
Phase 1: FIND & UNDERSTAND
What happens: The factory receives an objective (plain language). It finds the relevant seed, reads it, and states its understanding.
Producer: Seed Finder
Search for the seed. If institutional-memory search is configured (
/alignment-harness:harness-setup, checkpoint 4), use it with the objective keywords:node <plugin>/scripts/memory-search.js "{objective keywords}"Otherwise, search the project's own seed/plan folders and
git logdirectly for the objective's keywords, and say which method you used.If no seed exists: The factory STOPS and reports: "No seed found for this objective. Run /alignment-harness:align first to create one." The factory does NOT create seeds —
/alignment-harness:aligndoes. This separation prevents the factory from hallucinating intent.If seed exists: Read it. Extract:
- The objective in UX terms (what does the human get?)
- The full scope (every question that must be answered)
- The success metric (how do we know it's done?)
- Any stated constraints or risks
State understanding: Print:
FACTORY — Phase 1: FIND & UNDERSTAND Seed found: {path} Objective: {one sentence} Scope items: {count} Success metric: {from seed}
Checker: Alignment Verifier
The checker reads the seed independently and verifies the producer's extraction:
- Did the producer capture ALL scope items? (Count them in the seed, count them in the extraction)
- Is the objective stated in UX terms (what the human gets), not agent terms (what agents do)?
- Is the success metric specific enough to verify?
If gap found: Producer re-reads the seed and revises. Log: SELF-HEAL: Phase 1 — producer missed {N} scope items. Patching extraction template to enumerate scope items explicitly.
Phase 2: REFLECT & CLASSIFY
What happens: The factory reflects on the objective to understand what TYPE of work this is. The type determines which failure modes are most dangerous and which stages are most important.
Producer: Work Classifier
Run /alignment-harness:reflect on the objective. Specifically answer:
What type of work is this?
- RESEARCH — gathering and verifying facts from external sources
- ACTION — producing a checklist/plan that a human executes immediately
- DESIGN — creative work that produces a system, journey, or experience
- SYNTHESIS — combining outputs from other teams into a unified deliverable
- IMPLEMENTATION — writing code or building something technical
- DECISION — analyzing options and recommending one path
What is the primary constraint?
- THOROUGHNESS — must be exhaustive and verified (regulatory research)
- SPEED — must be done fast, lean stages (first sale sprint)
- CREATIVITY — must explore possibilities, persona feedback (CX design, brand)
- ACCURACY — numbers must be real, no assumptions (financial model)
- COMPLETENESS — must cover every item, no gaps (competitive intelligence)
What is the risk profile?
- HIGH STAKES — wrong output leads to legal, financial, or reputational harm
- MEDIUM STAKES — wrong output wastes time but is recoverable
- LOW STAKES — wrong output is easily corrected
What existing workflows or patterns are similar? If institutional-memory search is configured, search for prior workflows:
node <plugin>/scripts/memory-search.js "workflow {work type} {domain keywords}"Otherwise, check the project's own workflow-output folder (see Phase 6) for anything similar.
Print:
FACTORY — Phase 2: REFLECT & CLASSIFY
Work type: {type}
Primary constraint: {constraint}
Risk profile: {level}
Similar workflows found: {list or "none"}
Checker: Classification Auditor
If you have an institutional-principles oracle configured (see /alignment-harness:insight / /alignment-harness:oracles, and /alignment-harness:harness-setup checkpoint 5), query it to verify the classification:
<your configured oracle-query command> "What principles apply to {work type} work in the domain of {domain}? What failure patterns are most common?"
Cross-check: does the institutional knowledge agree with the classification? If it says "regulatory research has confirmation bias as the #1 risk" and the producer classified it as THOROUGHNESS-constrained — that's consistent. If the producer classified a fast, low-stakes sprint as THOROUGHNESS-constrained — that's a misclassification.
If no oracle is configured, the checker instead re-derives the classification independently from the objective and the seed alone, and says plainly that no institutional-principles oracle was consulted.
If misclassified: Producer re-reflects with whatever evidence surfaced. Log: SELF-HEAL: Phase 2 — misclassified {objective} as {wrong type}. Evidence: {source}. Reclassified as {correct type}. Patching classifier to weight {signal} higher for {domain pattern}.
Phase 3: LEARN FAILURE PATTERNS
What happens: The factory researches what specifically goes wrong with THIS type of work. Not generic failure modes — specific ones informed by institutional knowledge where available.
Producer: Pattern Researcher
Up to three research tracks, depending on what's configured:
Track A — Institutional memory (if configured):
node <plugin>/scripts/memory-search.js "failure hallucination error {domain keywords}"
Extract: What has gone wrong before with similar work? What patterns were caught? What corrections has the person made in the past?
Track B — Institutional-principles oracle (if configured):
<your configured oracle-query command> "What are the most dangerous failure modes when agents do {work type} work about {domain}? What principles prevent them?"
Track C — /alignment-harness:workflow-design Case Studies (always available):
Read the 5 canonical failure modes from that skill's own root-cause analysis:
- STALENESS — proposing things that already exist or using outdated data
- WRONG EVIDENCE — citing wrong source or misreading data
- CONTEXT GAP — missing a global mechanism that invalidates the premise
- CONFIRMATION BIAS — finding partial evidence and inflating it
- SCOPE MISFRAME — solving the wrong problem
For each: is this failure mode RELEVANT to the current objective? Rate HIGH/MEDIUM/LOW/NOT APPLICABLE.
If neither Track A nor Track B is configured, always include all 5 canonical failure modes from Track C as your baseline, and say plainly that domain-specific research wasn't available.
Derive domain-specific failure modes: Combine whichever tracks ran. Produce 3-7 failure modes specific to THIS objective. Each must have:
- A name
- What it looks like in this domain (specific example)
- Why it's dangerous (what happens if it's not caught)
Print:
FACTORY — Phase 3: LEARN FAILURE PATTERNS
Failure modes identified: {count}
1. {name} — {one sentence} [source: {which track surfaced it}]
2. ...
Checker: Pattern Verifier
For each failure mode, the checker asks: "Is this real or is the producer inventing problems?"
Verification method:
- If sourced from institutional memory (Track A): verify the search result actually says this
- If sourced from the principles oracle (Track B): verify the response actually says this
- If sourced from case studies (Track C): verify the relevance rating makes sense for this domain
If a failure mode is fabricated: Remove it. Log: SELF-HEAL: Phase 3 — producer fabricated failure mode "{name}" without source evidence. Removed. Patching researcher to require source citation for every failure mode.
If a real failure mode is MISSING: Add it. Log: SELF-HEAL: Phase 3 — checker found failure mode "{name}" that producer missed. Added. Patching researcher to always check this source for domain-specific risks.
Phase 4: DESIGN STAGES
What happens: The factory designs the workflow stages. Each failure mode maps to a stage that makes it structurally impossible.
Producer: Workflow Designer
Using /alignment-harness:workflow-design's stage library, select and configure stages:
Stage Selection Rules (based on Phase 2 classification):
| Work Type | Required Stages | Optional Stages | Skip |
|---|---|---|---|
| RESEARCH | Align, Research, Verify, Cross-Verify, Scope Guardian, Consumption Handoff | Adversarial, Reaction-check | UX Verify, Decompose |
| ACTION | Align, Blocker Scan, End-to-End Journey Check, Scope Guardian | Minimum Legal Check | Adversarial, Cross-Verify |
| DESIGN | Align, Research, Persona Feedback, Adversarial, Reaction-check | UX Verify | Heavy Verification |
| SYNTHESIS | Align, Dependency Check, Integration, Scope Guardian, Reaction-check | Cross-Verify | Primary Research |
| IMPLEMENTATION | Align, Research, Build, Test, UX Verify, Scope Guardian | Code Review | Adversarial |
| DECISION | Align, Research (both options), Verify, Adversarial, Reaction-check | Cross-Verify | UX Verify |
For each failure mode from Phase 3: Map to a stage that prevents it. If no existing stage covers it, design a custom stage.
For each stage, define:
- Input: what this stage receives from the prior stage
- Output: what this stage produces
- Evidence requirements: what PROOF must be present (not claims — artifacts)
- Failure conditions: when to reject and route back
- Agent role: who executes this stage
Assign agent roles following /alignment-harness:workflow-design principles:
- Orchestrator: dispatches, NEVER does work
- Specialized agents: one per stage concern
- Lightweight checks: format/URL validation, cheap mechanical checks
- Use whichever model tier your setup routes for each kind of work — a fast/cheap model for mechanical checks, a stronger model for design and judgment calls. There's no fixed mapping baked in here; match effort to how much judgment a stage actually requires.
- 1 item per agent for verification stages
- Sequential chaining for dependent stages, parallel for independent items
Three quality layers (ALL required):
- Intrinsic design — stages structured so correct behavior is easiest
- Compliance audit — dedicated agent checks all stage outputs
- Self-healing loop — compliance patches stage designs on failure
Print the full workflow design.
Checker: Design Auditor
The auditor performs THREE checks:
Check 1 — Seed Coverage: For every scope item in the sub-seed: does a stage answer it?
Scope item: "State-by-state regulatory map"
Mapped to: Stage 3 (Research), Track B
Status: COVERED
If any scope item is NOT covered → FAIL.
Check 2 — Failure Mode Coverage: For every failure mode from Phase 3: does a stage prevent it?
Failure mode: "Staleness"
Prevented by: Stage 4 (Verify) — requires dated, current sources
Status: COVERED
If any failure mode is NOT prevented → FAIL.
Check 3 — Principle Compliance:
Check against /alignment-harness:workflow-design's principles:
- Feedback interpreted as additive by default
- Nuance preserved, not compressed to binaries
- Orchestrator never does work
- Each stage shows evidence of compliance
- All three quality layers present
- UX verification via the actual system when relevant
- Consumption handoff verifies everything human touches
- Decompose comes after verify
- Per-item enforcement from scoring dimensions
- Self-healing loop active
- Pipeline can feed itself
- The reaction-check (if configured) resolves ambiguity before escalating to a human
If any check FAILS:
- Report the specific gap
- Designer revises the workflow
- Log:
SELF-HEAL: Phase 4 — design missing {check}. Patching stage selection rules to always include {stage} for {work type} workflows.
Phase 5: VALIDATE WITH A REACTION-PREDICTION CHECK
What happens: The complete workflow design gets passed through a reaction-prediction check — predicting how the person you're working with would react to it, before they've seen it. If you have one configured, this is the autonomous alignment gate. If you don't, this phase says so honestly and the workflow is saved marked as not checked this way — never marked "passed" when nothing actually ran.
How to Query (if configured)
Use whatever reaction predictor you've set up — /alignment-harness:jonathan-check2 generalized to the person you're actually working with, or /alignment-harness:smc-prediction-server's local reaction predictor. Ask it something shaped like:
TASK: Design a self-healing workflow for {objective}
WORKFLOW PRODUCED: {paste workflow summary — stages, failure modes, evidence requirements}
SEED IT FULFILLS: {paste seed scope}
Based on this person's own history:
1. What would they say first?
2. What would they challenge or find insufficient?
3. What would they want to see that's missing?
4. Would they approve this for autonomous agent execution?
Interpreting the Response
| Response says | Action |
|---|---|
| Approval signals ("thorough", "this covers it", "good") | Proceed to Phase 6, mark reaction_check: passed |
| Specific challenges ("what about X?", "missing Y") | Route challenge back to Phase 4 Designer for revision |
| Structural concerns ("this is too templated", "needs more nuance") | Route to Phase 2 for re-reflection, then re-design |
| Alignment concerns ("this doesn't match the intent") | Route to Phase 1 for seed re-reading |
Max loops: 2 revision rounds. If still failing after 2 rounds, save the workflow with a status: needs-human-review tag and report what the check flagged.
If nothing is configured
Mark the workflow reaction_check: not-run — no predictor configured in its front matter and proceed to Phase 6. Never mark it passed when nothing actually checked it — an unchecked design presented as checked is exactly the failure this whole harness exists to prevent.
Phase 6: PACKAGE & SAVE
What happens: The verified workflow gets saved as a reusable document.
Producer: Document Writer
Write the workflow document with this structure:
---
name: {workflow-name}
domain: {domain identifier}
parent_workflow: /workflow-factory
sub_seed: {path to sub-seed}
status: designed
version: 1
date: {today}
work_type: {from Phase 2}
risk_profile: {from Phase 2}
failure_modes_count: {from Phase 3}
stages_count: {from Phase 4}
reaction_check: {passed | not-run — no predictor configured | needs-human-review}
---
# Workflow: {name}
## Objective (UX Terms)
{what the human gets}
## Failure Modes
{table from Phase 3}
## Stages
{full stage descriptions from Phase 4}
## Agent Roles
{role table}
## Evidence Log Format
EVIDENCE LOG — {domain}
Stage | Item | Claim | Source | Status | Agent | Timestamp
## Self-Healing Protocol
When any stage fails:
1. Error logged with stage, failure mode, and specific evidence
2. Pattern traced to root cause in stage design
3. Stage skill patched to prevent recurrence
4. Next run cannot reproduce same failure class
## Factory Metadata
Classification: {Phase 2 output}
Failure patterns sourced from: {Phase 3 sources}
Auditor gaps found: {count from Phase 4}
Auditor gaps patched: {count}
Reaction-check loops: {count}
Reaction-check verdict: {summary, or "not run"}
Save to a folder in your own project — default docs/workflows/{domain-name}.md unless you've configured a different location. Tell the person the path you saved to.
Checker: Scope Guardian
Final check:
- Every scope item from the seed appears in a stage
- Every failure mode has a prevention stage
- The evidence log format is defined
- The self-healing protocol is present
- The file is saved to the correct directory
If any check fails → route back to Document Writer.
The Self-Healing Meta-Loop
The factory itself improves with every run. Here's how:
What Gets Patched
| When This Happens | Patch This |
|---|---|
| Phase 1 checker finds missed scope items | Phase 1 extraction template — enumerate explicitly |
| Phase 2 checker finds misclassification | Phase 2 classification rules — add the signal that was missed |
| Phase 3 checker finds fabricated failure mode | Phase 3 researcher prompt — require source citation |
| Phase 3 checker finds missed failure mode | Phase 3 research tracks — add the source that had it |
| Phase 4 auditor finds uncovered scope item | Phase 4 stage selection rules — add stage for this pattern |
| Phase 4 auditor finds uncovered failure mode | Phase 4 failure-to-stage mapping — add prevention |
| Phase 4 auditor finds principle violation | Phase 4 principle checklist — flag this for this work type |
| Phase 5 reaction-check raises a challenge | Phase 4 design — address the specific challenge |
Where Patches Are Logged
Every patch is logged in the workflow's Factory Metadata section AND in a running factory log, in the same folder as your saved workflows:
docs/workflows/_factory-log.md
Format:
## Patch Log
| Date | Domain | Phase | Gap | Patch Applied | Prevents |
|------|--------|-------|-----|--------------|----------|
This log is the factory's learning history. Future runs READ this log in Phase 3 to avoid repeating known issues.
Autonomy Protocol
The factory operates WITHOUT human intervention for most decisions. Here's when it acts autonomously vs when it escalates:
| Situation | Action |
|---|---|
| Seed exists, objective is clear | Proceed autonomously through all 6 phases |
| No seed found | STOP. Report to human: "No seed found. Run /alignment-harness:align first." |
| Classification uncertain (Phase 2) | Check the principles oracle if configured. If still uncertain, default to RESEARCH type with THOROUGHNESS constraint. |
| Failure modes uncertain (Phase 3) | Check whatever sources are configured. Include ALL 5 canonical failure modes as baseline regardless. |
| Design auditor finds gaps (Phase 4) | Fix autonomously. Only escalate if gap requires new stage type not in the library. |
| Reaction-check fails (Phase 5) | Revise up to 2 times. If still failing, save with needs-human-review tag. |
| Checker and producer disagree | Checker wins. Always. The producer revises. |
The Agent Uncertainty Protocol
CRITICAL: When any agent in the factory has uncertainty > 30% about what to do:
- STOP — do not proceed uncertain
- Search — check institutional memory (if configured) with the specific question
- Query — check the principles oracle (if configured) on the relevant domain
- Predict — check the reaction predictor (if configured) with the specific decision
- Only then proceed — with whatever enriched context was actually available
This protocol applies to EVERY agent in the factory, not just the orchestrator. A verification agent uncertain about whether a source is reliable → stops, searches whatever's configured, then decides. A designer uncertain about which stage to include → stops, checks whatever's configured, then decides. Where nothing is configured for a step, say so and proceed on your own best judgment rather than stalling.
How to Invoke
Single workflow:
/workflow-factory
Objective: "Audit a codebase's error handling for silent failures"
Seed location: docs/seeds/error-handling-audit.md
Output: docs/workflows/error-handling-audit.md
Batch (multiple domains):
/workflow-factory --batch
Seeds directory: docs/seeds/
Output directory: docs/workflows/
In batch mode, the factory runs Phase 1-2 for ALL seeds first (to classify work types), then runs Phases 3-6 in parallel for each seed. This allows the factory to learn from early workflows and apply that learning to later ones.
Test mode:
/workflow-factory --test
Objective: {any objective}
In test mode, the factory runs all 6 phases but does NOT save the output. Instead, it prints the full trace (every phase, every check, every patch) so the human can verify the factory is working correctly. This is also the recommended first run on a fresh install — a toy objective through --test shows every phase, check, and honest fallback before you trust it with real work.
Factory Version History
| Version | Date | Change |
|---|---|---|
| v1 | 2026-04 | Initial creation. 6 phases, producer/checker pattern, self-healing meta-loop, autonomy protocol. |
Self-Healing Patch Log (example — one factory's history; yours starts empty)
| Date | Phase | Gap | Patch Applied | Prevents |
|---|---|---|---|---|
| 2026-04 | All | Agents exhaust context researching before writing deliverable | Added Write-First Protocol: save draft within first 50% of context budget | Lost work from context exhaustion |
| 2026-04 | Phase 6 | Agents produce sub-artifacts but miss the main deliverable | SAVE instruction made explicit and FIRST priority in Phase 6 | Incomplete output |
Write-First Protocol (MANDATORY — learned from a real production failure)
Agents MUST save a draft deliverable within the first 50% of their context budget. The pattern that fails: research exhaustively → run out of context → deliverable never written. The pattern that works: read seed → design stages → SAVE the workflow file → THEN enrich with research if context remains. This was learned the hard way — two agents ran out of context mid-research and produced nothing; both succeeded once redeployed with this instruction.
Known Gaps for Future Enrichment
- Phase 3 research tracks could be broader — currently uses memory search + a principles oracle (if configured) + case studies. Could also query additional domain-specific sources as you accumulate them.
- Phase 4 stage selection rules are initial — they'll improve as the factory processes more domains and the self-healing log accumulates patterns.
- Batch mode parallelism needs testing on your own setup — does learning from early workflows actually help later ones?
- The Uncertainty Protocol should track how often agents stop to search — if it's more than half the time, the factory needs more context injected upfront.