← the whole session plugin/skills/initiative-playground/SKILL.md

Reusable pre-code playground for any major, user-impacting initiative. Before touching real code, model the whole thing as wikilinked Markdown files (runtime vs build-harness), anchor the person's verbatim intent, decompose quotes to plans, make every decision observable, run A/B tests in chat, log every correction back into the model, and refine the response on evidence until it's right. Only then build. Use when the person you're working with says "explore this potential / I want to build X / let's sandbox this / can we get a better response before we touch code" for anything that would impact users.

Initiative Playground — perfect the response before any code exists

What this is (feature-agnostic)

A repeatable method for de-risking any major, user-impacting initiative. Instead of improvising, you instantiate this pattern: build a code-shaped model in Markdown that you can actually operate and test — anchoring the exact words of the person you're working with, making every decision observable, A/B testing responses, and logging every correction back into the model — until the response is right in the sandbox. Only then is real code written.

The point: get the response right in a playground before touching the codebase, so users are never the test subjects and the eventual build inherits clean, evidence-earned boundaries.

Example from a real project (opt-in illustration — the shape, not the content, is what matters): a sealed system that classified a pattern in short pieces of text was built this way, with a 00-SEED.md root, a runtime/ folder, and a build-harness/ folder. If you want to see a filled-in instance before building your own, make a small neutral one first — for example a checkout-error-message redesign, worked through the same nine steps below — rather than reaching for someone else's private product data.

When to use

The person you're working with is starting a major initiative that impacts users, and wants to explore/refine the approach before committing code. Signals: "explore this potential," "I want to build X," "let's sandbox this," "can we get a better response first," "abstract this into a pattern."

The pattern — abstracted from the worked example

Each element below is GENERAL. The coaching example shows one instantiation in parentheses.

1. Anchor the verbatim intent (never paraphrase)

Create 00-SEED.md as the root. Paste the EXACT words of the person you're working with, every message, as a quote block. Keep "what was said" (verbatim) separate from "what I think it means" (labeled agent interpretation). If anything downstream conflicts with the verbatim block, the verbatim block wins. (Example from a real project — labelled opt-in illustration: several of the person's messages preserved verbatim, one ambiguous term kept alongside the agent's interpretation of it, clearly marked as interpretation, not fact.)

2. Nest the intent, formatted as a plan

Surface the holonic layers: overarching → parent → this-initiative → child UX outcomes. Format as a plan (what success looks like in the user's experience), not as code.

3. Decompose quotes → confirmed plans (zero loss)

For each direct quote, state: the verbatim, what 10/10 looks like, current reality, the gap, the plan + confidence %, and which FILE owns it. Carried in the seed.

4. Model the code architecture as wikilinked files — split runtime vs build-harness

Build the folder the way the CODE would be organized — natural boundaries, single-purpose, nested. Two top-level areas, always:

  • runtime/ — what the USER experiences (would ship to the app).
  • build-harness/ — what WE use to build + test (admin-only, never ships): the observable/admin surface, A/B tracker, evidence log, correction log.
  • the seal/safety boundary wraps both. Wikilink files to each other like code imports. The seed is the entrypoint/index that links to all. Empty-but-intentional dirs get .gitkeep (they grow only from evidence). (example: runtime/ = orchestrator→decision→protocols→principle-store; build-harness/ = surface, ab-tests, evidence-log, correction-log.)

5. Make every decision observable (one sentence each)

As you operate the model, each decision in the cascade emits ONE sentence: what the decision is, why, and which file governs it. Format: ▸ OBSERVABLE — [decision]: [choice] because [reason]. (gov: [file])

6. A/B test in chat, tracked in the harness

When the person you're working with says "A/B test," run a side-by-side and show each response. Always include the REAL current system (the baseline) as one arm. If no current system exists yet, say so plainly, use an unguided response as the stand-in baseline, and label it as a stand-in rather than a real baseline. Record every run in build-harness/ab-tests/runs/ with variants, input, both responses, and the verdict from the person you're working with. (Example from a real project: Arm A = the real production prompt; Arm B = the new candidate approach.)

7. Log every correction back into the model (the learning loop)

Every correction from the person you're working with is an EVENT: log it verbatim in build-harness/correction-log-and-self-update.md, then UPDATE the file it touches, both observable. Distinguish corrections to the RUNTIME/PRODUCT system from corrections to the AGENT's own behavior. This loop is itself a module.

8. Build from evidence only — the gate before code

A response/mechanism may NOT be codified until the person you're working with reports it actually worked in their real experience (for example, in a coaching or conversational product, a genuine shift in how the person felt; in another kind of product, a task actually completed or an error actually avoided). Log the evidence in build-harness/evidence-log/. NO real code until the sandbox response is confirmed right. (/governer's UX-breakage gate for user-impacting work → HUMAN-REQUIRED: browser evidence + explicit authorization before commit.)

9. Show compliance evidence, always

When the person you're working with gives a requirement, the response SHOWS proof it was met (commands + output), not a claim. If you use Obsidian or another linked-note vault (optional — see /alignment-harness:harness-setup), open every file added/edited there so the nested structure is visible. Otherwise, print the file tree and the diff so the same structure is visible without it.

The instantiation checklist (copy when starting a new initiative)

seeds/{initiative}/
├── 00-SEED.md                          # verbatim intent + nested intent + decomposition + wikilink index
├── runtime/                            # what the USER experiences
│   └── {modules at natural boundaries, single-purpose, wikilinked}
├── build-harness/                      # what WE use to build + test (admin-only)
│   ├── observable-surface-admin.md     # the admin witness panel
│   ├── ab-tests/{ab-test-tracker.md, runs/}
│   ├── evidence-log/{evidence-log.md, shifts/}
│   └── correction-log-and-self-update.md
└── {safety-frame}.md                   # the boundary both areas live in (if sealed/flagged)

Every file: intent header + "read the seed first ([[00-SEED]])" + the verbatim quote it owns + wikilinks.

Lifecycle

/align (confirm intent) → instantiate playground (this skill) → operate + A/B + correct in sandbox
→ evidence confirms the response is right → THEN /governer + a written plan (if you have a `/plan` or `/superpowers-writing-plans-v2` skill; otherwise write one inline) + build with subagents → /complete-seed

The playground sits BETWEEN alignment and code. It is where the response is perfected so that code, when written, is building something already proven to work.