← the whole session plugin/skills/flywheel-consultant/SKILL.md

Expert on self-improving agent flywheels — systems where session data gets extracted into structured oracles that steer future sessions better. An alignment flywheel built this way is the first working specimen of this pattern. Use this skill to understand how it works, audit signal quality, extend it to new data sources, diagnose degradation, OR to abstract the pattern and build a flywheel for a completely different domain. This is a case study AND a template.

Flywheel Consultant — An Alignment Flywheel (Case Study + Template)

This skill has two modes:

Mode 1 — Case Study: You are working on or auditing an alignment flywheel built this way — described below as a worked example. You need to understand the architecture, signal quality, oracle loop, or how to extend/diagnose it.

Mode 2 — Template Extraction: You are building a flywheel for a DIFFERENT domain — sales, coaching quality, customer support, product development, anything. The case-study flywheel is the working specimen; this skill teaches you to extract the underlying pattern and apply it.

The case study below is one proven instance of this pattern, built for one person's own workflow. Every number, file path, and identifier in it belongs to that one machine — none of it exists on a fresh install, and building your own version of it is itself a substantial, optional project (see "Setup This Piece Needs," further down). Read it to understand every architectural decision, which was earned through failure and correction; then either replicate the shape for yourself or skip straight to Mode 2's general template.


What the Flywheel Is

Sessions → extraction → oracles → better sessions → more extraction → stronger oracles.

The loop is not metaphorical. Here is the exact mechanism, in one working build:

  1. Agent sessions happen — the person works with Claude Code agents. In this build's setup, this meant thousands of session JSONL transcript files across dozens of project directories at ~/.claude/projects/. Yours will be however many you've accumulated — check with ls ~/.claude/projects/*/*.jsonl | wc -l if you're curious.

  2. Sessions contain signal markers — specific skills print recognizable output blocks that can be extracted verbatim. These are the 6 signal types (see below).

  3. Signal gets extracted into structured files, then loaded into whatever oracle/retrieval tool you use. This build uploads them to Google Drive and adds them as sources to NotebookLM notebooks; any tool that can ingest a folder of text files and answer questions against it works the same way.

  4. The oracle notebooks/indexes become oracles — in this build, 4 NotebookLM notebooks query all extracted data and answer questions about this project's intent, alignment patterns, and agent behavior.

  5. Oracle loop steers agents — /jonathan-check3 (this plugin's oracle-loop skill, if installed) presents current scope to the highest-tier oracle, gets back a predicted response from the person plus a next action plus a SCOPE_COMPLETE or SCOPE_INCOMPLETE signal. The agent executes, loops until scope is done.

  6. Sessions continue → more signal → the corpus grows → oracles improve → steering accuracy compounds.

None of steps 3–5 require NotebookLM specifically or even require a hosted oracle at all. If you haven't set one up, the honest fallback is: an agent reasoning directly from the raw extracted files (grep, read, summarize) instead of a semantic-search oracle. It's slower and less precise, but it uses the exact same extracted signal and never silently pretends an oracle answered when none exists.


The 6 Signal Types and Their Extraction Markers

Every data store is built from a specific marker pattern in session transcripts. These markers are the extraction delimiters — without them, data is invisible to the flywheel.

Signal Type Marker What It Captures
Coherence checks COHERENCE_CHECK on its own line (within 8 lines of 1. or **1.**) Real-time alignment decisions: what the agent noticed, what it was uncertain about, how it resolved gaps
Reflect outputs ## AGENT_REFLECTION Self-reflection from /reflect: target, location, gap, initiative, leverage, questions
Align sessions ★ Alignment_Calibration_System_Initiated Full /align intent mapping sessions — nested intent discovery, certainty levels, memory-search results
Scope declarations ## SCOPE_DECLARATION / ## END_SCOPE_DECLARATION Agent contracts before implementation: decomposed UX intent statements, boundary, verification, certainty
Work complete ## WORK_COMPLETE /communicate-what-you-finished outputs — UX stories brought to reality, what's incomplete, verification commands
Skill sequences skillInvocations array in your compaction records Ordered skill invocations per session, read from wherever you store session compactions (see the agentic-session-compactions skill — local file by default, a hosted API only if you built one)

The 7 Data Stores (one example layout)

In this build, all files live at <project>/docs/intent/session-intelligence/ (under whatever this project calls its intent docs). If you build your own version of this, put them wherever this project keeps generated docs — the exact location doesn't matter, only that extraction and the oracle both agree on it.

File What It Contains Size (this build, for scale)
coherence-checks-raw.md Coherence check blocks pulled from JSONL transcripts ~1.5MB in this corpus
reflect-raw.md /reflect outputs ~540KB
align-raw.md /align sessions ~376KB
skill-sequence-corpus.md Sessions with governer score + scope + ordered skill invocations 28KB
compaction-intent-index.md Parent/child intent pairs from all compactions 257KB
scope-declarations-raw.md Verbatim SCOPE_DECLARATION blocks 35KB
work-complete-raw.md Verbatim WORK_COMPLETE blocks 9KB

If you upload these to an oracle tool with its own per-file IDs (Google Drive, etc.), track the ID-to-file mapping in your own notes — there's nothing generic to say about a specific ID, since it's assigned by that tool the moment you upload.

Format discipline is critical. Each file uses a flat, marker-based format so a retrieval tool can find individual records without context overflow:

  • Coherence checks: {date} {session_id[:8]}\n{raw block}\n\n---
  • Skill sequences: {date} {session_id[:8]}\ngov:{score}\nscope: {title}\nux: {ux_title}\nskills: skill1(xN) → skill2(xN)\n\n---
  • Compaction intents: {date} {session_id[:8]}\nparent: {overarchingIntentTitle}\nchild: {targetUxIntentTitle}\n • {decomposed statements}\n\n---

The 4 Notebooks (one example split)

In this build this lives in a NOTEBOOK-ARCHITECTURE.md file alongside the data stores above, documenting which notebook has which sources.

Notebook Purpose Sources
1 — Intent Nesting Answers: what are the main initiatives, how does child UX intent connect to parent strategic intent, what patterns exist across sessions compaction-intent-index.md
2 — Generalist Agent Oracle Answers: what was the actual intent in session X, what do alignment failures look like, what scope declarations agents make All coherence, reflect, align, scope, work-complete files
3 — Skill Sequence Prediction Answers: given this scope type, what skill sequence completes it, what does a completed session look like structurally skill-sequence-corpus.md
4 — Everything ("God Mode") Answers: what would the person say, what's the next action, is this scope complete — the master oracle powering /jonathan-check3 All 7 data sources

You don't need 4 separate notebooks to get value from this — one combined index is a fine starting point. The split exists because a narrower index gives more precise answers for narrower questions.


The Oracle Loop — /jonathan-check3

The full-context oracle loop is the payoff of the flywheel. It runs autonomously until the oracle signals the scope is done. See /jonathan-check3 (this plugin's own copy of that skill, if installed) for the mechanics; the shape is:

Step 1 — Present current state to oracle:
  CURRENT SCOPE: {one sentence}
  WORK DONE SO FAR: {what was built}
  GOVERNER SCORE: {N/100}
  OPEN QUESTIONS: {anything uncertain}
  → Ask: what would the person say? what's the next action? SCOPE_COMPLETE or SCOPE_INCOMPLETE?

Step 2 — Print oracle response observably

Step 3 — If SCOPE_INCOMPLETE: execute the next action, loop back to Step 1
          If SCOPE_COMPLETE: stop, present to the person

Step 4 — Never silently skip a loop

The oracle can do this because, in this corpus, it has access to over a thousand real alignment decisions, hundreds of self-reflection outputs, over a hundred intent-mapping sessions, dozens of skill-sequence records, and hundreds of session-intent pairs — all from real work with that agent setup. It pattern-matches across all of them to predict what the person would actually say. Your own oracle's quality is a direct function of how much of your own real session history you've fed it — there's no shortcut to that volume, though a smaller corpus still beats none.


The 6 Operating Principles

These are the hard-won rules behind this build (originally surfaced by querying a principles oracle during the flywheel's design):

1. Preserve causal DNA — extraction must be verbatim Paraphrased or summarized extraction destroys the signal. The oracle works because it can match against the person's actual words, corrections, and frustrations verbatim. Every extraction block must be copied literally from the transcript, not interpreted.

2. Predictable H2 markers are the only reliable extraction delimiters ## AGENT_REFLECTION, ## SCOPE_DECLARATION, ## END_SCOPE_DECLARATION, ## WORK_COMPLETE — these exist specifically because regex can reliably find them. When adding new signal types, always create a new H2 marker that will appear on its own line. Never rely on prose proximity or semantic similarity for extraction.

3. Signal before compression The extraction pipeline must run before compaction, not after. Compaction summarizes and destroys the fine-grained signal. The markers must be in the full session transcript, not the compacted version.

4. Store extracted signal somewhere durable, not just in session memory Whatever cross-session memory mechanism this project has (a local compactions folder by default — see agentic-session-compactions — or your own hosted store if you built one) is the canonical persistence layer. Don't rely on any short-lived in-session memory as the place signal survives between sessions.

5. Seed + track pattern — extraction is a project, not a one-time script The flywheel needs a seed file for each data-extraction initiative and a tracking mechanism to know what's been processed. New sessions are added continuously — the pipeline must be re-runnable and able to identify what's new since the last run.

6. Self-correct the source, not just the output When an oracle response reveals a drift or misalignment, the fix is not to patch the agent's output — it's to identify which marker or data source produced the bad signal and fix the extraction rule. The flywheel self-improves by improving its inputs, not its outputs.


The Gold Standard: reflect → align → greenlight

Not all signal is equal. The single highest-quality signal the flywheel can receive is the reflect→align→greenlight sequence — and it's the only type that contains explicit human validation.

Here's why this trio is categorically different:

  • /reflect — the agent steps back and examines its own work against the intent. This is the agent asking "am I actually building the right thing?"
  • /align — the person and the agent co-construct a shared understanding of intent. Nested intent gets surfaced, named, and made specific.
  • greenlight — the person explicitly confirms items in the intent map (a confirmation mark, an ID confirmation, "yes that's right"). This is a human saying "the agent's model of my intent is correct."

The greenlight is what makes this trio irreplaceable. It's the only moment where the oracle can be certain the agent understood correctly and the person agreed. Every other signal type is agent-generated, potentially self-serving, and unconfirmed by a human. Coherence checks are triggered during real work but not explicitly approved. Work-complete blocks are useful but the human didn't confirm the work matched the intent.

Implication for the flywheel: when you see a session with a reflect→align→greenlight sequence followed by a scope declaration and work-complete — that session is worth triple the weight of a session where the agent just ran governer and started building. Prioritize refreshing the oracle with sessions containing this sequence.

Implication for signal design: when adding new signal types to the flywheel, ask: "Does this contain explicit human confirmation, or just agent output?" The former feeds the oracle's confidence in predictions. The latter feeds pattern recognition but not ground truth.


Signal Quality — What Degrades the Flywheel

The flywheel degrades when signal quality drops. These are the known failure modes:

Paraphrased extraction — copying the gist instead of verbatim text. The oracle loses the ability to match on the person's actual language. Detection: query the oracle with a known exact phrase from a session and see if it retrieves the right context.

Marker drift — skills change their output format without updating the extraction regex. A skill that used to print ## AGENT_REFLECTION now prints # Agent Reflection — the extractor misses it entirely. Detection: run extraction and compare record counts to what you'd expect given your session volume; a sudden drop to near-zero for one signal type is the tell.

Gap in coverage — new sessions accumulate but extraction only ran once. The oracle is steering from stale data. Detection: check the most recent date in the data stores vs today's date. If more than 2 weeks old, run re-extraction.

Source size overflow — a data store grows past your oracle tool's per-source limit (NotebookLM's is roughly 50MB per source). The notebook silently drops the overflow. Detection: check that all records are queryable (query for a known recent session and verify it returns).

Wrong notebook/index — an agent queries the narrow generalist oracle when it needs the full "everything" oracle for full-context prediction. Detection: check which notebook/index is configured wherever the oracle-loop skill stores that setting.


How to Extend the Flywheel

When adding a new data source to the flywheel:

  1. Identify the signal — what does an agent print to chat that captures valuable knowledge? It must be printed verbatim, not computed silently.

  2. Create a marker — add an H2 markdown header that will appear on its own line. Example: ## LEVERAGE_ASSESSMENT. This is the extraction delimiter.

  3. Update the skill — modify the skill file that produces this output to print the marker before and after the block.

  4. Write the extraction script — grep the JSONL files for the marker, extract the block, format it in the flat {date} {session_id}\n{content}\n\n--- pattern.

  5. Run extraction — produce the raw data store file wherever you're keeping this project's generated intent docs.

  6. Load it into your oracle tool — if you've connected a document-sync integration (Google Drive or similar, via whatever MCP or connector you've set up), use that; for a small file, hand-uploading through the tool's own UI works too. There's no single canonical path here — use whatever you already have wired up, and if nothing is wired up, this step waits until it is.

  7. Add to the relevant notebook/index — determine which of your notebooks/indexes should have this as a source, and add it through that tool's own UI.

  8. Update your own architecture notes — add a row to your own file table (mirroring the "7 Data Stores" table above), documenting the purpose and destination.


Files to Know (one example layout — yours will differ)

  • NOTEBOOK-ARCHITECTURE.md — wherever you keep the data stores above — maps your notebooks/indexes, their sources, and any tool-assigned ids, plus your own creation instructions
  • jonathan-check3's SKILL.md — this plugin's own copy, if installed — the oracle loop skill (the "everything" notebook/index id lives here once you set one up)
  • A principles oracle, if you've built one — query it for operating principles before making architecture decisions
  • Your data stores directory — wherever step 5 above wrote to — all raw data files live here

Setup This Piece Needs

None of the above is active until you've actually built it — this whole flywheel is an optional, substantial project layered on top of the harness, not something a fresh install gets for free. To build your own version:

  1. Pick an oracle tool (NotebookLM is what this build used; any tool that can ingest a folder of text and answer semantic questions against it works the same way).
  2. Add extraction markers to the skills you want to feed it (start with /reflect, /align, /declare-scope, /communicate-what-you-finished if you have them).
  3. Run extraction against your own ~/.claude/projects/*/*.jsonl history, building the 7 data stores above.
  4. Load them into your oracle tool, note whatever id it gives you, and wire that id into /jonathan-check3's skill file (or the equivalent oracle-loop skill you're using) wherever it stores that setting.

Until you've done this, /jonathan-check3 and this flywheel have nothing to query — say so plainly rather than pretending an oracle answered. The jonathan-check2 skill (a lighter, one-shot response predictor) is a reasonable fallback if you want prediction without building the full flywheel first.


Signal Quality Hierarchy — What the Oracle Actually Depends On

Not all signals are equal. Ranked by quality, from highest to lowest, based on this build:

1. Coherence check blocks The crown jewel. These are captured at the exact moment of misalignment — real friction, real correction in the session. They're involuntary signal, which makes them more truthful than intentional outputs. In this corpus, sessions with more emotionally charged language in them correlated with more coherence failures — worth tracking if you want an early-warning signal, though the specific correlation will be personal to how each person writes when frustrated. If you had to pick one signal source to power the oracle, this is it.

2. Compaction intent pairs The parent intent + child UX intent for every session. Tells the oracle what people were actually trying to accomplish. The strongest predictor of what a new session SHOULD be trying to accomplish.

3. Align session outputs Full intent-mapping sessions where nested intent was surfaced and validated. High signal because they contain the person's actual words describing what they wanted, at moments of deliberate clarity.

4. Reflect outputs Self-reflection moments — the agent stepping back and examining its own work. These surface what the agent modeled as the gap, the target, and the leverage. Useful for understanding agent reasoning quality over time.

5. Skill sequences Context + governer score + ordered skill invocations. The only signal type that enables NEXT-ACTION PREDICTION. Without this, the oracle can say "the person would say X" but not "therefore do Y."

6. Scope declarations Agent contracts before implementation. Useful for auditing whether what was scoped matched what was built.

7. Work complete blocks Completion summaries. Often the lowest count of the seven because the producing skill isn't invoked consistently. Each one is high quality; coverage is the gap.


Known Signal Gaps — Markers Missing from High-Signal Skills

These are skills that can produce valuable output but whose content is invisible to a flywheel unless they carry an extraction marker. If you're building your own flywheel and want these skills to feed it, add the markers first:

Skill Signal Type Gap Recommended Fix
/jonathan-check2 Response predictions + challenge list Uses visual separator lines, no ## H2 marker Add ## JONATHAN_CHECK / ## END_JONATHAN_CHECK
/jonathan-check3 Oracle loop responses + SCOPE_COMPLETE signals Uses visual separator lines, no ## H2 marker Add ## JONATHAN_CHECK_V3 / ## END_JONATHAN_CHECK_V3
/decompose Decomposed UX intent statements + confidence scores Has an opener marker but no closing output block Add ## DECOMPOSITION_OUTPUT / ## END_DECOMPOSITION_OUTPUT
/insight Notebook query responses Verbatim oracle responses printed but no ## marker Add ## ORACLE_QUERY / ## END_ORACLE_QUERY

Impact of these gaps: every response prediction, every oracle-loop iteration that reaches SCOPE_COMPLETE, every decomposed UX statement, and every oracle query response is currently lost to the flywheel until these markers are added — these tend to be some of the highest-frequency, highest-quality outputs in the system.


Oracle Cadence — The Refresh Problem

The oracle does not auto-refresh. This is the most operationally dangerous gap in this kind of system.

When new sessions happen — and they happen every day — the new coherence checks, reflect outputs, and skill sequences are NOT automatically added to your notebooks/indexes. The oracle continues to steer from stale data unless someone manually:

  1. Re-runs extraction scripts on the new JSONL files
  2. Loads the updated data stores into the oracle tool (replacing or appending to the existing sources)
  3. Re-syncs or re-adds the notebook/index sources

Refresh cadence recommendation:

  • Weekly for high-frequency signals (coherence checks, skill sequences)
  • Monthly for lower-frequency signals (align sessions, work complete)
  • After any major alignment correction session (same day)

How to know when the oracle is stale: query it for a known recent session by date. If it can't retrieve anything from the last 2 weeks, the data stores need refreshing.

The extraction scripts, in this example layout, live at: wherever the data stores from earlier live — re-run the build scripts against the JSONL files, reload the updated files into your oracle tool, and it will use the new versions on next query.


The Flywheel as a Template — Building Flywheels for Other Domains

This pattern is domain-agnostic. This section teaches you to extract it and apply it anywhere.

The Universal Flywheel Pattern

1. Identify what knowledge matters in your domain
   → What information, if captured from every session, would make future sessions better?
   
2. Design the signal events
   → What moments produce that knowledge? (alignment corrections, closing decisions, 
     evaluation outputs, customer reactions, pricing negotiations, etc.)
   
3. Create extraction markers
   → What H2 header will appear when each signal event fires?
   → Rule: must be on its own line, must be unambiguous, must wrap the full output block
   
4. Build the extraction pipeline
   → Grep the session transcripts for markers
   → Format into flat, marker-delimited data stores
   → Upload to oracle (NotebookLM or equivalent)
   
5. Wire the oracle loop
   → Define the oracle query: current state + what matters → predicted next action
   → Loop until the oracle signals the session goal is achieved
   
6. Maintain the signal
   → Refresh cadence
   → Marker coverage audit
   → Signal quality monitoring

What Makes a Good Signal Event

A signal event is worth capturing if:

  • It contains knowledge a future agent couldn't easily reconstruct from context alone
  • It happens recurrently (the same type of event fires across many sessions)
  • Its content is specific enough to be useful (not generic, not always the same)
  • Capturing it doesn't meaningfully slow down the primary session

A signal event is NOT worth capturing if:

  • It's always the same text (no variance = no learning)
  • It requires the agent to stop and generate summary content (interrupts the session)
  • It duplicates information already persisted to a database

Example Domains Where This Pattern Applies

Domain High-Signal Events Extraction Marker
Sales Pricing objection + resolution, competitor comparison, deal structure decisions ## DEAL_DECISION
Customer support Escalation decisions, root cause determinations, resolution patterns ## RESOLUTION_PATTERN
Coaching quality Session turning points, breakthrough moments, technique evaluations ## COACHING_ASSESSMENT
Product development Scope change decisions, design tradeoffs, user feedback interpretations ## PRODUCT_DECISION
Financial analysis Model assumptions, scenario interpretations, risk flags ## ANALYSIS_JUDGMENT

What This Case Study Teaches

This flywheel reveals three non-obvious things about flywheel construction:

1. Involuntary signal beats intentional signal The coherence check is the most valuable signal because it's captured at the moment of friction — not curated afterward. When designing a new flywheel, find the moments of friction in your domain. Those are your highest-quality signal events, not the structured summary outputs.

2. The oracle's confidence depends on marker precision A vague marker produces noisy retrieval. COHERENCE_CHECK near a numbered list is precise. A ## NOTES block with mixed content is not. The more precisely the marker identifies the specific type of knowledge, the higher the retrieval quality.

3. Frequency of extraction is a design decision This flywheel extracts from JSONL transcripts, which means extraction requires a pipeline run, not live capture. For some domains, real-time persistence to a database is more appropriate than batch extraction. The choice depends on how quickly the knowledge needs to be queryable.


Signal Quality Audit Procedure

Run this when you suspect oracle quality has degraded, or after 2+ weeks without extraction:

1. Check marker counts
   grep -c "COHERENCE_CHECK" /path/to/new-sessions/*.jsonl
   → Expected: a handful per session. If 0, the coherence gate hook may have stopped firing.

2. Check date coverage
   → Query oracle: "What happened in the session from [last week's date]?"
   → If it returns nothing, data stores are stale.

3. Check signal variance
   → Run a sample extraction of the last 20 sessions
   → Are coherence checks still varied? Or does the same text appear repeatedly?
   → Low variance = the skill is producing templated output, not authentic signal.

4. Check marker integrity
   → Did any skill updates remove or rename the markers?
   → Search the skill files for the expected marker strings.

5. Check source size
   → Confirm each source file is under your oracle tool's per-source limit
   → Large files may be silently truncated (NotebookLM's limit is roughly 50MB per source).

  • /jonathan-check3 — the oracle loop this flywheel powers, if you've built one. Use for autonomous steering and completion gates on high-stakes work.
  • /jonathan-check2 — one-shot response prediction. Use before presenting work to the person, whether or not the full flywheel exists.
  • /reflect — produces ## AGENT_REFLECTION blocks. Every reflect output feeds the flywheel.
  • /align — produces ★ Alignment_Calibration_System_Initiated blocks. Every align session feeds the flywheel.
  • /declare-scope — produces ## SCOPE_DECLARATION / ## END_SCOPE_DECLARATION blocks. Every scope declaration feeds the flywheel.
  • /communicate-what-you-finished — produces ## WORK_COMPLETE blocks. Every completion feeds the flywheel.
  • /governer — scores tasks. Scores are recorded in compactions and, if you extract them, feed the skill-sequence-corpus.
  • /insight — queries whatever oracles you've set up. Each query should eventually feed back into the flywheel once ## ORACLE_QUERY markers are added.
  • /decompose — produces UX intent statements. Once ## DECOMPOSITION_OUTPUT markers are added, every decompose output feeds the flywheel.