← the whole session plugin/skills/ask/SKILL.md

Use BEFORE asking the person any clarifying question. Replaces naked questions with a criticality-gated decision flow: synthesize the context the question emerges from, predict the likely answer with a confidence number, score how much it matters using the governer score already run for the session, then route based on the combined criticality. Low-criticality questions get decided by the agent and announced. Medium ones get an institutional-memory pass before reasking. High-criticality ones get the full briefing template and wait for a go-ahead.

/ask

When this fires

Any moment you're about to ask a clarifying question, present a multi-choice menu, request a decision, or otherwise pause work to wait on the person. This skill is the gate BEFORE the question lands.

Includes: "Should I proceed with X or Y?", AskUserQuestion tool calls, "Which approach do you want?", "Want me to do A?", "Do you mean X?", "Should I commit?"

Excludes: status updates, progress reports, requests for explicit destructive authorization (those have their own gates), and technical announcements that don't actually require an answer.


Phase 0 — Quick check: do you actually need to ask?

Before scoring anything: could you make a reasonable call and continue? If the session is running in a mode that says work without stopping for clarifying questions, the answer is almost always yes — make the call, announce it, continue.

If you're still genuinely uncertain, proceed to Phase 1.


Phase 1 — Internal scaffolding (think these, don't write them as prose yet)

Hold these four questions in mind:

  1. The overarching thing I'm working toward — what is the scope-level goal of this session, and how does what the person asked for connect to it?
  2. The challenge that triggered this question — what specific obstacle, fork, or ambiguity made me reach for a question?
  3. What I predict the person would say — based on what's in memory, in the session, what's already been said: what's the single most likely answer? State it as a specific predicted reply, not a category.
  4. My calibration — how confident am I, 0-100, that my predicted answer matches the person's actual intent?

Phase 2 — Score how much this matters

Reuse the risk score already computed for this session rather than inventing a parallel one — that's what /alignment-harness:governer is for.

Step 2.1 — Check for an existing session score

alignment-harness status

If a score already exists for the scope this question is about → reuse it: "score: {N}/100 (inherited from this session's governer run)."

If the question is about a new scope not yet scored, run one:

alignment-harness score --task "<one sentence: what this question is about>" --category <category> --ac <N> --hp <N> --ew <N>

This runs entirely locally — no server, no key required. If you'd rather estimate by hand for a quick in-the-moment question, average these four (1-100), and if any single one is ≥80, raise the overall to at least that:

  • degreeToWhichThisTaskAffectsWhatUsersExperience
  • degreeToWhichAgentMistakesWouldCompoundAcrossFutureSessions
  • degreeToWhichGettingThisWrongWouldSpreadToOtherSystems
  • degreeToWhichThisCouldBreakSomethingAlreadyWorking

Step 2.2 — Compute criticality

GRAVITY = the score from step 2.1 (0-100)
CONFIDENCE = how confident you are your predicted answer matches the person's intent (0-100)
UNSURENESS = 100 - CONFIDENCE
CRITICALITY = (UNSURENESS × GRAVITY) / 100

Criticality scale: 0 = "I'm certain and it doesn't matter anyway" → 100 = "I have no idea and it matters a lot."


Phase 3 — Route based on criticality

One consistent rule, no overlapping bands:

Criticality < 30 — Just decide and continue

The question doesn't actually need to be asked. Make the call yourself. Announce it in one line so it's observable, then keep working.

[auto-decided, criticality={N}/100] {one-sentence decision}. Continuing.

Example: [auto-decided, criticality=12/100] Proceeding with Test 1 first (TDD discipline requires RED phase before fix). Continuing.

Do NOT actually present the question. Do NOT use AskUserQuestion. Just decide and move.

Criticality ≥ 30 — Search first, then reassess

The question might matter, but you may be missing context that would resolve it on its own. Search institutional memory against the question or the scope it's inside — see the agentic-find skill: if set up, run it; otherwise grep your own past sessions (grep -l "<keywords>" ~/.claude/projects/*/*.jsonl) and this project's docs and git log.

Re-run Phase 2 with whatever you found. If criticality now drops below 30 → decide and continue, per the block above. If it's still 30 or above → go to Phase 4 and brief the person. Don't skip the search step even when your first-pass criticality is already high — a two-minute grep that resolves the question is always cheaper than interrupting someone.


This is a peer-to-peer communication, not a form. If you've set up a voice/communication skill for this person (how-to-talk-like-the-founder is one example of one — it works for anyone once it's calibrated to your own person's voice), load it before composing. No jargon, no menus, no headers shouting at the reader — one person telling another what's going on and what they need a moment of input on.

Template:

Hey — running into a question here I want to share so we stay on the same page.

I'm working on {x — the immediate task} in service to {y — the overarching goal}, and to do that
I'm doing {a — the specific approach I picked}, because {how a connects to x and y}.

But when doing {a} I'm seeing some options here, and my prediction is that you'd want
{predicted answer} because {reason that prediction connects to what I know about your prior
stated intent}.

So here's the question, along with my prediction of your answer:

  {question, in plain language, one sentence}
  {predicted answer, one or two sentences}

My confidence on the prediction is {N}%. The gravity of getting this wrong is {N}% because
{what would actually break, or what intent it could violate}. Combined criticality is {N}%.

Reply with "yes" or "looks right" if my prediction matches, or correct me if not.

After printing this, STOP. Do not start work on the predicted answer. Wait.


If you're tracking this kind of thing, log each invocation somewhere durable — a JSONL file under alignment-harness records ask-telemetry works fine if you don't have anything fancier:

{"ts":"<iso timestamp>","event":"AskGated","sessionId":"<id>","gravity":N,"confidence":N,"criticality":N,"route":"auto-decided|search-reassess|full-briefing","question":"<first ~120 chars>"}

Over time this lets you see: how often the gate auto-decides vs escalates, how often predicted answers turn out to match what the person actually wanted, and whether your criticality cutoff (30, above) is calibrated right for how you personally want to be interrupted — raise it if you're being asked too often, lower it if the agent is deciding things you'd rather weigh in on.


Phase 6 — What NOT to do

  • Do not invent a parallel scoring system. The governer score is canonical for this harness. If the score feels wrong, fix the governer inputs — don't add a second formula here.
  • Do not present a menu of options. Either auto-decide or predict the answer. Never "(A) X (B) Y (C) Other."
  • Do not ask the question without predicting the answer first. A naked question skips the whole point of this gate.
  • Do not use the built-in question-popup tool for criticality < 30. Auto-decide and announce instead.
  • Do not skip connecting the question to the overarching goal. The person needs to see WHY this question matters, not just the question in isolation.
  • Do not write the four Phase 1 scaffolding questions out as prose to the person. They're internal. Synthesize into the briefing template.

Examples

Example A — Auto-decided (criticality 4)

Working on a TDD-discipline task. About to write Test 1 (RED phase). Could either write all three tests at once or just Test 1 first.

Gravity: 72/100 (the parent task is high-leverage). Confidence in prediction: 95% (the person explicitly invoked "TDD" — RED phase first is what that means). Unsureness: 5. Criticality: (5 × 72) / 100 = 3.6.

Prints: [auto-decided, criticality=4/100] Writing Test 1 alone first per TDD RED-GREEN discipline. Test 2 comes after fix lands. Continuing. No question asked.

Example B — Full briefing (criticality 62)

About to commit. Notices the commit touches a file the person has historically been careful about. Unclear whether existing scope authorization extends to it.

Gravity: 88/100 (touches auth code). Confidence: 30% (the person's prior stance here is ambiguous). Unsureness: 70. Criticality: (70 × 88) / 100 = 61.6 → 62.

Prints the full briefing template, waits.

Example C — Search reassess

Choosing between two library options. Both look reasonable, no strong initial signal.

Gravity: 50/100. Confidence: 50%. Unsureness: 50. Criticality: (50 × 50) / 100 = 25 → under 30, would auto-decide as-is. But say gravity were higher — 60/100 with the same confidence gives criticality 30, right at the search threshold: run the institutional-memory search first regardless of which side of 30 you land on when it's this close.

Search turns up an old note: this person prefers minimal dependencies. Confidence in the prediction jumps to 85%. Unsureness drops to 15. New criticality: (15 × 60) / 100 = 9 → well under 30 → auto-decide, announce, continue.


Why this skill exists

Agents tend to over-ask: pausing for confirmation on things they could decide, presenting menus when they should predict, and — when they DO need input — asking the question in isolation without the context that makes it answerable. This skill replaces the binary "should I ask" decision with a calibrated one, using the same risk score as every other gate in the harness, so the ask/decide threshold isn't a separate invented rule. And when it does ask, the briefing template forces context, prediction, and stakes into one message that takes a few seconds to scan instead of minutes to reconstruct.

Discoverability

This skill is invoked:

  1. Directly: /ask.
  2. By skill chains: any skill about to call the question-popup tool should invoke this first.
  3. By self-recognition: any time you think "I should ask about X" — that thought is the trigger.