← the whole session plugin/skills/smc-frustration-detector/SKILL.md

A reference pattern for detecting frustration spikes in a person's messages and logging a telemetry event when they happen. Wire it into your own project's UserPromptSubmit hook if you want it running on every message — it is not wired in automatically by installing this skill.

Frustration Detector

When someone's messages carry noticeably more frustration than their own normal level, it's usually a sign the agent has drifted from what they meant — misread the intent, ignored it, or something upstream is breaking and they're reacting to the fallout. This pattern catches that moment and logs the full state of what was going on, so it can be looked at afterward instead of just felt and forgotten.

This ships as a working reference script (scripts/frustration-detector.js, no dependencies) that you can call from your own project's UserPromptSubmit hook. It is not something this skill wires into any hook automatically — you decide whether and where to run it.

How It Works

  1. Marker counter — counts frustration markers (a configurable list — see below) per message in each session
  2. Baseline — your own resting rate, measured from your own history, not assumed
  3. Threshold trigger — fires when the session's rate exceeds 1.5× baseline with a minimum of 5 marker occurrences
  4. Telemetry event — logs a structured event on first fire per session
  5. No spam — fires once per session; later high-frustration messages are still counted but don't re-fire

Fails open, always. A bug in the detector must never block or crash the hook it's wired into — every code path in the reference script catches its own errors and exits quietly rather than interrupting the message.

Markers — not just profanity

Not everyone shows frustration the same way. The reference script's default list includes common profanity, but also phrases like "again", "I already told you", "that's not what I meant" — edit the marker list in the script to match how you actually write when annoyed. There's no single right list; the point is comparing your own rate against your own baseline, whatever markers you pick.

Setting your baseline (do this before it can fire anything)

Until you set a baseline, the detector runs in observe-only mode — it counts and logs quietly and says so once per session, but never fires a spike. An uncalibrated threshold isn't tuned to anyone; presenting it as if it were would be worse than saying plainly that it isn't set up yet.

# Scan your own past sessions for this project and get a suggested number
node <path-to-this-skill>/scripts/frustration-detector.js calibrate

# Confirm it
node <path-to-this-skill>/scripts/frustration-detector.js set-baseline 0.02

If you have little or no history yet, calibrate again after a week or two of real use — the suggested number gets more meaningful with more messages behind it.

Wiring it in

Add a call to check mode from your project's UserPromptSubmit hook, piping the hook's own JSON payload to it on stdin:

cat "$HOOK_INPUT_JSON" | node <path-to-this-skill>/scripts/frustration-detector.js check

Put this early in your hook, before anything that depends on injected context — that way frustration detection stays independent of whether your other alignment machinery is even running.

Telemetry Event

Fired when the threshold is breached (written as one JSON line to the folder alignment-harness records frustration-detector prints, events.jsonl):

{
  "type": "FrustrationSpikeDetected",
  "session_id": "claude-ABCD1234",
  "marker_count": 8,
  "marker_per_msg": 0.024,
  "baseline": 0.015,
  "pct_above_baseline": 60,
  "timestamp": "2026-04-05T14:23:17Z"
}

If you want a fuller dump of system state at the moment of the spike (recent errors, governer score trend, which hooks fired, whether your memory-search or oracle setup was active), add that yourself where the reference script logs the event — the shape is up to you and what your own setup actually has running.

Files

  • Reference script: scripts/frustration-detector.js (this skill folder) — the whole detector: counting, baseline, state, event logging
  • Config: the folder alignment-harness records frustration-detector prints — config.json (your baseline and thresholds) and events.jsonl (fired spikes)
  • Per-session state: a file per session under your OS temp directory, cleaned up naturally (temp files, not meant to persist)

To Verify

# Set a low baseline so it's easy to trigger for testing
node <path-to-this-skill>/scripts/frustration-detector.js set-baseline 0.02

# Feed it several frustrated-looking messages in the same test session
for i in 1 2 3 4 5 6; do
  echo '{"prompt":"this is broken again, i already told you this is wrong", "session_id":"test-frustration-1"}' | \
    node <path-to-this-skill>/scripts/frustration-detector.js check
done

# Check the event landed
cat "$(alignment-harness records frustration-detector)/events.jsonl"

You should see exactly one FrustrationSpikeDetected line for test-frustration-1, not one per message.

Troubleshooting

False positives (markers matched in code examples or quoted text):

  • The reference detector is naive — it counts raw string/phrase matches, without checking whether they're inside a code block or a quote.
  • Tighten the marker list, or add context checks, if this matters for your use.
  • The trade-off as shipped: a higher false-positive rate in exchange for not missing real signal — an early-warning system, not a precise one.

Baseline seems wrong after a while:

  • Recalibrate: calibrate again with more history behind it, then set-baseline with the new number.
  • People's baselines can genuinely shift (a stressful week, a new project) — that's useful information, not just noise to filter out.

Why This Matters

When someone gets frustrated with an agent, it usually means one of:

  1. The agent didn't understand the intent (comprehension failure)
  2. The agent understood but ignored the intent (alignment failure)
  3. Something is cascading wrong assumptions faster than correction can happen (runaway)
  4. Underlying tooling is breaking and the agent is fumbling around it

A frustration spike is an early-warning signal that enables:

  • Real-time intervention (the person sees an alert and can redirect immediately)
  • Post-mortem analysis (the logged event shows what was going on at the moment it happened)
  • Ongoing calibration (a baseline that drifts means something about the working relationship or the system changed)

This detector is not meant to judge the person. It's meant to surface when agent behavior has diverged from their intent with enough severity that they expressed frustration about it — that's the signal that something in the system needs fixing.