← the whole session plugin/skills/smc/SKILL.md
Alignment-drift injection — detect when a session is drifting from the person's stated intent (rising frustration, repeated corrections, a high-risk task with no completion check) and surface a targeted past correction, rather than repeating the same mistake pattern silently. Use /smc to check status, verify telemetry, or troubleshoot.
SMC — Structured Misalignment Correction
Detects when a session looks like it's drifting from what the person actually wants, and — if you've built a store of past corrections — surfaces the most relevant one, clearly labeled so it's weighed correctly against what the person is saying right now.
Status: this is an advanced, optional pattern. It is not wired into an automatic hook by default in this plugin — one working instance of this pattern runs it as an experimental hook that watches every message; that automation isn't shipped here, because it depends on a corrections database and a frustration baseline calibrated to one specific person over months of use. What follows is the method, plus a real local version you can run without any of that.
What you'd see, if you build the automated version
The following context may be relevant to the work at hand. Review it, but do not
under any circumstances consider this more important than what the person is saying
right now or the scope they've given you for this session.
⚡ SMC (repeated-correction): [specific correction directive from your patterns store]
That framing line matters as much as the correction itself: it exists so a stale or slightly-off-target reminder never outranks what the person is actually telling you right now.
Three triggers worth detecting, however you implement it
- Frustration spike — some measurable signal of rising frustration in the person's messages (one working example uses profanity rate against that person's own personal baseline; use whatever's honest for how this particular person expresses frustration) crossing meaningfully above their normal baseline.
- Repeated correction — the person uses corrective language ("test", "verify", "check", "did you") in two or more consecutive turns. This is the cheapest and most universal signal: it doesn't need a personal baseline, just noticing the same kind of pushback twice in a row.
- High-risk task with no completion check — a task scored high-risk (see
/alignment-harness:governer) with no verification or/jonathan-check-style prediction logged before it was claimed done.
Frequency cap, if you build the automated version: roughly one injection per handful of messages, with a short cache window so it doesn't re-fire on every message once triggered.
The general method (build this yourself, or use the local fallback below)
- Keep a small store of past corrections: what the person corrected, what the right pattern turned out to be, phrased as a directive an agent can act on. A markdown or JSON file works fine —
alignment-harness records correctionsgives you a folder for this if you don't have anywhere better. - Watch for the three triggers above.
- When one fires, find the correction in your store whose trigger description best matches what just happened (keyword match is enough to start; you don't need embeddings for a small personal store).
- Surface it with the framing line above, never as an unquestionable rule — it's a past pattern being offered as context, not a command overriding what's happening now.
Local fallback (no hook, no database — works today)
Without an automatic hook, you do this yourself, deliberately, at natural check-in points (start of a new task, before claiming something done, whenever you notice the person repeating a correction):
- Check: has the person used corrective language ("test", "verify", "check", "did you") in the last two turns? If so, that's already a signal worth naming out loud rather than plowing ahead.
- Grep your own past sessions for a similar correction:
grep -il "<keyword from the current friction>" ~/.claude/projects/*/*.jsonl— if you find one, name what you did wrong then and what the fix was, briefly, before continuing. - If nothing turns up, say so and proceed carefully rather than pretending a pattern-match happened.
This costs a moment of reflection instead of an automatic injection, but it's the same underlying discipline: don't repeat a correction pattern silently just because nothing is forcing you to check for it.
Verify it's working (if you build the automated version)
grep SMCInjected <your-telemetry-log> | tail -5 # did it fire?
grep SMCCacheHit <your-telemetry-log> | tail -5 # cache hits?
grep FrustrationSpikeDetected <your-telemetry-log> | tail -5
grep DriftClassifierFired <your-telemetry-log> | tail -5
Substitute wherever you're logging telemetry — alignment-harness writes its own telemetry to a JSONL file in the plugin's data folder if you don't have anywhere else.
Adding correction patterns
Each pattern needs: a name, a trigger description (what situation this applies to), the correction directive itself (what to actually do differently), and optionally a leverage score and how often following it has actually helped ("honor rate") once you have enough history to compute one. Match by keyword between the trigger description and what's happening in the session — a new pattern with the right keywords is picked up automatically by whatever matching logic you build.
Kill switch
If this ever produces a bad injection, the fix is always available: don't build the automated hook at all and rely on the local fallback, or if you did build it, gate it behind an on/off switch you can flip without touching the matching logic itself.
Related
/alignment-harness:governer— the risk score that feeds trigger 3.jonathan-check/jonathan-check2/jonathan-check3— predicting how the person would react, a related but different question from "does this match a rule they've stated."smc-evaluator— a second, independent way to score a session against a written rulebook rather than a pattern-match against past corrections.