← the whole session plugin/skills/error-fix-safety-eval/SKILL.md
Safety evaluation for error-log fixes — classifies fix risk by type and blast radius, and determines whether it's safe to fix autonomously or needs a human's go-ahead. Use before writing any fix for an error you found in logs or reports, to avoid hours of rollback when an 'obvious' fix breaks a critical flow.
Error Fix Safety Evaluation
When you find an error (in logs, a bug report, or a crash) and understand the root cause, run this evaluation BEFORE writing code. The goal: never break a user's experience to fix a log entry.
Why this matters
Error logs are symptoms. Fixes change code. Code changes have blast radius. A null guard on a non-critical path is free. A null guard on a payment webhook can silently eat charge failures. Same pattern, wildly different risk. This framework forces you to think about WHERE a fix lands, not just WHAT it does.
Step 1: Classify the fix type
Read the fix you're about to make and classify it into exactly one category:
| Category | What it means | Examples |
|---|---|---|
| Defensive addition | Adds a guard/catch on a path that currently crashes. Existing success path unchanged. | try-catch around an API call, null check before property access, fallback default value |
| Dead code removal | Removes code that provably does nothing. No execution path changes. | Self-referencing map entry (A -> A), unreachable branch, no-op fallback |
| Declaration reorder | Moves existing code earlier/later. No logic change. | Fix a temporal-dead-zone bug by moving a const above where it's used, reorder imports |
| Presentation only | CSS, copy, layout. No data flow change. | Centering a modal, changing a label, adjusting padding |
| Log suppression | Tightens the condition under which an error/warning is logged. UX unchanged. | Only log once data is fully loaded, skip logging a known-transient state |
| Behavioral change | Changes what the code DOES — different return values, different control flow, different API responses. | Changing a gate condition, modifying a response shape, altering auth logic |
| Data flow change | Changes what data moves where — new params, different DB writes, altered event payloads. | Adding a parameter to a function, changing what gets saved to the DB |
Step 2: Identify the blast radius
Where does this code run in the user's journey? Define your own zones for your own product — these are a starting template:
| Zone | Risk level | What usually lives here |
|---|---|---|
| Auth/payment | CRITICAL | Login, registration, token refresh, payment webhooks, subscription checks, access-tier resolution |
| Core product action | HIGH | Whatever your product exists to deliver — the one thing a user came here to do |
| Enrichment | LOW | Analytics, tracking, experiment lookups, metadata fetching — nice-to-have, not load-bearing |
| Admin/internal | MINIMAL | Admin UI, internal tooling, dev-only features |
| Presentation | MINIMAL | CSS, layout, copy, visual-only components |
Example (opt-in — replace with your own): in a coaching product, "core product action" is a coaching conversation actually sending and receiving a response — the one thing the whole app exists to do.
Step 3: Apply the decision matrix
Cross-reference fix type with blast radius:
| Fix Type \ Zone | Auth/Payment | Core Product Action | Enrichment | Admin/Internal | Presentation |
|---|---|---|---|---|---|
| Defensive addition | PROPOSE | AUTO | AUTO | AUTO | AUTO |
| Dead code removal | PROPOSE | AUTO | AUTO | AUTO | AUTO |
| Declaration reorder | AUTO | AUTO | AUTO | AUTO | AUTO |
| Presentation only | — | — | — | AUTO | AUTO |
| Log suppression | PROPOSE | AUTO | AUTO | AUTO | AUTO |
| Behavioral change | PROPOSE | PROPOSE | AUTO | AUTO | AUTO |
| Data flow change | PROPOSE | PROPOSE | PROPOSE | AUTO | AUTO |
AUTO = fix it, commit it, log it. No human needed.
PROPOSE = write the fix plan as a proposal (see /how-to-submit-and-track-proposals if you have it set up, otherwise a markdown file in alignment-harness records proposals works fine). Do NOT implement yet — leave it for a human to greenlight.
Step 4: Verify defensive coding rules
Before committing ANY fix, verify these invariants:
- Fail-open for access: if the fix touches anything that determines what a user can see or do, the failure mode MUST grant access, not deny it. A free user seeing a premium feature for a few seconds is a far smaller problem than a paying user getting locked out.
- No new constraints: the fix must not introduce a condition that could block a user who wasn't blocked before. Example: adding
if (!someId) return res.status(400)is WRONG when a sensible default would let the request through instead. - Existing behavior preserved on the success path: the fix only adds a fallback for the failure path. When everything works normally, the code behaves identically to before.
- No silent data loss: if the fix catches an error, it must not swallow data that downstream code needs. Return a safe default, not
undefined/void.
Step 5: Document your evaluation
Before committing, state in your reasoning (not in code):
Safety eval: [fix type] in [zone] → [AUTO/PROPOSE]
Fail-open: [yes/no — what happens on failure?]
Success path changed: [yes/no]
Blast radius: [what could break if this fix is wrong?]
Quick reference: common patterns
These come up repeatedly. Worth having memorized (the exact field/endpoint names below are one product's; substitute your own equivalents):
| Pattern | Classification | Decision |
|---|---|---|
| Add try-catch to a non-critical API call | Defensive + Enrichment | AUTO |
| Add try-catch to an auth/payment call | Defensive + Auth | PROPOSE |
| Guard a null field with a sensible default | Defensive + Enrichment | AUTO |
| Guard a null field by rejecting the request (400) | Behavioral + Core | PROPOSE (violates fail-open) |
| Remove a self-referencing fallback map entry | Dead code + Core | AUTO |
| Move a variable declaration above where it's used | Declaration reorder | AUTO |
| Change CSS to center a modal | Presentation | AUTO |
| Suppress an error log during a normal loading state | Log suppression + Core | AUTO |
| Change what an auth "who am I" endpoint returns | Data flow + Auth | PROPOSE |
| Add a field to a payment webhook handler | Data flow + Auth | PROPOSE |
| Silence a noisy third-party SDK error | Log suppression + External | AUTO (silencing the log, not changing behavior) |
Integration with error/log triage
If you have a log-triage skill installed (system-log-triage is one example; use whatever you have, or plain reasoning over your own error reports if you don't), this skill runs AFTER that triage identifies a fix, and BEFORE implementation — triage decides WHAT to fix, this skill decides WHETHER you can fix it on your own:
error/log triage → identifies root cause → designs fix
↓
error-fix-safety-eval → classifies fix → AUTO or PROPOSE
↓
AUTO: implement, commit, log it as fixed (a "known-fixed" list of some kind, if you keep one,
helps future triage recognize the same error next time)
PROPOSE: write the proposal, leave the fix for a human
When in doubt
If you can't clearly classify a fix, or if the matrix says AUTO but something feels off, PROPOSE. The cost of a false PROPOSE is one human review cycle. The cost of a bad AUTO is a production incident.