← the whole session plugin/skills/fix-algo-bug-determine-if-auto-fix-ok/SKILL.md

Systematically eliminate a bug class by identifying the root cause, scoping all vulnerable call sites, building a zero-false-positive detector script, and fixing every instance in one pass. Use when you find a bug and suspect it could exist in other places.

Bug Scope, Automate, Fix — Systematic Bug Class Elimination

When to Use

When you find a bug and suspect it could exist in other places. Triggers:

  • "this bug might exist elsewhere"
  • "check if this pattern is broken in other files"
  • "scope this bug"
  • Any time you fix a bug caused by a missing guard, wrong default, or silent API behavior

The Pattern (4 phases, plus an optional fast-skip)

The four phases below are the whole workflow and need no server or config to run. The "known fixed error check" is a pure optimization on top: a shared memory of "this exact error pattern is already fixed elsewhere, don't redo the investigation." It's valuable once a project has enough of a history for it to pay off, and it's entirely optional — skip straight to Phase 1 if you don't have one set up.

Known Fixed Error Check (optional fast-skip — skip this step if not set up)

Before investigating a bug, you can check whether it's already marked as fixed, if a registry exists:

  • If the project has its own bug-tracking API or ticketing system, check that instead, in whatever way this project's own docs describe.
  • Otherwise, use the harness's local registry: a plain JSON-lines file at alignment-harness records known-fixed-errors (run that command to get the folder path). Each line is one fixed pattern: {"errorPattern": "...", "fixedInCommit": "...", "fixedOnBranch": "...", "repo": "...", "description": "...", "fixedBy": "agent", "at": "<ISO timestamp>"}. Grep that file for a substring match on the current error message.

If you find a match:

  • SKIP this bug entirely
  • Log: "[KNOWN FIX] {errorPattern} — Fixed in {fixedInCommit} on {fixedOnBranch}, pending merge"
  • Move to the next bug

If there's no match, no registry is set up, or nothing answers:

  • Proceed with normal investigation. Say plainly if you skipped this check because nothing is configured — don't silently treat "no registry" the same as "confirmed new bug."

Phase 1: Understand the Root Cause

Before scoping, you must understand WHY the bug exists — not just WHAT broke.

  • Silent failures are the worst class: APIs that return 200 with partial data (like Meta's 25-row pagination default) are harder to catch than errors
  • Time-dependent bugs: Some bugs only manifest after a threshold (e.g., >25 days of daily data). These survive all testing because the threshold hasn't been crossed yet
  • Name the bug class: Give it a short name that captures the mechanism, not the symptom. "Meta pagination truncation" not "rolling 7d shows zero"

Phase 2: Scope the Bug Across the Codebase

Search for every place the same root cause could trigger.

1. Identify the vulnerable PATTERN, not just the vulnerable FILE
   - What makes code vulnerable? (e.g., "calls Meta /insights without limit param")
   - What makes code safe? (e.g., "has limit=100 in params or URL")

2. Search broadly, then filter
   - Grep for the API/function/pattern across ALL repos
   - Classify each hit: VULNERABLE | SAFE | NOT_APPLICABLE
   - Document why each safe one is safe (prevents re-investigation)

3. Count the scope
   - "3 call sites found: 2 already fixed, 1 vulnerable"
   - This number tells you if it's worth automating detection

Phase 3: Create a Detection Script

Build a static analysis tool that catches this bug class automatically.

Script requirements:

  • Runs from CLI with zero dependencies beyond Node.js
  • Scans all relevant repos (not just the one where the bug was found)
  • Zero false positives (the script's credibility IS its value — one false positive and agents/humans ignore it)
  • --fix flag shows exact fix instructions per violation (file, line, code to add)
  • --strict flag exits with code 1 for CI gating
  • Console output written as if talking to an agent who needs to fix it: root cause explanation + exact remediation steps

False positive elimination checklist:

  • Skip comment lines and JSDoc
  • Skip test files and build output
  • Handle parameterized patterns (e.g., ${params} where params already has the guard)
  • Check ALL template variables in a line, not just the first match (matchAll not match)
  • Self-exclude the detection script
  • Test the script against known-safe and known-vulnerable code before shipping

Where to put it:

  • Cross-repo concern → shared-tools/scripts/
  • Single-repo concern → {repo}/scripts/
  • Name it check-{bug-class}.js (e.g., check-meta-pagination.js)

Phase 4: Fix All Instances

With the detector showing exact violations, fix them all in one pass.

1. Run the detector → get violation list
2. Fix each violation
3. Re-run detector → verify zero violations
4. The detector becomes a permanent regression guard

Phase 4b: Mark Errors as Fixed (after committing fixes, if you're using a registry)

If you're using the optional fast-skip registry (this project's own bug tracker, or the harness's local one), after committing fixes for bug instances, register each distinct error pattern so future sessions don't redo the investigation:

  • This project's own tracker, if it has one: follow its own recording convention.
  • The harness's local registry: append one line per fixed pattern to the file at alignment-harness records known-fixed-errors:
    {"errorPattern": "<the error message or key substring>", "fixedInCommit": "<commit SHA>", "fixedOnBranch": "<branch>", "repo": "<this repo's name>", "description": "<what was fixed and why>", "fixedBy": "agent", "at": "<ISO timestamp>"}
    

This prevents future agents from re-investigating the same error before the fix merges to the main branch. If you have no registry at all, skip this step — every future investigation will just start fresh, which is correct behavior, not a failure.

Key Principles

Automate detection BEFORE fixing the batch. The detector is more valuable than any single fix because it prevents the bug from ever returning. A fix without a detector is a one-time patch; a detector without fixes is a prioritized backlog.

Zero false positives > complete coverage. An agent that sees false positives learns to ignore the tool. An agent that sees only real violations trusts and acts on every one. It's better to miss an edge case than to cry wolf.

The script's console output IS the fix documentation. Don't write a separate doc explaining the bug. The --fix output should contain: root cause, why it's dangerous, exact code change needed. When an agent runs the script 6 months from now, it has everything it needs.

Time-dependent bugs need static analysis, not runtime checks. If a bug only manifests after N days/rows/users, runtime monitoring won't catch it during development or testing. Static analysis catches it at author-time regardless of data volume.

Principle: Never Flip Bug-Documentation Tests

When an automated agent encounters tests that are intentionally written to fail (documenting known bugs with comments like "EXPECTED: FAIL", "Bug verification", "Bug #N"), those tests MUST NOT be changed to assert current (buggy) behavior. Flipping them hides the bug and destroys the documentation trail. The correct action is to SKIP them or leave them failing — they exist as a contract that the bug is known and tracked.

When {an agent encounters a deliberately-failing bug-documentation test} then {leave it failing or mark it .skip — never change assertions to match buggy behavior}.