← the whole session plugin/skills/validate-against-live-data/SKILL.md

Pattern for validating code correctness by auditing against real production data.

Validate Against Live Data

Tests prove your logic works on data you imagined. This is the method for finding out whether it works on the messy, real data your actual users have: write a read-only script that pulls every real record from wherever the truth about your users lives (a payments provider, an auth provider, a CRM, a production or staging database, even an export), finds the matching record in your app's own database, runs the actual production function against it, and compares its answer to what the real record implies. Start with 20 records to shake out the script, then 50–100, then all, fixing what each round reveals — and keep the script afterward as a regression check.

"Live data" means your own system of truth and your own database, whatever those are. The method doesn't care which; it cares that the check runs the real function on real records, read-only, and reports an accuracy number with every mismatch shown. The examples below use Stripe and MongoDB because that's the stack where this pattern was built and proven — substitute your own systems throughout; nothing about the method requires those two specifically.

When to Use This Pattern

Use this whenever code makes decisions about real users or real money:

  • Access tier resolvers (who gets premium vs free)
  • Billing/entitlement systems (who is paying, who is trialing)
  • Subscription state machines (status transitions)
  • Permission gates (who can access what)
  • Pricing logic (what does a user owe)
  • Any resolver that reads from an external system (Stripe, Firebase, etc.) and makes a determination

The core question: "If I run this function against every real customer, does it return the right answer for all of them?"

Unit tests prove logic works in isolation. This pattern proves it works on the actual messy data in production.

Why Unit Tests Are Not Enough

Real data has properties that test fixtures never capture:

  • Users with multiple Stripe customers (email reuse, test-then-live migrations)
  • Stale DB fields that contradict the external system
  • Edge statuses like incomplete_expired, paused, past_due that nobody wrote tests for
  • Records matched by email fallback instead of primary key
  • External system records with no matching internal record at all
  • Data created by old code paths that no longer exist

You cannot anticipate these in test fixtures. You must discover them empirically.

The Progression Strategy

Round 1: Start Small (20 records)

Purpose: Shake out script bugs and find the most obvious resolver failures.

node scripts/audit-resolver-accuracy.js 20

At this scale you will find:

  • Script connectivity issues (wrong DB, wrong API key)
  • Obvious resolver bugs (entire status categories returning wrong tier)
  • Missing error handling in the resolver itself

Fix everything found. Then proceed.

Round 2: Medium Scale (50-100 records)

Purpose: Find issues that only appear with more data diversity.

node scripts/audit-resolver-accuracy.js 50
node scripts/audit-resolver-accuracy.js 100

At this scale you will find:

  • Multi-entity resolution issues (user has 2+ Stripe customers)
  • Email fallback matches (DB has wrong primary key but email still works)
  • Status edge cases that only exist on a few real accounts
  • External records with no internal match (orphaned Stripe customers)

Round 3: Full Audit (ALL records)

Purpose: Achieve 100% coverage and confirm production-grade reliability.

node scripts/audit-resolver-accuracy.js all
node scripts/audit-resolver-accuracy.js all quiet   # failures-only output for large runs

At this scale you will find:

  • Long-tail edge cases (1 in 200 accounts)
  • Performance issues in the resolver (API rate limits, N+1 queries)
  • The true accuracy number for the system

Each round reveals issues that were invisible at smaller scale. Do not skip rounds.

How to Structure the Audit Script

Template Structure

scripts/audit-<system>-accuracy.js

Every audit script follows this skeleton:

#!/usr/bin/env node
/**
 * <SYSTEM> ACCURACY AUDIT
 * =======================
 * Queries <EXTERNAL SYSTEM> for all real records, matches to internal DB,
 * runs the REAL <resolver/function> against each, compares to expected output.
 *
 * Usage:
 *   node scripts/audit-<system>-accuracy.js           # Last 20 (default)
 *   node scripts/audit-<system>-accuracy.js 50        # Last 50
 *   node scripts/audit-<system>-accuracy.js all       # ALL records
 *   node scripts/audit-<system>-accuracy.js all quiet # Failures only
 */

// ─── STEP 0: Environment Guard ─────────────────────────────────────────────
// ALWAYS verify you are pointing at the right external system.
// For billing audits, REQUIRE a live key — never audit test data by accident.

// ─── STEP 1: Fetch from External System (paginated) ────────────────────────
// Paginate through ALL records. Never rely on a single page.
// Show progress for large fetches.
// Respect rate limits with small delays between pages.

// ─── STEP 2: Fetch Secondary Records ───────────────────────────────────────
// Some systems have multiple record types (subscriptions + one-time purchases).
// Fetch all relevant types.

// ─── STEP 3: Build Audit List (deduplicated) ───────────────────────────────
// Dedupe by primary key (e.g., customer ID).
// For each record, compute the EXPECTED output based on its external state.
// This expected mapping is the test oracle.

// ─── STEP 4: Run the Real Resolver Against Each Record ─────────────────────
// For each external record:
//   a. Find matching internal record (try primary key, then fallback keys like email)
//   b. Run the ACTUAL resolver/function (not a mock, not a simplified version)
//   c. Compare resolved output to expected output
//   d. Log: match/mismatch, evidence, resolver confidence, match method

// ─── STEP 5: Summary Report ────────────────────────────────────────────────
// Total records, found in DB, not in DB
// Correct resolutions (passes)
// Mismatches/errors (failures) — with full detail per failure
// Special categories (multi-entity hits, email fallback matches)
// Accuracy percentage
// Exit code: 0 if clean, 1 if any failures

Key Implementation Details

Environment Guard: For billing/money systems, require the live credential explicitly, and refuse to run with a test credential — you are auditing real data, not test fixtures. Whether this should default to "require production" or allow a staging pass first is a real dial: default to requiring production for money and access decisions, and let the person choose a staging option if they want one first. The Stripe example below is illustrative — substitute your own provider's test/live convention:

// Illustrative — this is Stripe's own test/live key convention, not a universal one
const LIVE_KEY = process.env.STRIPE_SECRET_KEY_PRODUCTION || process.env.STRIPE_SECRET_KEY;
if (!LIVE_KEY || !LIVE_KEY.includes('_live_')) {
  console.error('ERROR: No live credential found. This audit MUST run against real customer data.');
  process.exit(1);
}

Pagination: External APIs paginate. Always auto-paginate to get ALL records. Show progress counters for long fetches.

let allRecords = [];
let hasMore = true;
let cursor = undefined;

while (hasMore) {
  const batch = await externalApi.list({
    limit: 100,
    ...(cursor ? { starting_after: cursor } : {}),
  });
  allRecords.push(...batch.data);
  hasMore = batch.has_more;
  cursor = batch.data[batch.data.length - 1]?.id;
  await delay(100); // Respect rate limits
}

Deduplication: One external entity may appear in multiple record types. Dedupe by the primary identifier before running the resolver.

Matching Strategy: Try the canonical key first, then fall back to secondary keys. Track HOW each record was matched — this itself reveals data integrity issues.

let user = await User.findOne({ stripeCustomerId: item.customerId });
let matchedBy = 'primary_key';

if (!user && item.email) {
  user = await User.findOne({ email: item.email });
  matchedBy = 'email_fallback';  // This is a data integrity signal
}

Run the REAL function: Import and call the actual production resolver. Do not reimplement its logic in the audit script. The whole point is to test the real code path.

const { resolveEntitlementState } = require('../resolvers/gen3/resolveEntitlementState');
const resolution = await resolveEntitlementState(user, { isLoggedIn: true });

Rate Limiting: The resolver itself may call external APIs. Add small delays between resolver calls in large audits to avoid hitting rate limits.

Exit Code: Exit with code 1 if any failures. This lets the script serve as a CI gate or regression check.

What to Look For in Results

Mismatches

The primary signal. A mismatch means the resolver returned a different answer than what the external system state implies. Every mismatch needs investigation — it is either a resolver bug or a wrong expectation in the test oracle.

Not-Found Records

External records with no matching internal record. These may be:

  • Test accounts created directly in the external system
  • Deleted users whose external records were not cleaned up
  • Users who started checkout but never completed registration
  • A data integrity gap worth understanding even if not a resolver bug

Multi-Entity Hits

Users matched via fallback (e.g., email search found a subscription on a different Stripe customer than the one stored in the DB). These indicate:

  • Users who have created multiple external accounts
  • Migration artifacts (test-mode customer vs live-mode customer)
  • The resolver's fallback logic is doing real work on real data

Email Fallback Matches

Records found only by email, not by the stored primary key. This means the DB has stale or wrong primary keys. The resolver may be compensating, but the underlying data should be fixed.

Resolver Errors

The resolver itself throws an exception for certain records. These are bugs that only surface on specific data shapes.

Handling Environment Differences

Test vs Live Keys

Many external systems have separate test and live modes (Stripe, Firebase, etc.). The audit script must be explicit about which mode it targets.

For billing audits: always require the live key. Auditing test data tells you nothing about production accuracy.

For non-billing systems: you may want to audit both. Run the script twice with different env vars.

Database Environment

The audit connects to a real database. Ensure MONGODB_URI (or equivalent) points to the right environment. For production audits, this should be the production database.

Safeguards

The audit script is READ-ONLY. It must never write to the database or external system. It only reads and compares. This makes it safe to run against production at any time.

Check this, don't just assert it. Before running a script like this against a live source, read back through it yourself for any write call — database updates/inserts/deletes, or POST/PUT/PATCH/DELETE calls to the external API. If you find one, remove it or get the person's explicit approval before running against production. The whole pattern is only safe because it never writes, and that shouldn't rest on nobody double-checking.

Use a read-only credential if the system supports one, and say so to the person if you're about to run with a broader-permission credential than that.

If you don't have live access yet

The whole point of this pattern is the person's own real data — nothing about any other project's setup can stand in for it. If there's no credential available, or you can't identify which function actually makes the decision in question:

  • Say plainly what this is for: "Tests only prove your logic on data we imagined. This runs your real function against every real customer, read-only, and tells you how often it gets the answer right — with every miss shown. To do that I need a read-only way into where the truth about your users lives, and I need to know which function you want checked."
  • Look for what's already inspectable: payment/auth/access code in the repo, the environment variable names it expects (from .env.example or config — names only, never values), and past sessions where the person debugged billing or access issues (~/.claude/projects/*/*.jsonl) — these usually name the function and the edge cases that matter.
  • Propose the function, the source of truth, and the expected-answer rule (the "oracle" — e.g. "an active subscription should resolve to paid"), and ask the person to confirm each, and to supply a read-only credential themselves.
  • If nothing is available at all, run against whatever sample is available — a staging database, a data export, or a handful of records the person pastes in — and label the result plainly as "not production; accuracy on N sample records," never as a verdict on real users. If there's truly nothing to run against, say the claim "this works on real users" is unverified and name exactly what would verify it.

After the Audit

The Script Becomes a Permanent Tool

Once the audit script works and shows 100% accuracy:

  1. Keep it in the repo as a regression tool
  2. Run it before deploying resolver changes to catch regressions
  3. Run it periodically (monthly or after data migrations) to catch drift
  4. Use it as evidence when someone asks "can we trust this code?"

Documenting Results

After a full audit pass, the summary is your proof of correctness. This is an illustrative shape with invented numbers — yours will reflect your own system:

Total records in system of truth:   267
Found in your database:             241
Not found in your database:          26
Correct resolutions:                241
Mismatches/Errors:                    0
Accuracy:                           100%

This is stronger evidence than any number of unit tests.

Keep a dated record of each run. Save the summary (counts, accuracy, mismatches with evidence) under alignment-harness records live-audits, named for the date and the function audited, so "can we trust this code?" has an answer with a date on it, and a later run can show whether accuracy drifted.

Example (opt-in reference — the shape to copy, not a requirement)

One reference implementation audits an entitlement-state resolver against every real Stripe subscription and lifetime purchase in a real product:

  • Supports 20 / 50 / 100 / all / all-quiet modes
  • Requires a live Stripe key (refuses to run in test mode)
  • Finds users by their Stripe customer id, with an email fallback
  • Runs the real resolver function against each user
  • Reports mismatches, not-found, multi-customer hits, and email fallback matches
  • Exits with code 1 on any failure

Running it for real found 3 genuine bugs across 4 rounds: an unhandled status value at round 1 (20 records), a need for multi-customer resolution discovered at round 2 (50 records), and confirmed 100% accuracy after fixes at the full round. That progression — small round finds the obvious bug, medium round finds the structural one, full round confirms — is the part worth copying; the specific resolver and provider are just one project's stack.

Checklist for New Audit Scripts

  • Environment guard: require the correct API key mode
  • Paginate through ALL external records (not just first page)
  • Fetch all relevant record types from the external system
  • Dedupe by primary identifier
  • Compute expected output for each record (the test oracle)
  • Find matching internal record with primary key + fallback strategy
  • Run the REAL resolver/function (not a reimplementation)
  • Compare resolved vs expected, log full evidence on mismatch
  • Track match method (primary key vs fallback) as a data quality signal
  • Add rate limiting delays for large audits
  • Show progress for large runs
  • Support quiet mode (failures-only output)
  • Summary with counts: total, found, not-found, pass, fail, accuracy %
  • Exit code 1 on any failure
  • Script is purely read-only (never writes to DB or external system)