← the whole session plugin/skills/validate-load-bearing-claims-against-reality/SKILL.md

Use when you are about to report or act on a load-bearing conclusion — especially one a governer scored 90+ — that you reached by reading code, docs, or any single source of inference. This skill is the discernment that turns operating-certainty into validated truth by measuring the actual reality your claim describes. Triggers — high-leverage conclusion, "the code shows", "the DB does X", "this resolves to", "the gate fires", reporting a finding you have not measured against the thing it is a claim ABOUT, any 90+ governer task before the completion claim.

Validate Load-Bearing Claims Against Reality

The target — what a 10/10 looks like

The intent of this skill is one capability: when a conclusion carries real weight, you reach the state where two fully independent perspectives on the same truth agree — what the code/doc CLAIMS is happening, and what reality ACTUALLY CONTAINS when you measure it directly — so that a false conclusion becomes nearly impossible, and you can hand the person you're working with a claim that is not "my careful reading says" but "I measured it and reality confirms it."

Picture the finished state precisely. You have a conclusion — say, "the database does not override the live coaching prompt." You have arrived at it carefully: you traced the code end to end, you ran adversarial checks, you even attached honest confidence scores and named what you had not yet verified. That is good work. It is also, still, one perspective — the code's account of itself. The 10/10 is the move you make next: you find the cheapest available window into the actual reality the claim is about — the real persisted value, the live response, the stored record on a real object that a real user touched — and you measure it against what the code said. When the stored prompt on a live conversation comes back byte-identical to what the code generates, your certainty stops being a careful inference and becomes a measured fact. That is the target: certainty that has been validated or rejected from fully outside itself.

The reason this matters more than any rule about it: a load-bearing claim that is only code-deep is the inception point of compounding hallucination. You report it, the person builds a decision on it, the next agent inherits it as settled, and three steps later real decisions are being made on something nobody ever measured. The single measurement — often one query, often sixty seconds — is what severs that chain at its root. The gravity is not "you might be wrong." The gravity is "the real data was one cheap query away and you chose your own certainty instead." When you can see that clearly, you will not stop at the code, because stopping will feel like what it is: leaving the most decisive evidence on the table.

Why flagging the gap is not the same as crossing it

There is a subtle, important distinction the worked example below turns on. It is possible to do everything almost right — to reach a careful conclusion, to honestly label it "not yet runtime-verified," to even offer the reality-check as a suggested next step — and still fall short of the target. Naming a gap is not closing it. The honest caveat protects you from claiming false certainty, but it does not give the person the validated truth they actually need; it hands them the unfinished work and the obligation to ask for the last step. At high leverage, the target is not "be transparent that I haven't measured it." The target is "measure it." Transparency about an un-crossed gap is the floor. Crossing the gap is the intent.

The worked example — a real shape of near-miss, genericized (originally 2026-05-30)

This skill exists because of a specific real moment on a real project; the details below are genericized (the real incident named a specific product's internal data and internal names, which don't belong in a shared plugin) but the shape — and the lesson — is exactly what happened.

The person had a scary suspicion: that a piece of runtime configuration actually in effect was being silently overridden by a database-stored version, different from the code that had been carefully aligned to a known-good baseline — meaning the system might currently be running on the wrong version of itself without anyone knowing. The agent took it seriously and did genuinely rigorous work: a multi-agent adversarial workflow, one side building each claim and the other side prompted to refute it, the whole code path traced from the entry point through to where the configuration is actually used. It concluded, at ~90% confidence with the adversarial side conceding: the live behavior comes from the source files; the admin editing tool is a separate workbench that doesn't feed the live path; the database does not override the code. It reported that to the person with explicit confidence scores and explicit "NOT runtime-verified" caveats, and offered the reality-check as a next-step command.

That was still a failure of duty. Not because the conclusion was wrong — it turned out to be right — but because the agent stopped at the code's account of itself and handed the person its certainty instead of measured truth, when the measurement was trivially available. The person named the gap exactly: "you could pull up a real record and find the actual value attached to it, and look at what the code says it should be, and see if it's the same thing… that completely different perspective about hard reality, measured against what the code is saying — that's a big gap I never see you think about."

So the agent crossed it. One small script, about sixty seconds: it pulled a real, live record's actually-stored value and diffed it against what the code generates for the same case. The two were byte-identical at the head, including a distinctive quirk (a specific misspelling or unusual phrasing) that was present in both the stored value and the code — strong evidence they came from the same generation path, not independent copies. The only divergence was an expected per-instance injected middle section. The agent then went further and pulled several older, historical records from a known-good period — and they confirmed the same generation mechanism was running then too, with the only difference being a labeling detail that changes intentionally per instance (like a version tag or a named variant). That last detail matters enormously: a naive "diff one code output against one stored record" would have screamed MISMATCH at that labeling detail and manufactured a false alarm. Measuring multiple real records revealed that detail was a legitimate per-instance variable, not a divergence. Reality didn't just confirm the conclusion — it refined it in a way code-reading alone would have gotten wrong in both directions.

The lesson is not "the agent was almost wrong." The lesson is the shape of the discernment: the conclusion and the measurement are two different perspectives on one reality, and at high leverage your job is to make them meet.

What this looks like as a way of seeing (not a checklist)

When a claim is forming and it carries weight, hold two questions in view at once. First: what is this claim actually ABOUT — what real thing in the world would be true or false if I'm right? (A persisted field. A live API response. A row in a collection. A rendered DOM. A real user's actual state.) Second: what is the cheapest honest window onto that real thing? The gap between "what the code says it does" and "what the thing actually contains" is the gap you are responsible for. Most of the time the window is cheap — a query, a curl, a single read of a real record. When it is cheap and the claim is load-bearing, the only coherent move is to look through it. When it is genuinely expensive or impossible, that is when you report the claim as inference with its confidence and its named gap — but only after you've confirmed the gap can't be cheaply closed, not as a substitute for the looking.

Finding where "reality" lives, on a project you don't know yet

This skill's whole method depends on having some cheap, honest window onto the real thing a claim is about — a database, a staging or production URL, a log stream. On a fresh project you may not know what those are yet. Before assuming none exist:

  1. Look at what's already there: .env/config files for connection strings and service URLs (read them for the names of what's configured — which database, which environment — never echo secret values), package.json scripts (a db:console or logs:tail script is a strong hint), existing test fixtures, and your own past sessions in this project (~/.claude/projects/*/*.jsonl) for databases or endpoints referenced before.
  2. Confirm with the person before using any of it: "I found what looks like a database connection and a staging URL — can I use those read-only to check claims against real data?"
  3. If nothing turns up and the person doesn't know either, that's a real gap, not a failure — report the claim as inference, state your confidence, and name specifically what you'd check if the window existed ("I could verify this by querying X, but that isn't set up yet"). Never present that inference as though it had been measured.

The governer duty — mandatory at score ≥ 90

When the governer scores a task 90 or above, this validation is not optional polish; it is a duty owed to the governer contract, because at that leverage a false load-bearing claim compounds across every agent and decision downstream. The intent of making it mandatory is to guarantee that the most consequential claims in the system are the ones held to the highest standard of evidence — measured reality, not careful inference.

Concretely, before reporting completion or a load-bearing finding on any ≥90 task, the agent makes the validation visible to the admin:

  1. Name the load-bearing claim and the reality it is about. State, in plain language, the one-or-few conclusions the person's decision will rest on, and name the actual real-world thing each is a claim about (the persisted value, the live response, the record). This is the intent-surfacing step — it makes the gap explicit before it's crossed.

  2. Show the measurement that crossed the gap. Produce the actual evidence that you measured reality against the claim — the query you ran, the real record you pulled, the diff you computed — and what it showed. Two perspectives, shown to agree (or shown to disagree, which is even more valuable). This is the difference between "my reading says" and "I measured it."

  3. If you did not cross the gap, name it as a duty failure and cross it now. If you find yourself about to report a ≥90 load-bearing claim on inference alone when the reality was cheaply measurable, recognize that as a failure of duty to the governer — not a small one — and go measure it before reporting. Make that visible too: the admin sees that the gap existed, that it was caught, and how it was closed. The point is not self-flagellation; it's that the system demonstrates the discernment working rather than hiding the near-miss.

The aim of all three is a single experience for the person: when an agent hands them a high-leverage conclusion, they can trust it is validated against reality, and they can SEE the validation, without having to be the one who remembers to ask "but did you actually check?"

How this composes with the rest

This is downstream of /governer's overallGovernerScore (score at or above the mandatory threshold below makes this mandatory) and pairs with /verification-gate and /hallucination (which it operationalizes for the specific case of inference-vs-measurement on weighty claims). Nothing currently makes /governer call this skill automatically — until that two-way wiring exists, invoking this skill at 90+ is on you. It is the same root principle as the standing instruction, from whoever you're building alignment with, that nothing is a fact until reality — not a doc, not a confident-sounding trace — confirms it. When in doubt about whether a claim is "load-bearing enough," ask whether anyone will build on it; if yes, it earns the measurement.

If you actually internalized this: the next time you catch yourself writing "the code shows X" about something that matters, you'll feel the pull to go open the real thing and check — and that pull, not any rule on this page, is the skill.