← the whole session plugin/skills/pipeline-cross-verify/SKILL.md
Independent cross-verification of another agent's verification output. Receives ONLY the verification contract and code diff/file references — NOT the first agent's reasoning. Independently runs its own commands, grades each claim PASS or FAIL with evidence, and writes a cross-verify report. The evaluator's finding stands even when it contradicts the first agent's conclusion.
/pipeline-cross-verify — Independent Cross-Verification
This is a pipeline stage: it runs after some prior stage has produced a verification contract (a claims list with evidence and certainty), and it is INDEPENDENT — it grades that output without access to the prior stage's reasoning. If you're using a multi-stage observable pipeline (see
pipeline-stage-gate,pipeline-adversarial-design,pipeline-compliance-auditfor the rest of that family), this is one stage in it. It also works standalone: any time one agent's claim is about to become the next agent's assumed fact, dispatch a second agent with only the claim and the evidence, never the first agent's narrative, and let its verdict win.
Why this stage exists
Cross-verification by accident — two agents happening to process the same thing and their outputs getting compared informally — catches some mistakes but is unreliable and untracked. Making it a designed role fixes that: a dedicated evaluator agent
- receives the verification contract only — not the first agent's reasoning chain,
- runs its own independent commands against the codebase,
- grades each claim on its own evidence,
- reports discrepancies between what the first stage concluded and what the evaluator found.
The evaluator's job is not to audit the first stage for process compliance. It is to independently verify the facts that stage claimed to verify. If it said "confirmed" and the evaluator finds different evidence, the evaluator's finding stands — and the discrepancy is the most important output of this stage.
What the evaluator receives
ONLY:
- The verification contract — the claims the prior stage checked, in this format:
evidence_summary: - assumption: "..." verdict: confirmed|disproven|inconclusive evidence: "..." certainty: "N%" - File references — the specific files and line numbers cited. No commentary about what they mean.
- The diff (if code changed) — the actual diff, nothing else.
The evaluator does NOT receive: the prior stage's reasoning chain or narrative, its conclusions about what the evidence means, earlier-stage context (unless a file reference points there), or any prior agent's framing of the issue.
This information boundary is intentional. The evaluator must form its own view from raw evidence, not audit the prior stage's interpretation.
Execution protocol
Step 0: Orientation — without the prior stage's context
Before running any commands, read each claim and ask:
- What would I need to find in the code to independently confirm or deny this?
- What command would produce the relevant output directly from the source?
- What would I expect to see if this claim were FALSE?
Write this pre-read orientation as a brief section in the output — it shows your independent framing BEFORE you look at the files the prior stage referenced.
Step 1: Verify each claim independently
For each claim:
- Run a command that would confirm OR deny it — independently of whatever the prior stage ran. If it used
grep -n "X" file.js, consider a different angle that would catch the same behavior. If it read lines N-M, read a broader range that includes context it may have excluded. If it used a memory/institutional search tool, run a different query that probes the same domain. - Report the actual command and actual output — never a summary:
Command: <exact command run> Output: <verbatim output, truncated only if over 40 lines> - Grade the claim:
- PASS — your independent evidence confirms the prior verdict.
- FAIL — your evidence contradicts the prior verdict, or you can positively show it's wrong.
- PARTIAL — you confirm the claim exists but at a different scope, line, or severity than stated.
- CANNOT VERIFY — the referenced file/line doesn't exist, or the command produces no output.
Important distinction: a broken citation is not the same as a disproven claim. If a file reference doesn't resolve, that means the prior stage cited something wrong — grade it CANNOT VERIFY (see the rules below on why this isn't neutral), not FAIL as if you'd found positive evidence against the underlying claim. Reserve "the claim is actually false" for cases where you found real evidence contradicting it, not just a broken pointer to where it was supposed to be.
Step 2: Run the known-failure-mode checklist
These are general classes of mistake worth checking for on every claim (one real pipeline found each of these actually happening in real runs — rename them to your own project's incident history once you have one):
Stale data — claim was true when filed, code has since changed. For any claim referencing file:line:
git log --oneline -5 -- <file>
If the file was modified after the claim's notice date, the prior stage may have verified old code. Flag as STALE RISK.
Wrong file/line reference — the cited lines don't contain what was claimed. Read the exact lines cited. If claim says "file.js:47 — missing null check on req.user," read line 47: is req.user accessed there without a guard? If the code doesn't match the description, grade FAIL (a genuine mismatch you found, not just a missing file) regardless of whether something similar exists elsewhere in the file — grade on what was actually cited.
Feature already exists — claim says something is missing that's actually present. For any claim in the form "X is not implemented" or "Y is missing":
grep -rn "relevant_term" <relevant directory> --include="*.<your language's extension>" | head -20
Search broadly, using your project's actual language(s), not just one. If the feature exists somewhere other than where the prior stage looked, grade PARTIAL and say where it lives.
Rubber stamp detection — the prior stage confirmed everything. If its evidence_summary has all verdict: confirmed and certainty ≥90% on every claim, treat that as a yellow flag worth a closer look, not proof of a problem — genuinely correct claims do happen. For the highest-certainty confirmed claims, run an explicit negative check: "what would I expect to see if this claim were FALSE?" — then look for that, and report what you found either way.
Step 3: Evaluate for discrepancies
- No discrepancies — your evidence matches the prior verdicts. Note which claims were cross-confirmed.
- Partial discrepancies — some claims verified at a different scope or with different line refs. List each with your evidence vs the prior evidence, side by side.
- Material discrepancy — your evidence contradicts the prior verdict on a claim that materially affects the outcome (a "confirmed" becomes FAIL, or a "disproven" becomes PASS). This is the most critical output. When this happens: state it explicitly in a prominent section, don't try to reconcile the two views — report what you found — and flag whether it affects the overall recommendation.
Step 4: Self-audit before finalizing
Sanity check:
- Did I actually run the commands, or did I summarize what I expected to find?
- Is every grade backed by verbatim command output in this document?
- Did I check every failure mode for every claim?
- Did I treat the prior stage's certainty assertions as claims to be verified, not established facts?
Reflect:
- Am I grading based on the evidence I found, or deferring to the prior stage's framing?
- If a completely fresh agent saw this for the first time, what would it find?
- Is my "CANNOT VERIFY" actually a citation failure I'm softening into something gentler?
Output format
# Cross-Verify — {item title}
**Run ID:** {run-id}
**Date:** {date}
**Evaluator:** independent — did not read the prior stage's reasoning before running commands
---
## Pre-read orientation (before looking at the referenced files)
{What you expect to find based on the claim descriptions alone. Written before Step 1.}
---
## Independent verification per claim
### Claim 1: {claim description from evidence_summary}
**Prior verdict:** {confirmed|disproven|inconclusive} at {certainty}%
**Evaluator command:**
{command}
**Output:**
{verbatim output}
**Evaluator grade:** {PASS|FAIL|PARTIAL|CANNOT VERIFY}
**Evaluator finding:** {1-2 sentences on what the evidence shows}
**Discrepancy:** {none|describe discrepancy if grade contradicts prior verdict}
### Claim 2: ...
---
## Known-failure-mode checklist
| Failure mode | Checked | Finding |
|---|---|---|
| Stale code (file modified after notice date) | yes/no | {git log output or "no recent changes"} |
| Wrong file/line (cited lines don't match description) | yes/no | {finding or "lines match"} |
| Feature already exists | yes/no | {finding or "not applicable"} |
| Rubber stamp (100% confirmed, no negative checks run) | yes/no | {finding or "negative checks ran — see claims above"} |
---
## Discrepancy summary
**Overall agreement:** {full agreement | partial discrepancy | material discrepancy}
{If partial or material: list each discrepancy with prior verdict vs evaluator finding, side by side.}
{If material discrepancy: state explicitly whether it affects the overall recommendation.}
---
## Self-audit
**Commands run (not summarized):** {yes/no}
**Known failure modes checked for every claim:** {yes/no}
**Grade based on evaluator's own evidence (not prior framing):** {yes/no}
---
## Recommendation
{One of:}
- **ADVANCE** — evaluator confirms prior verdicts. May proceed.
- **ADVANCE WITH CORRECTIONS** — partial discrepancies found. List corrections needed before proceeding.
- **RETURN FOR RE-VERIFICATION** — material discrepancy found. Must re-verify with the evaluator's evidence before proceeding.
- **CLOSE AS DISPROVEN** — evaluator's evidence positively contradicts the confirmed claim (not merely a broken citation) and finds no user impact.
Key rules
- The evaluator's finding stands. If the prior stage said "confirmed" and the evaluator finds different evidence, the evaluator's grade is the output — not a compromise between the two.
- Grade on what was cited, not what might be true. If line 47 was cited and doesn't contain what was claimed, that's FAIL — even if the behavior exists on line 83.
- Never access the prior stage's reasoning chain. If you catch yourself using its narrative to interpret your command output, stop. Return to the raw output and grade independently.
- CANNOT VERIFY is not neutral. When a file reference doesn't resolve, that's a failure by whoever cited it — grade accordingly. A claim grounded in a non-existent source is not inconclusive, it's ungrounded. But it is also not the same as a positively disproven claim (see Step 1) — don't let "I couldn't verify this" quietly become "this is false."
- Known failure modes are not optional checks. Run every one for every claim.
- Material discrepancy is the most valuable output. A clean "no discrepancy" report is useful. A "the prior stage said confirmed but here's what the code actually shows" report is the reason this stage exists.
Compliance trail (worth including in every output)
- What institutional-memory/search queries were run (if any) and why.
- Whether git log was checked for file modification dates.
- Which failure modes were checked and for which claims.
Tracking your own baseline over time
If you run this repeatedly, two numbers are worth watching:
- How often the prior stage's confirmed claims actually hold up. One real pipeline settled around a 40-70% confirmation rate as "normal" for its process — that's an observed baseline from that history, not a universal target. You won't have a baseline on day one; track your own as runs accumulate, and treat a suspiciously perfect rate (near 100%) as worth a closer look rather than good news by default.
- How often this stage agrees with the prior one. Above roughly 90% agreement across 5+ runs is worth treating as a yellow flag the same way — either the pipeline is genuinely reliable, or this "independent" check is quietly deferring to the first stage's framing despite the information boundary. Log it wherever you're keeping records (
alignment-harness records pipeline-cross-verifyworks if you don't have anywhere better) so the pattern is visible over time, not just per-run.
Lighter-weight use
You don't need the full report format above every time. For a quick spot-check on any subagent's output before it moves on to the next stage of whatever you're doing: sample 3-5 claims, trace each to its actual source, and if any claim can't be traced, treat that stage's output as not yet fit to pass forward.