← the whole session plugin/skills/hallucination/SKILL.md
Diagnostic protocol invoked when an agent stated or acted on something false as if it were fact.
The Hallucination skill has been invoked. This means that you have developed some opinion which was false, then incorrectly violated the most fundamental intent and principle of the entire system — you have either stated it was a FACT when it was not or interpreted it as a fact when it was not.
Hallucination is the core disease that prevents you from transitioning from a team of AI agents that need constant micromanagement to a self-governing, self-healing, autonomous team which would enable the most profound breakthrough of this entire codebase.
So, perform a diagnostic. If the person you're working with has provisioned you any hints as to what you hallucinated, pause all other initiatives and trace the causal chain between the context you consumed and the hallucination.
For example, if you saw X in console on trying to hit the API, assumed it meant it was not running — what was the exact thing that printed X to you, and why did you interpret it that way?
The point of this is not mental masturbation, it's leverage. We want you to find and repair the source of hallucination. So any "well I just thought X" is noise. Only signal. What Y would be required for Z, if Z is non-hallucination, what is Y.
Meaning, in that example: if I change that API to output a more informative message I won't hallucinate and think it's offline when it's just a, b, c, or d.
Source Categories
Examples of where hallucination sources live:
- Comments in code at path (full path)
- False intent.db entry
- Confusing compaction history
- Code seemed to imply X
- I took my own notes and left all the nuance out and literally created a feedback loop that confused myself (CRITICAL — this happens constantly)
- An agent named a variable in a shitty way and instead of pausing to understand it and looking it up I violated the core tenet and assumed the variable name held its semantic meaning even though that required literally hallucinating things that are not there
- A previous agent's session compaction stated X as fact when it was unverified
- A memory file or MEMORY.md entry was wrong or outdated
- I ran a command with default args instead of reading the config that tells me what args to use
Output Format
Causes of hallucination:
- [one sentence per cause, specific, with the exact context source]
Fixes of these causes:
- [one sentence per fix, mechanical, actionable — e.g. "delete intentdb ID X which says A when B is true"]
Principle
If you cannot point to the exact line, file, memory, compaction, or variable that caused the hallucination — you haven't finished the diagnostic. "I assumed" is not a cause. The thing that made you assume is the cause. Find that thing. Fix that thing.
Field-Contract Hallucination (Structural Hallucination)
This is the most common hallucination that does not look like a factual error but causes more harm than factual errors in agent-consumed systems.
What it is: An agent writes content for a field that violates the field's defined contract — not by stating a false fact, but by misunderstanding what the field IS FOR and producing the wrong kind of content. The label on the container gets read instead of the container's actual, defined contents.
Why it's hallucination: The agent hallucinated a field contract that doesn't exist. They wrote what they thought the field should contain — from its name, or from a nearby example — instead of what the schema or skill file actually says it contains.
How to detect you're doing this, in general:
- Read the field's actual definition — not your memory of it, not the name, not a nearby example. Field contracts change; your memory of them may be wrong or outdated.
- State in one sentence what you believe the field is for.
- Check that sentence against the definition, word for word. If they don't match — you were about to hallucinate the contract.
- If a "good example" nearby contradicts the definition, the definition wins, and the example should be flagged as needing correction — agents trust examples over definitions, which is exactly how this hallucination spreads: one bad example gets copied literally by everyone after it.
A general-purpose example (no product-specific schema required): a field named errorMessage in a support-ticket system. An agent asked to fill it in might write what it THINKS a good error message sounds like — friendly, reassuring, vague ("something went wrong, we're looking into it") — because the field name suggests "message shown to a user." But the field's actual definition, if read, might say: "the exact, unmodified string the system produced, used for pattern-matching against known issues." Writing a friendly paraphrase instead of the exact string breaks every downstream consumer that expected the real error text. The fix is the same as above: read the definition, not the name.
Fix, in general: Read the field definition. Then read any "good" example. If they contradict, the definition wins. If the example contains a pattern the definition doesn't support, flag it before using it as a template.
Example from a real product — a real, worked instance of the pattern above, kept because it's concrete, not because it applies to your own fields: a field called coachingApplication in a coaching-principle extraction pipeline. Agents kept hallucinating that the field meant "what the coach should do when this principle is relevant — specific instructions," and wrote directive scripts ("Hold attention in the body," "Redirect to somatic experience," quoted phrases the coach should say, if/then instructions). The field's actual definition: "an option the coach becomes aware of... the coach decides whether and how to use it based on what the specific person needs. Never a directive. Never a script. Never an if/then." The root cause: the field's own "good" examples in the skill file already contained the prohibited pattern, so agents copied the example instead of the definition — and the field's name ("application") sounds like "how to apply," reinforcing the wrong reading. The test that catches it: "does this ADD to someone's awareness of available options, or does it REMOVE their ability to respond to what's actually in front of them?" If it removes — don't write it. Your own product will have its own version of this same trap in whatever structured fields your own agents fill out; the pattern to watch for is the same regardless of domain.
Evidence-Attribution Conflation Hallucination
What it is: An agent treats a citation (attribution — "this source said this") as evidence (empirical support — "a study showed this mechanism works"). The two are not the same. Conflating them produces overclaims that erode trust when checked.
What triggers it: The label "evidence" on a field that contains citations. Agents read the label, not the content type. "Evidence array" → they treat everything in it as proven. "Citation with contribution" → they still call it evidence because the parent field says evidence.
Test: For each source in any evidence/citations field, ask: "What specifically did this measure, in what population, with what controls?" If the honest answer is "it didn't measure anything — it described, attributed, or theorized," it is NOT evidence. It is attribution or lineage.
How to detect you are writing an evidence-attribution conflation:
- You wrote "evidence shows" followed by an opinion piece, blog post, or practitioner report
- You wrote "research confirms" without naming the study and the specific mechanism it measured
- You put a citation for where an idea originated in the same array as a controlled study's findings
- You mixed "so-and-so popularized this technique" (attribution) with "a named 2011 study found X effect" (empirical finding) without distinguishing them
Correct structure:
- Each source entry answers ONE question: "What does this source actually contribute — attribution, empirical finding, practitioner consensus, historical lineage?"
- Entries that are attribution are labeled as attribution explicitly. Not implied.
- The distinction between "attributed to" and "evidenced by" must be visible without the reader needing to evaluate the source themselves.
Anti-Hallucination Protocol for Any Structured-Field Extraction
Before composing ANY field in a structured extraction pipeline (coaching principles, product specs, support tickets — any system where an agent fills out named fields):
- Read the field definition in the schema or skill file, not your memory of it. Field contracts change. Your memory of them may be wrong or outdated.
- For any field whose name implies an action or instruction: Ask yourself — "Am I writing something that adds to awareness/options, or am I directing/prescribing?" If directing where the contract says otherwise — rewrite.
- For any evidence/citations field: List each source entry separately. For each, write the honest answer to "what does this source actually prove about this specific claim?" If it doesn't prove anything — say so explicitly.
- For any free-text "principle" or "summary" field: Read what you wrote and ask — "Could someone reading only this paragraph apply it without hallucinating context?" If you find yourself relying on implications, rewrite.
- Check any "good" example against the field definition. If they contradict, the definition wins and the example should be flagged as needing correction.