Field note / FNL-02Field noteRev. 01 / 2026

Evidence over certainty.

Every answer should carry its sources.

Confidence is less valuable than traceability. A recommendation should carry the sources, challenges, and assumptions that allowed it to survive.

In June 2023 a federal court sanctioned two lawyers and their firm five thousand dollars for a brief resting on six judicial decisions that did not exist1. The decisions had been produced by a language model, complete with quotations, citations, and the names of real judges. When one of the lawyers grew suspicious and asked the model whether the cases were real, it said they were.

That case is remembered as a novelty. It was the start of a series. A database that tracks only the decisions in which a court itself found reliance on hallucinated material — not accusations, findings — recorded roughly two hundred cases by the middle of 2025 and 1,933 by August 20262. In late 2025 a global consultancy issued a partial refund to the Australian government after a commissioned report was found to contain fabricated references and a fabricated quotation from a federal court judgment, with the firm acknowledging that a generative AI system had been used in its preparation3. None of this is mysterious in origin. When bibliographic citations generated by language models were audited, 55 percent from one model generation were fabricated outright; its successor cut that to 18 percent4. The improvement is real, and it is not the finding. The finding is that a better model reduced the failure without removing it.

Court decisions involving AI-hallucinated material, 2023 to 2026A rising curve of court decisions in which a tribunal found reliance on AI-hallucinated material: near zero in mid 2023, roughly two hundred by mid 2025, then accelerating to seven hundred nineteen in January 2026 and one thousand nine hundred thirty-three by August 2026. The counts are a floor: only cases a court caught and wrote up are included.Q2 2023MID-2025≈200 CASESAUG 20261,933 CASESDECISIONS CITING HALLUCINATED MATERIALA FLOOR, NOT A CEILING — ONLY CASES A COURT CAUGHT AND WROTE UP
Fig. 01 / Hallucinated material in courtCounts from the database, dated

Look at the shape these incidents share. Nothing in the fluency of an answer distinguishes a supported claim from an invented one — the fabricated brief read exactly like a real one, which is what the model was built to achieve. Certainty is a property of the prose. Support is a property of the trail behind it, and the trail is the only place the difference shows.

Provenance is an older discipline than it looks

It would be comfortable to treat this as a machine problem. The record says otherwise. In medicine, a meta-analysis of 46 studies covering some 32,000 quotations found that 16.9 percent of quotations do not accurately reflect the source they cite, roughly half of them major errors5 — written by people, reviewed by people, published anyway. More than ten thousand papers were retracted in 2023, a record6. And when a hundred published psychology findings were systematically re-run, 97 percent of the originals had reported statistically significant results; 36 percent of the replications did7.

The fields that live with this did not respond by asking authors to sound more careful. They engineered the requirement. Medicine grades its confidence in evidence by properties of the evidence — study design, consistency, directness — rather than by the strength of anyone’s assertion8. Data stewardship turned provenance into infrastructure: findable, accessible, interoperable, reusable9, with a web standard for expressing where a piece of data came from and what produced it10. The lesson generalises. Where answers matter, the trail stops being scholarly politeness and becomes a load-bearing requirement.

What carrying sources means

In Ypotheto, a hypothesis’s record keeps its source trail as structure, not as formatting. Every claim in a recommendation carries references to the evidence it rests on — what the source is, where it can be found, and the excerpt that mattered. A claim that arrives with no resolvable source does not survive into the record. That rule is blunt on purpose: the alternative is a record in which supported and unsupported claims are indistinguishable, which is the failure the court cases have in common.

One detail is worth pausing on because it looks pedantic and is not. A citation in the record is keyed to the source it names, never to a position in a numbered list. Reference lists merge, and merged lists renumber; a marker that means “source three” can come to point at a different source without any visible change. A number that silently changes referents is worse than no citation at all — it is a trail that leads somewhere wrong while looking intact. And when a source is later removed, the marker that named it is stripped as a broken link rather than taking the surrounding reasoning down with it. The reasoning may still hold; it simply no longer claims that support.

The claim
One statement, scoped to one criterion, owned by one hypothesis
The sources
Each named and resolvable, with the excerpt that mattered kept beside the reference
The challenges
What was raised against the claim, and what it survived
The assumptions
What the claim rests on that no source settles, stated rather than implied
The basis
How well the claim was measured — computed from the evidence, never written by the model
Fig. 02 / The source trailWhat a recommendation carries

Confidence is authored; measurement is computed

There is a harder version of the problem than fabricated references. A system can cite real sources and still misstate how sure it is. Language models, when asked to verbalise their confidence, tend toward overconfidence — a tendency that persists across elicitation methods and diminishes only partially in stronger models11. Stated confidence is generated text. It is subject to everything that makes generated text fluent, including the pull toward sounding sure.

The design consequence is a split in authority. The model may author a score — the argument for why a candidate looks strong is reasoning, and reasoning is its job. It may not author how well the candidate was measured. That is arithmetic over the evidence: what was actually observed, by what kind of source, and how directly. The basis of every score is computed and is never empty, because an empty basis is where unsupported confidence hides. The same split governs values the system fills in when evidence cannot be found: an estimate is marked as an estimate for as long as it lives, its confidence is capped, and no arithmetic can promote it into a measurement. A guess is not allowed to clear an evidence bar that was set for measurements.

Independence is a property of the data path

Corroboration has its own failure mode, older than any of this: several apparent sources, one origin. Journalism calls it circular reporting. Two reports that agree because both copied the same wire story are one source wearing two mastheads, and agreement between them justifies nothing.

The machine version is quieter. A research system reaches its evidence through tools, and two tools with distinct names can draw on the same upstream source — at which point their agreement is real and worth nothing. So independence is computed from the data path rather than counted from the tools: high confidence requires measured evidence from more than one separate origin, and a model’s own synthesis across tool outputs never counts as corroboration, because it read the same outputs the measurements came from.

Schematic: apparent corroboration against independent corroborationTwo panels. On the left, three tools all draw on a single upstream source, so their agreement counts as one independent source. On the right, two tools each draw on a separate origin, so their agreement is real corroboration. The caption beneath reads: agreement counts only across separate origins.THREE TOOLS, ONE UPSTREAMTOOLTOOLTOOLONE ORIGINONE INDEPENDENT SOURCETWO TOOLS, TWO ORIGINSTOOLORIGINTOOLORIGINTWO INDEPENDENT SOURCESAGREEMENT COUNTS ONLY ACROSS SEPARATE ORIGINS
Fig. 03 / IndependenceSchematic — corroboration needs separate origins

A trail a person can walk

All of this serves one sentence: every claim traces back to a source you can check yourself. The value of the trail is not that every reader walks it. It is that walking it is possible and cheap — the excerpt is already beside the reference, the reference already resolves — so checking a claim costs a minute rather than an afternoon. A trail that is expensive to walk is decoration with better manners.

The trail also changes what disagreement is. When two lines of reasoning collide, sources turn the collision from rhetoric into something tractable: the point where their evidence diverges can be found, read, and judged. That is the subject of the third note.

The lower claim

None of this makes an answer right. It makes an answer checkable, which is the property the confident failures above were missing. A recommendation that carries its sources, its challenges, and its assumptions can be inspected by someone who was not there and defended in front of someone who is skeptical. A confident paragraph can only be believed.

Sources

  1. 01United States District Court, S.D.N.Y.. Mata v. Avianca, Inc., Opinion and Order on Sanctions. 678 F. Supp. 3d 443, 2023. law.justia.com
  2. 02D. Charlotin. AI Hallucination Cases. damiencharlotin.com, database, 2026. damiencharlotin.com
  3. 03Accounting Times. Deloitte to refund government after using AI in $440k report. accountingtimes.com.au, 2025. accountingtimes.com.au
  4. 04W. H. Walters and E. I. Wilder. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 13, 14045, 2023. nature.com
  5. 05C. Baethge and H. Jergas. Systematic review and meta-analysis of quotation inaccuracy in medicine. Research Integrity and Peer Review 10, 13, 2025. researchintegrityjournal.biomedcentral.com
  6. 06R. Van Noorden. More than 10,000 research papers were retracted in 2023 — a new record. Nature 624, 2023. nature.com
  7. 07Open Science Collaboration. Estimating the reproducibility of psychological science. Science 349, aac4716, 2015. science.org
  8. 08G. H. Guyatt, A. D. Oxman, G. E. Vist et al.. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ 336, 924, 2008. pmc.ncbi.nlm.nih.gov
  9. 09M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg et al.. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 3, 160018, 2016. nature.com
  10. 10World Wide Web Consortium. PROV-Overview: An Overview of the PROV Family of Documents. W3C Working Group Note, 2013. w3.org
  11. 11M. Xiong, Z. Hu, X. Lu, Y. Li, J. Fu, J. He and B. Hooi. Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. ICLR, 2024. arxiv.org