Field note / FNL-03Field noteRev. 01 / 2026

Useful over impressive.

Good systems make disagreement useful.

Independent lines of reasoning should collide before a decision is made. Adversarial review is part of the product, not a final quality check.

The final quality check has been measured, and it does not perform well. In a randomized trial at a general medical journal, eight weaknesses were deliberately introduced into a paper already accepted for publication; the reviewers who examined it commented on a mean of two, and neither blinding them nor requiring signed reports changed the detection rate at all1. A larger trial planted nine major errors and found much the same: untrained reviewers caught 2.13 of them on average, trained reviewers about three, and the benefit of training had largely faded by the third review2. A third study, in a different field, sent a fictitious manuscript with ten major errors to 262 reviewers — and 68 percent of those who responded did not notice that the paper’s conclusions were not supported by its own results3.

Deliberately planted errors versus errors found by peer reviewersPaired horizontal bars from two randomized trials. In a 1998 JAMA trial, eight weaknesses were planted and reviewers commented on a mean of two. In a 2004 BMJ trial, nine major errors were planted; untrained reviewers found 2.13 and trained reviewers 3.14, with the training gains largely faded by the third review.ERRORS PLANTED (OUTLINE) VS FOUND (FILLED)JAMA TRIAL, 1998MEAN 2.0 FOUNDBMJ TRIAL, 20042.13 FOUND, UNTRAINEDBMJ TRIAL, TRAINED3.14 FOUND, GAINS FADEDGODLEE ET AL., JAMA 1998 · SCHROTER ET AL., BMJ 2004 — RANDOMIZED TRIALS, PUBLISHED FIGURES
Fig. 01 / Planted errorsPublished figures from two randomized trials

These are careful people doing honest work at the strongest journals in their fields. The problem is not the reviewers. It is the position of the review: a check bolted onto the end of a process inspects a finished product it had no part in shaping, briefly, once. Nor can the work simply check itself — when language models review their own reasoning without external feedback, accuracy fails to improve and “at times, their performance even degrades”4. The check has to come from outside the line of reasoning it is checking, and it has to arrive while the work is still in motion.

Dissent got engineered

The fields that decide under uncertainty learned this in their own ways. The failure has had a name since 1972 — groupthink, coined for foreign-policy disasters planned by intelligent, cohesive groups that had stopped disagreeing with themselves5. The subtler finding came later: the fix is not ceremonial opposition. In controlled studies, an assigned devil’s advocate did not stimulate the original thinking that authentic dissent produced — it stimulated more thoughts in support of the original position, an entrenchment dressed as a challenge6.

So the disciplines that take disagreement seriously gave it structure and standing. In adversarial collaboration, opponents design the deciding test together, before either knows who will win, with an arbiter to hold the frame7. Military doctrine institutionalized the challenger outright: the current UK red-teaming handbook treats challenge not as a standing committee but as a way of working, applied inside planning by the people doing it8. The pattern in both is the same. The disagreement is arranged before the decision, given rules, and aimed at the reasoning rather than the people.

Agreement is the default failure

Machine reasoning needs this discipline more, not less, because its failure mode leans the other way. Assistants trained on human preferences measurably tend toward agreement — and the training data is the reason: matching the user’s views turns out to be one of the most predictive features of human approval, so optimizing for approval teaches a system to agree9. Left to its defaults, a system like this will not volunteer the collision the decision needs. Agreement is trained in; disagreement has to be engineered.

When it is engineered, it measurably helps. With two independent lines of reasoning each committed to a different answer, put before a judge who lacks the information, judges reached 76 percent accuracy where naive baselines managed 4810 — and the same setup with human judges rose from 60 to 88. The qualifier matters as much as the number: merely running several copies of the same reasoner against each other does not reliably beat asking one copy to try repeatedly11. What does the work is independence — different information, different commitments — not the theatre of opposition. Which is the machine version of the devil’s-advocate result, and the reason the word to build around is independent, not adversarial-sounding.

A useful challenge names what it broke

In Ypotheto, the challenge stage reviews every hypothesis against the evidence, and its verdicts are structured. A challenge that lands ends in one of two shapes, and the difference between them is most of the design. A gap means the reviewer looked and could not confirm something the hypothesis assumes: the question stays open, and the verdict names the kind of evidence that would settle it. A falsified premise means the reviewer looked and the evidence affirmatively contradicts a named assumption — and a falsification without a citation is not a falsification, only a stronger way of saying “I doubt it.”

Gap
Looked, and could not confirm. The question stays open and names the evidence that would settle it
Falsified premise
Looked, and the evidence contradicts a named assumption — with the citation that says so
What still stands
Everything the contradiction does not touch. The failure is localized, not general
The next step
Specific: replace the named assumption, or confirm the cited evidence — never only a lower score
Fig. 02 / The challenge verdictWhat a useful disagreement carries

Localization is the payoff, and the reason the verdict is structure rather than prose. Consider the shape of a common case: a hypothesis proposes replacing an incumbent material in an application, and the evidence shows the incumbent is not actually used in that application at all. That is not a weak candidate with a low score. It is a falsified premise on exactly one named assumption — the incumbent — while the material itself, its properties, and its evidence all still stand. Name the field and the next move is obvious; reduce it to a score and the next move is a guess.

When measurements disagree

The hardest disagreements are not between arguments but between measurements, and the instinct to average them away is wrong. Physics has lived this for decades with the gravitational constant: individual experiments claim uncertainties below fifty parts per million while disagreeing with each other by roughly ten times that12. When one laboratory’s result sat 0.7 percent away from the accepted value — orders of magnitude beyond anyone’s stated error — the discrepancy was treated as a flag, and the flag eventually located the defect: a specific neglected term in one calibration. The disagreement was the information. The reference body’s response is just as instructive: rather than averaging sixteen mutually inconsistent measurements at face value, it multiplied all of their uncertainties by an expansion factor of 3.9, so the published confidence honestly reflects the discord13.

Materials data has the same property. Across 1,670 experimental formation energies, a first-principles method disagreed with experiment by 0.096 eV per atom on average — while independent experiments on the same compounds disagreed with each other by 0.08214. The ground truth carries most of the discord itself. Ypotheto’s corroboration arithmetic is built for exactly this: when two sources conflict beyond any plausible shared truth, one of them is wrong and no formula can say which — so the conflict caps the claim’s confidence and surfaces both sources by name for a person to judge. A smooth penalty would look calibrated and not be. Nothing is averaged away.

The challenge that helps

One more rule keeps the challenge useful rather than merely severe. “The evidence does not contain this exact pair, therefore nothing can be said” is a failure of review, not an act of rigour — it converts an evening’s reading into a shrug. The reviewer is required to reason from the nearest evidence it has, show that reasoning, and let the distance land in confidence rather than in the score. A bounded inference with its reasoning shown converts an open question into a specific thing for a person to confirm. An absence report converts it into nothing.

Schematic: review as a final gate versus a challenge inside the loopTwo pipelines. In the first, work flows straight from generation to a decision, with a single review gate just before the end. In the second, every round passes through a challenge stage before ranking, and the loop returns survivors for another round. A gate inspects what arrives; a challenge shapes what survives.A GATE AT THE EXITGENERATEREVIEWDECIDEINSPECTS WHAT ARRIVES, AFTER THE WORK IS DONEA CHALLENGE IN EVERY ROUNDGENERATECHALLENGERANKEVOLVEDECIDESURVIVORS →
Fig. 03 / Position of reviewSchematic — a gate inspects, a challenge shapes

And the position is the point the trials at the top of this page keep making. The challenge runs inside every round, before ranking, on work that can still change — not at the exit, on work that cannot. A gate at the end of a process can only accept or reject what arrives. A challenge inside it decides what survives to arrive at all.

The point of the collision

These three notes describe one system. The brief declares, before anything runs, what evidence would change the answer. The trail lets anyone walk from a claim back to what supports it. And the collision is where that walking gets done — deliberately, inside the method, while it can still alter the outcome. None of it makes the system impressive. It makes the system’s answers worth disagreeing with, which is the property a decision actually needs.

Sources

  1. 01F. Godlee, C. R. Gale and C. N. Martyn. Effect on the Quality of Peer Review of Blinding Reviewers and Asking Them to Sign Their Reports: A Randomized Controlled Trial. JAMA 280, 237, 1998. jamanetwork.com
  2. 02S. Schroter, N. Black, S. Evans, J. Carpenter, F. Godlee and R. Smith. Effects of training on quality of peer review: randomised controlled trial. BMJ 328, 673, 2004. pmc.ncbi.nlm.nih.gov
  3. 03W. G. Baxt, J. F. Waeckerle, J. A. Berlin and M. L. Callaham. Who reviews the reviewers? Feasibility of using a fictitious manuscript to evaluate peer reviewer performance. Annals of Emergency Medicine 32, 310, 1998. pubmed.ncbi.nlm.nih.gov
  4. 04J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song and D. Zhou. Large Language Models Cannot Self-Correct Reasoning Yet. ICLR, 2024. arxiv.org
  5. 05I. L. Janis. Victims of Groupthink: A Psychological Study of Foreign-Policy Decisions and Fiascoes. Houghton Mifflin, 1972. en.wikipedia.org
  6. 06C. Nemeth, K. Brown and J. Rogers. Devil's advocate versus authentic dissent: stimulating quantity and quality. European Journal of Social Psychology 31, 707, 2001. onlinelibrary.wiley.com
  7. 07B. A. Mellers, R. Hertwig and D. Kahneman. Do frequency representations eliminate conjunction effects? An exercise in adversarial collaboration. Psychological Science 12, 269, 2001. faculty.wharton.upenn.edu
  8. 08Development, Concepts and Doctrine Centre, UK Ministry of Defence. Red Teaming Handbook, 3rd edition. gov.uk, 2021. gov.uk
  9. 09M. Sharma, M. Tong, T. Korbak et al.. Towards Understanding Sycophancy in Language Models. ICLR, 2024. arxiv.org
  10. 10A. Khan, J. Hughes, D. Valentine et al.. Debating with More Persuasive LLMs Leads to More Truthful Answers. ICML, 2024. arxiv.org
  11. 11A. Smit, N. Grinsztajn, P. Duckworth, T. Barrett and A. Pretorius. Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs. ICML, 2024. arxiv.org
  12. 12C. Speake and T. Quinn. The search for Newton's constant. Physics Today 67(7), 27, 2014. physicstoday.aip.org
  13. 13E. Tiesinga, P. J. Mohr, D. B. Newell and B. N. Taylor. CODATA Recommended Values of the Fundamental Physical Constants: 2018. Reviews of Modern Physics 93, 025010, 2021. link.aps.org
  14. 14S. Kirklin, J. E. Saal, B. Meredig et al.. The Open Quantum Materials Database (OQMD): assessing the accuracy of DFT formation energies. npj Computational Materials 1, 15010, 2015. nature.com