Build with intent.
The brief comes before the model.
A useful research system begins with the decision a person needs to make: the target properties, the constraints, and the evidence that would change the answer.
In November 2023 an autonomous laboratory reported synthesising dozens of new inorganic compounds, selected from computed stability data and made without human intervention. The pipeline ran. The results were reported with confidence. A re-analysis published the following year concluded that no new materials had been discovered in the work, and that roughly two thirds of the reported successes were likely already-known disordered forms of the ordered compounds that had been predicted1. The original authors disagree, and the disagreement is still live.
Set the scientific question aside and look at the shape of the failure. No component misbehaved. The predictor predicted, the robot synthesised, the analysis produced a result. What went wrong was upstream of all of it, in what the word novel had been taken to mean. A parallel critique of a large computational screening effort makes the same point from the other direction: it found little evidence of compounds satisfying all three of “novelty, credibility, and utility” at once2. Three requirements. A specification that names one of them will be met, and will not be useful.
This is the failure mode that governs machine-assisted research, and it is not a failure of capability. A system given an under-specified problem does not stall or complain. It returns something confident and well-formed, and the defect is invisible until someone checks.
Substitution punishes vagueness
Materials substitution is an unusually unforgiving place to be imprecise, because the constraint set is much wider than the property that prompted the search. A replacement has to perform, but it also has to be processable on existing equipment, available at volume from a supply that is not itself restricted, permitted under the regulation that forced the substitution, and affordable. Fail any one and the candidate is finished, however well it scores on the rest.
The National Research Council put this at the centre of its definition of criticality in 2008: the harder, more expensive, or slower it is to substitute a material, the more critical that material is3. Criticality, in other words, is a property of the substitution problem as posed — including its deadline. The timescale is the part that is easy to underestimate. Moving a new material from discovery to market has conventionally taken twenty years or more, which is why the stated ambition of national materials programmes has been to do it twice as fast at a fraction of the cost4, a target the National Academies has framed as halving a ten- to twenty-year cycle5.
Meanwhile the forcing functions arrive on their own schedule. The European Union now lists thirty-four critical and seventeen strategic raw materials6. And the scope of a restriction is itself contested for years before anyone has to comply with it.
A restriction proposal covering more than ten thousand substances was published in February 20237. The consultation returned more than five thousand six hundred comments; a narrower proposal followed in August 2025; committee opinions came in March 2026, one final and one still in draft. Three years, and the argument was about which substances were in scope. Anyone who began substitution work in 2023 was specifying against a moving target — which is an argument for writing the specification down, not for waiting.
What a brief has to declare
In Ypotheto the brief is not a preamble. It is built and then reviewed as its own step, and no hypothesis is generated until a person has read it and edited it. That ordering is the whole design: the expensive part of the work is committed only after the specification it will be judged against exists in a form someone has agreed to.
Most of the brief is what you would expect — the goal, the framing, the incumbent material, the candidates already in view, what is explicitly out of scope. The part that carries the argument is smaller. Every evaluation criterion has to declare which direction counts as better, and what kind of evidence would settle it.
- Direction
- maximize · minimize · target · avoid · classify · describe
- Evidence mode
- literature · local tool · external tool · human review · mixed
- Answer shape
- The fields a candidate must carry to be evaluable at all, kept deliberately few
- Weight
- What this criterion is worth relative to the others, read by the scoring arithmetic
The second row is the one that matters most, and it is the field note’s third clause made literal. “The evidence that would change the answer” is not a principle here. It is an enumerated input, set before any agent runs, and it decides whether a claim on that criterion can rest on the literature or has to come from a measurement or a person.
The constraint classes are also independent, which is easy to state and easy to forget. A candidate can be excellent on properties and finished on supply.
Read the columns as the brief. A constraint class the brief never named is a column nobody checked, and a shortlist that ranks well on four of five is not a shortlist.
A better model does not fix an unclear question
The intuition that capability is the binding constraint is testable, and it does not survive. Widely used language models have been shown to swing by as much as seventy-six accuracy points on the same task from changes to prompt formatting alone — separators, spacing, casing — with the sensitivity persisting in larger models and after instruction tuning8. If presentation moves a result that far, the specification is not a detail of the setup. It is the dominant term.
The same conclusion arrives from machine learning practice as underspecification: a pipeline can produce many models that score identically on held-out data and then behave very differently in deployment, because the specification never distinguished between them9. And the defect is frequently in the evaluation rather than the model — a survey of machine-learning-based science documented data leakage across seventeen fields and nearly three hundred papers10. Leakage is a specification defect. It means the thing being measured was not the thing intended.
Fluent restatement is not a specification
Ask a model to write the criteria for success and it will very often return a well-written paraphrase of the goal. It reads like a specification. It contains no bar, so nothing can be judged against it.
This is worth guarding against structurally rather than by asking more nicely, because the output is fluent and the absence is hard to see. A definition of success that restates the opportunity leaves any later review stage assessing whether the work is sufficient against a restatement of the question it started from. So a success definition in the brief has to carry something measurable — a count, a threshold, a named comparison — or be left empty. An empty field is honest, and it is visible. A paragraph that agrees with itself is neither.
Prompts request; schemas constrain
The most useful thing to know about instructions written in prose is that they are requests. A requirement can be stated at the top of a run and restated in the plan built from it, and still not appear in the output — not because the reasoning was poor, but because nothing was checking. Prose does not enforce.
This is where traceability stops being an aspiration and becomes a mechanism. When verifiable citation was first measured rigorously in generative search systems, only about half of generated sentences were fully supported by the sources attached to them, and around three-quarters of citations actually supported the sentence they were cited for11. Grounding a system in sources does not, by itself, make its claims traceable. The trail has to be a structural requirement, checked, or it is decoration.
The consequences of the alternative are not hypothetical in this field. A widely circulated working paper reporting that an AI tool had substantially increased materials discovered at a large firm was withdrawn, with the hosting archive citing concerns about the validity of the data12; the institution stated it had no confidence in the provenance or reliability of that data. The figures had already been repeated in the press before anyone could check them, because there was nothing to check them against.
So the requirements that must hold are carried by the brief’s schema rather than by its wording. A field a candidate must have is checked, and the check is deliberately narrow: it is satisfied by a value, or by a recorded verdict that the value could not be obtained. There is no placeholder, because a placeholder is a hole that looks like data. And because a requirement that cannot be met excludes every candidate rather than the weak ones, the required set is kept small on purpose.
What this actually buys
None of this makes a research system correct. It makes it answerable. A brief that names its criteria, their direction, the evidence that would settle them, and the shape of an acceptable answer produces a shortlist whose reasoning can be inspected by someone who was not there.
That is a lower claim than the field usually makes, and it is the one worth making. The systems that failed publicly did not fail for want of capability. They failed because nobody could tell, from the output, what had actually been asked.
Sources
- 01J. Leeman, Y. Liu, J. Stiles, S. B. Lee, P. Bhatt, L. M. Schoop and R. G. Palgrave. Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis. PRX Energy 3, 011002, 2024. link.aps.org
- 02A. K. Cheetham and R. Seshadri. Artificial Intelligence Driving Materials Discovery?. Chemistry of Materials 36, 3490, 2024. pubs.acs.org
- 03National Research Council. Minerals, Critical Minerals, and the U.S. Economy. National Academies Press, 2008. nationalacademies.org
- 04National Science and Technology Council. About the Materials Genome Initiative. mgi.gov, 2021. mgi.gov
- 05National Academies of Sciences, Engineering, and Medicine. NSF Efforts to Achieve the Nation's Vision for the Materials Genome Initiative. National Academies Press, 2023. nationalacademies.org
- 06European Chemicals Agency. ECHA publishes PFAS restriction proposal. echa.europa.eu, 2023. echa.europa.eu
- 07European Union. Regulation (EU) 2024/1252 on critical raw materials. EUR-Lex, 2024. eur-lex.europa.eu
- 08M. Sclar, Y. Choi, Y. Tsvetkov and A. Suhr. Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. ICLR, 2024. arxiv.org
- 09A. D'Amour, K. Heller, D. Moldovan et al.. Underspecification Presents Challenges for Credibility in Modern Machine Learning. Journal of Machine Learning Research 23, 2022. jmlr.org
- 10S. Kapoor and A. Narayanan. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4, 100804, 2023. cell.com
- 11N. F. Liu, T. Zhang and P. Liang. Evaluating Verifiability in Generative Search Engines. Findings of EMNLP, 2023. aclanthology.org
- 12arXiv administrators. Withdrawal notice, arXiv:2412.17866. arXiv, 2025. arxiv.org