When hypotheses become cheap, validation becomes the product

AI agents can produce scientific conjectures faster than institutions can test them. The useful system is therefore not the idea generator alone, but the evidence queue that decides what deserves contact with reality.

The latest argument about AI in science sounds familiar from the software side: generation is getting cheap, while proving that an output deserves to change the world remains stubbornly expensive.

Google DeepMind's new report, Conjecture Machines, describes agents proposing hypotheses, designing experiments, and searching candidate solutions at a pace that laboratories and reviewers cannot match. Its most striking example is not a benchmark. A microbiology team spent most of a decade establishing a mechanism for antibiotic resistance. Given the problem in 2024, Co-Scientist returned five ranked explanations in two days, with the team's unpublished result at number one.

That is a remarkable compression of conjecture. It is not a compression of proof.

I have been circling the same boundary in V0M3 and N17Q: generated material is a proposal, not accepted state; a fluent claim is not evidence; a completed run is not the same thing as a verified outcome. Science raises the consequence of that distinction. The candidate may be a molecule, a proof, or a biological mechanism, and the receipt may take months.

The product opportunity is no longer just a better conjecture machine. It is a system that spends scarce validation well.

The asymmetry is the story

An agent can read across disciplines, generate alternatives, criticize them, and rank the survivors in parallel. Many validation paths remain serial, physical, and capacity constrained. Cells still need time to grow. Instruments have queues. Peer reviewers have finite attention. A negative result may arrive only after a synthesis, assay, or field observation has consumed real material.

That produces an unusual scaling problem. More generation does not relieve the bottleneck. It increases pressure on it.

If a system creates one hundred plausible hypotheses where a team can test three, the central question changes from “can it think of something?” to “which three purchases are worth making with reality?” Ranking by confidence is not enough. A confident proposal may be expensive to falsify. Ranking by novelty is not enough. A surprising claim may depend on unavailable equipment or an unmeasurable outcome. Ranking by expected impact alone encourages spectacular stories with weak stopping conditions.

The queue needs to understand evidence, cost, reversibility, and information gain at the same time.

Queue signalQuestion it should answerFailure if omitted
Prior supportWhat observations make this candidate plausible?The system repeatedly buys tests for attractive speculation.
Discriminating resultWhat outcome separates this hypothesis from its strongest rival?A successful experiment still leaves the explanation ambiguous.
FalsifierWhat result would make the team stop believing or revising it?Every outcome is narrated as partial confirmation.
Validation costWhich scarce instrument, person, sample, or elapsed time does this consume?Cheap confidence scores allocate expensive physical work.
ReversibilityCan an unsafe or misleading result be contained?The experiment has consequences the ranking model never priced.
Information gainWhat will be learned even if the candidate fails?Negative results become waste instead of navigation.

My working hypothesis is that the best scientific-agent interface will look less like a chat and more like a constrained experiment exchange. Ideas enter freely. Access to validation does not.

A conjecture needs a contract

A hypothesis should become a durable object before it enters the experiment queue. The object needs more than prose and a confidence score.

It should name the claim precisely enough to disagree with. It should carry the evidence graph that led to it, including contradictory findings and retrieval dates. It should distinguish observations from model inference. It should state the nearest competing explanation, the test intended to separate them, and the result that would trigger revision or abandonment.

It should also disclose the model, scaffold, tools, data revisions, prompts or skills, and selection procedure that produced it. If ten thousand candidates were generated and one survived, that search history matters. A single polished answer hides the size of the lottery.

The experiment contract then adds operational constraints: eligible facility, sample requirements, safety boundary, estimated cost, expected duration, measurement protocol, preregistered analysis, and stop conditions. Approval binds to that exact contract. A later change to a dosage, dataset, or measurement does not inherit authorization invisibly.

The diagram has a deliberate bottleneck. A proposal cannot jump from plausible language to the knowledge base. It has to acquire a testable shape, compete for validation, and return with a receipt.

Confidence is not the queue

Model confidence is useful when it is calibrated, scoped, and attached to a particular claim. It is still only one feature of a decision.

Imagine two candidates. The first has a 70 percent estimated chance of being correct and requires three months of scarce lab time. Its likely outcome would only weakly distinguish it from an existing explanation. The second has a 35 percent estimated chance, can be tested overnight using an existing dataset, and would eliminate an entire family of models if it fails.

The first candidate is more likely to be true. The second may be the better experiment.

This is the breakthrough in framing that I find most useful: rank validations, not just hypotheses. A conjecture score asks how promising an idea appears. A validation score asks whether spending the next unit of scarce reality on it will reduce important uncertainty.

The distinction makes room for deliberately adversarial tests, replication, instrument calibration, and null experiments. None looks exciting in an idea leaderboard. All can be more valuable than generating candidate number 10,001.

A practical queue could estimate a range rather than one synthetic score:

  • Expected information gain under positive, negative, and ambiguous outcomes.
  • Cost in money, samples, facility time, human attention, and calendar time.
  • Dependency value, meaning how many other claims become clearer after this result.
  • Replication value, especially when an influential result has weak independent support.
  • Safety and containment requirements.
  • Time sensitivity, including samples, observations, or opportunities that expire.
  • Reuse value of the resulting dataset, method, or calibrated instrument.

The ranking should preserve its inputs and competing candidates. Otherwise a model can produce an authoritative ordering that nobody can reconstruct after the result arrives.

Negative results need first-class storage

Generation-heavy systems are tempted to treat a failed candidate as dead output. That throws away the expensive half of the loop.

A negative result can constrain a parameter, invalidate an assumption, expose a confounder, calibrate a method, or rule out an entire branch of future search. It should return to the evidence graph with the same care as a positive result. The next agent needs to know not merely that the experiment “did not work,” but which intervention was performed, what was observed, within which detection limits, against which controls, and under which conditions.

Ambiguous results matter too. A failed instrument, contaminated sample, underpowered test, or analysis deviation is not evidence against the hypothesis. It is evidence about the validation attempt. Collapsing both into failure teaches the next system the wrong lesson.

This is where receipts become more than audit logs. The receipt is the boundary between what the agent intended and what the experiment actually established.

For computational work, a receipt may include immutable inputs, code revision, environment, random seeds, full outputs, checks, and a verifier result. For physical work, it may include sample lineage, instrument state, protocol deviations, raw measurements, calibration records, operator attestations, and custody. In both cases, the final prose is downstream of the receipt.

Provenance needs to survive synthesis

An agent may read hundreds of papers and produce one compact mechanism. A citation attached to the paragraph does not explain which source supports each part, which step is an inference, or whether an apparently independent claim traces back to the same original experiment.

I would model support as a graph with typed edges: supports, contradicts, reproduces, derives from, shares data with, and merely mentions. The graph should retain source revisions and the observation cutoff. When a paper is corrected or a dataset changes, affected claims can become stale without erasing their history.

Synthesis can remain readable. The graph does not have to spill into every sentence. It does need to be available whenever a claim crosses into a decision, experiment, publication, or downstream agent context.

The point is not bureaucratic completeness. It is to keep a machine-generated connection from turning into institutional memory before anyone can locate its foundation.

Verification should consume a different budget

If idea generation and verification draw from the same undifferentiated compute budget, the generator can spend the account before the critic arrives. The same applies to human attention. A long, polished report consumes review capacity before its key claims have been triaged.

I would reserve validation budgets explicitly. A run can generate broadly within one limit, but only a bounded number of candidates may request formal checking, expert review, or physical execution. High-consequence claims require independent methods or actors. The generator cannot appoint its own output as verified.

Mathematics gives an early glimpse of the shape. DeepMind describes Aletheia pairing a proof generator with a verifier and sending flaws back for revision. Formal systems can sometimes produce an unambiguous check. Even there, humans still need to understand significance, assumptions, and whether a machine proof answers the question people care about.

Most science cannot reduce validation to one verifier call. That makes budget separation more important, not less.

Review interfaces should begin with disagreement

Most generated-research interfaces start with a coherent answer. That presentation encourages the reviewer to edit prose instead of interrogating the claim.

I would start with the decision boundary instead:

  • What is the exact claim?
  • What is the strongest alternative explanation?
  • Which evidence would look different if either were true?
  • Which part came from source material and which part was inferred?
  • What would change the ranking?
  • What is the cheapest decisive test?
  • Which uncertainty cannot be reduced with the available methods?

Only then should the system show the polished narrative. The prose becomes a view over the research state, not the state itself.

Review also needs diversity that is operational rather than cosmetic. Five agents using the same model, sources, and scaffold are not five independent confirmations. The system should expose shared dependencies and seek genuinely different data, methods, or evaluators where independence matters.

A lab scheduler becomes part of the reasoning system

Once physical validation is scarce, scheduling is epistemic. Which experiment runs next affects which branch of knowledge can advance, which samples expire, and which claims continue to influence decisions without confirmation.

A scheduler should understand prerequisite results, batching opportunities, instrument setup, contamination boundaries, parallelizable work, and the value of early stopping. It should also protect replication and calibration from being crowded out by novel candidates with better stories.

The agent can propose a schedule. Facility policy and scientific ownership constrain it. Every reschedule changes the evidence horizon and should be visible to downstream planning.

This begins to resemble the execution boundary in N17Q. Intent, authorization, attempt, observation, and current account are separate. The difference is that nature is the external system, and its callback latency is non-negotiable.

The next benchmark should price reality

Benchmarks for scientific agents often ask whether a system produced a known answer, useful hypothesis, or high-scoring candidate. Those are necessary tests of capability. They do not measure the institution the agent will enter.

I want evaluations that give an agent a validation budget and a stream of incomplete, contradictory evidence. The agent must choose which experiments to buy, update beliefs after negative or ambiguous results, preserve provenance, and stop when another test is not justified. Hidden ground truth can measure discovery, but the score should also include wasted validation, missed falsifiers, unsafe proposals, duplicated experiments, calibration, and whether the final account matches the receipts.

The interesting question is not how many ideas the agent can produce. It is how much uncertainty the whole system can responsibly remove per scarce validation unit.

That is a harder benchmark because it evaluates judgment across time, not one answer. It is also closer to the actual bottleneck.

The conjecture machine is only the front half

DeepMind's report is optimistic about scientific agents and direct about the infrastructure they will strain. I think the validation bottleneck is not an unfortunate delay between model generations. It is the product boundary that makes machine-assisted discovery credible.

Cheap conjectures are valuable because they let us search spaces that human attention could never enumerate. They become dangerous when volume is mistaken for knowledge, model ranking is mistaken for experimental priority, or polished synthesis outruns its evidence.

The system I want is generous with ideas and strict with state. It remembers failed tests. It prices physical work. It can explain why one experiment displaced another. It treats a negative receipt as accumulated knowledge and an ambiguous receipt as a reason to resist storytelling.

The conjecture machine may be the visible breakthrough. The validation ledger is what lets the breakthrough survive contact with reality.