As AI moves from analyzing data to generating hypotheses, proofs and experiments, the bottleneck in science may be shifting from intelligence to physical validation.
For most of the past decade, the role of artificial intelligence in science was relatively easy to describe. AI was a tool. It could classify an image, fit a model, search a chemical space or predict the structure of a protein. The scientist still decided what question mattered, designed the experiment and interpreted what the result meant.
That distinction is becoming harder to maintain.
A new generation of AI systems is beginning to participate in the research process itself. They search literature, generate hypotheses, write experimental code, analyze results, revise their assumptions and, increasingly, decide what to try next. In mathematics, AI systems are producing formal proofs that computers can verify. In biology, they are searching genomic datasets for previously unknown molecular machinery. At experimental facilities, agents are beginning to operate scientific instruments and respond to unexpected conditions.
The change is subtle but important. AI is moving from being a scientific instrument to becoming something closer to a research actor.
That does not mean an AI system has become a scientist in the human sense. It does not choose a career, become curious about nature or decide which problems society should care about. But if science is understood operationally as a loop of observation, hypothesis, experimentation, analysis and revision, machines are beginning to perform a surprisingly large portion of that loop.
The implications may be larger than simply making existing researchers more productive.
If AI can eventually generate scientific ideas much faster than laboratories can test them, the fundamental bottleneck of scientific progress could move. For centuries, one of the scarcest resources in science has been human intellectual attention. In an AI-driven research system, the scarce resource may increasingly become something else: access to experiments, instruments, samples, physical materials and reality itself.
AI has already transformed several scientific disciplines without becoming an autonomous researcher.
AlphaFold is perhaps the clearest example. Protein structure prediction once required extensive experimental work and specialized computational methods. Deep learning made it possible to predict structures at extraordinary scale. Yet the conceptual relationship remained familiar: researchers posed the problem, the model produced a prediction and humans decided what to do with it. That model of collaboration is now changing.
In May 2026, researchers published Robin, a multi-agent system designed to automate important parts of experimental biological research. Robin combines agents for literature search and data analysis, allowing the system to generate hypotheses, propose experiments, interpret experimental results and generate updated hypotheses. Researchers used the system to identify potential therapeutic candidates for dry age-related macular degeneration.
That distinction matters because Robin is not simply answering a question posed to it. It is participating in an iterative process.
The traditional image of AI in science is a researcher sitting in front of a powerful computational instrument:
Human question → AI answer.
Agentic scientific systems begin to look different:
Question → hypothesis → experiment → result → interpretation → revised hypothesis.
Humans still remain deeply involved, particularly where physical experimentation is required. But more of the intellectual machinery between those points can now be delegated.
A similar pattern is emerging in AI research itself. A Nature paper published in March described an “AI Scientist” capable of generating research ideas, writing code, executing experiments, analyzing results, producing figures, writing papers and conducting a form of automated peer review. One manuscript produced by the system passed the first stage of review at a machine-learning workshop, although the authors were careful to describe substantial limitations and the continued importance of human oversight. The interesting question is therefore no longer whether AI can perform individual research tasks. It clearly can. The question is how much of the research loop can eventually be connected.
Mathematics may be the most revealing environment for understanding this transition because it removes one of science's hardest problems: physical verification. An idea in experimental biology might require months of laboratory work before anyone knows whether it is correct. A mathematical proof can, at least in principle, be checked immediately. That makes mathematics unusually well suited to autonomous AI research.
In September, OpenAI reported that an internal system had generated a proposed solution to the Navier–Stokes existence and smoothness problem, one of the Millennium Prize Problems, accompanied by a formal proof in Lean. OpenAI said the system used for the result was significantly more capable than GPT-6 Astra and emphasized that the work involved collaboration between mathematicians and AI researchers.
Anthropic has been pursuing a similar direction. In September it published work on a complete computer-checked formalization of Fermat's Last Theorem. The significance of these systems is not simply that AI can produce unusually difficult mathematics. Formal verification changes the economics of exploration. Human mathematicians are constrained by time. A researcher cannot seriously pursue millions of possible proof strategies. An AI system potentially can explore a much wider search space, provided that unsuccessful paths are inexpensive and successful ones can be automatically verified.
This creates a powerful architecture: generation at machine scale + verification at machine scale. The second half is crucial.
Generative AI became useful because models became remarkably good at producing plausible outputs. Scientific discovery requires something stronger. Plausibility is not enough. Science needs mechanisms that separate interesting speculation from correct results. Formal mathematics offers exactly that. A proof assistant such as Lean does not care whether a proof was written by a Fields Medalist, an undergraduate or an AI system. If the formal argument satisfies the rules, the proof checks. That makes mathematics an unusually clean laboratory for AI-driven discovery. The model can generate. The verification system can reject.
In many other sciences, nature itself performs the verification—and nature operates much more slowly.
The transition from mathematics to biology exposes the central challenge of AI science. In September, Anthropic announced that Claude had identified a previously uncharacterized enzyme system containing CRISPR-like repeats. According to Anthropic, the system emerged from a project in which Claude explored large DNA datasets, generated hypotheses and worked with scientists who subsequently tested those hypotheses experimentally. Anthropic announced a dedicated life-sciences research group and laboratory alongside the result.
This is a different kind of achievement from solving a benchmark. A language model is not merely retrieving something already described in a paper. It is navigating a biological dataset, noticing patterns and proposing that something unknown may exist. But notice what happens next. The AI can generate the hypothesis quickly. The laboratory still has to determine whether the hypothesis corresponds to biological reality.
That gap may become one of the defining features of AI-driven science. Suppose a conventional research group can produce ten serious hypotheses in a year. A sufficiently capable AI system might eventually produce ten thousand.
At first this sounds like an extraordinary acceleration in science. But if the laboratory can still test only ten of them, the scientific system has not accelerated by a factor of one thousand. Instead, the bottleneck has simply moved downstream. The scarce resource becomes experimental throughput. This is why developments in laboratory automation may ultimately matter as much as improvements in scientific reasoning models.
Researchers are already experimenting with agents that interact directly with scientific equipment. A 2026 Nature Machine Intelligence paper demonstrated an AI agent capable of autonomously performing X-ray sample alignment at a synchrotron beamline. The agent planned actions, issued commands to instruments, interpreted observations and adapted when conditions changed. The workflow was first developed in a virtual environment and then deployed on a real beamline.
That may sound like a narrow technical achievement, but it points toward a much larger architecture. The full scientific loop requires two worlds. The first is digital: literature, datasets, models, simulations, mathematics and reasoning. The second is physical: robots, microscopes, particle accelerators, chemical reactors, biological samples and measurement systems.
AI is progressing extremely quickly in the first world because digital actions are cheap. An agent can read another paper, rewrite a program or run another simulation almost instantly.
Physical science obeys different constraints. Cells must grow. Chemical reactions take time. Materials must be fabricated. Telescopes have limited observing windows. Clinical trials can take years. Instruments cost money and laboratories have finite throughput. Even highly automated laboratories cannot escape physics. This creates an unusual asymmetry. AI could potentially expand the supply of scientific ideas far more quickly than the infrastructure required to verify them. The result may be a new kind of scientific abundance: far more hypotheses than humanity can test.
There is another change underway that receives less attention but could substantially alter how science works. Scientific knowledge has traditionally been stored in papers. A paper is static. It describes what researchers did and leaves future scientists to reproduce the methods themselves.
A Nature paper published in September introduced Paper2Agent, a system designed to transform scientific papers into interactive AI agents. Instead of simply reading a paper, another system can access its methods and data through standardized tools and reuse them in new research workflows. In one demonstration, the agent combined genetic data with AlphaGenome tools to generate and computationally test a candidate mechanism related to ADHD risk, while explicitly leaving the biological hypothesis for subsequent experimental validation.
This points toward an interesting possibility. Scientific literature may gradually evolve from a library of documents into a network of executable knowledge. A future researcher—or research agent—might not simply read a method section. It could invoke the method. It might not simply cite an earlier analysis. It could rerun it against new data. And instead of a scientific paper being the terminal output of one research project, it could become a reusable component inside thousands of subsequent research processes. That would fundamentally change the velocity with which scientific knowledge compounds.
The conventional story of scientific progress is that insight is scarce. There are more possible questions than researchers. More datasets than people can analyze. More chemicals than chemists can synthesize. More papers than anyone can read. AI attacks exactly this scarcity. Models can search enormous bodies of literature, generate candidate explanations, write code, analyze results and maintain parallel lines of investigation at a scale no research team could practically match.
But solving one scarcity usually reveals another. If intellectual exploration becomes cheap, then experimental validation becomes relatively expensive. A useful way to think about the scientific process may therefore be as a pipeline whose bottleneck is moving. For much of modern scientific history: Ideas → experiments → discoveries was constrained heavily by the availability of human researchers capable of producing and pursuing good ideas. An AI-heavy research environment may look more like: millions of machine-generated hypotheses → limited experimental capacity → verified discoveries. At that point, science begins to resemble an infrastructure problem.
Which laboratories have robotic automation? Which institutions have access to synchrotrons, sequencing facilities, telescopes, cleanrooms and biological samples? Which experiments can be simulated accurately enough to reduce the need for physical testing? Which systems can automatically prioritize the 100 hypotheses worth testing out of the million that AI produced? These may become more important questions than how quickly another language model improves on a reasoning benchmark.
The most important consequence of scientific AI may therefore have little to do with whether a machine deserves to be called a scientist. The terminology is less interesting than the underlying economics. Scientific intelligence has historically been expensive. Training an expert takes decades. A researcher can read only so much, run only so many analyses and seriously explore only a limited number of ideas. AI is beginning to make parts of that intelligence reproducible at software scale.
If the trend continues, generating a hypothesis may eventually become cheap. Testing it will not. This is why the next phase of AI-driven science will probably depend as much on robotics, laboratory automation and scientific infrastructure as on better foundation models. A research agent capable of producing thousands of credible experiments is only as powerful as the physical system capable of executing them.
The future scientific laboratory may therefore look less like a room full of researchers individually conducting experiments and more like a continuous loop connecting models, simulations, automated instruments and humans who decide which questions matter. The AI will propose. Machines will test. Humans will choose the direction. And reality will still have the final vote. If AI can generate scientific ideas faster than we can test them, the bottleneck in science will no longer be intelligence. It will be the speed at which we can ask the physical world whether those ideas are true.
10/02/2026