When Agents Run Computations: Schrödinger’s Robert Abel On Where The Human Still Decides
Generating drug candidate ideas has become inexpensive, and the supply of hypotheses now exceeds any organization’s capacity to evaluate them. Schrödinger’s answer, launched at the end of July, is Bunsen: an agentic AI co-scientist that runs the company’s own simulation software on a scientist’s behalf. It is in early access with selected customers, including a strategic software agreement with Bristol Myers Squibb, and full commercial release is expected before the end of 2026, with NVIDIA and Google Cloud supplying the compute.
The company is an interesting place for this to originate. Schrödinger built its reputation on force fields, sampling algorithms and free energy methods rather than on learned models, and its chief executive, Ramy Farid, was until recently among the sector’s more vocal skeptics about AI-driven discovery — STAT ran a 2024 interview with him under the headline “AI in drug discovery is ‘nonsense,’ but call Schrödinger ‘AI’ if you want”. He withdrew the position when Bunsen launched, telling Endpoints News: “I was highly skeptical, to be honest, and couldn’t be more pleased to be wrong. This is transformative.”
To understand what that means in practice, I spoke with Robert Abel, Ph.D., Executive Vice President and Chief Scientific Officer at Schrödinger, about the prototypes that didn’t work, how far the agent has already spread through the company’s own computing infrastructure, which decisions it now makes that used to require a trained modeler, where the human has to stay, and why the experiment that would settle the evidence question can’t realistically be run.
Andrii: Bunsen is advertised as a system that plans and executes computational studies, not just an advanced search tool, so to speak. In practice that likely means it makes methodological decisions which until now belonged to a trained computational chemist/bioinformatician. Could you describe what that looks like across a real workflow, from the initial scientific question to the interpreted result, and where the human remains in the loop?
Robert: Bunsen, our agentic AI co-scientist, helps researchers understand scientific objectives, develop computational strategies, execute sophisticated molecular discovery workflows and interpret results. By combining AI with physics-based simulation, Bunsen allows researchers to apply their expertise at greater scale, enabling them to explore more scientific possibilities, prioritize the most promising opportunities with greater confidence and accelerate discovery decisions.
When a scientist asks Bunsen to perform a complex scientific task, such as running a virtual screen to find novel hits or pursuing a de novo design campaign, Bunsen first puts together a detailed work plan. The scientist reviews that plan, provides feedback and suggests revisions.
Once the scientist is comfortable with the plan, Bunsen can execute the modeling campaign without continuous supervision – work that might previously have required days or weeks of effort by a modeler. The scientist can still inspect intermediate outputs, monitor progress, spot-check results and interact with Bunsen throughout the process.
It's important to be precise about what is being automated. The methods Bunsen draws on — our force fields, sampling algorithms and free energy methods, and the protocols that govern how they are applied — are the product of decades of work by a large team of outstanding scientists, validated extensively against experimental data. Bunsen is not inventing new methods, rather, it is deploying that body of work, selecting among protocols we have validated and operating within parameter ranges we have characterized. The scientist still defines the scientific objective and ultimately determines whether Bunsen has interpreted the question correctly and whether the approach is fit for purpose.
We’ve recently found that roughly 80% of the computational modeling jobs running on our HPC are now being run by Bunsen. These include free energy calculations, molecular dynamics simulations, quantum calculations and other sophisticated atomistic modeling. We’ve moved well beyond these systems being advanced search tools. Bunsen is doing serious work.
Andrii: Attempts to build reasoning systems for science are not new. Expert systems such as MYCIN and Internist-1 generated comparable expectations in the 1970s and 1980s and eventually collapsed, largely because the computational and data environment of that time was immature. Today the environment is different. What, in your view, makes an agentic co-scientist viable now, and which parts of that old problem do you consider still unsolved?
Robert: We actually started developing Bunsen-like agentic systems at Schrödinger in late 2023, shortly after GPT-4 launched. For the next year and a half, we consistently found that these systems would make mistakes on relatively simple tasks that prevented them from being useful in a realistic scientific setting, even though these systems were vastly more powerful and flexible than the expert systems developed in the 70s and 80s.
By late 2025, our internal testing suggested agentic systems had reached a step-change in capability that made the idea of an agentic co-scientist much more feasible. The underlying models had become significantly stronger, driven by improvements in model architectures, training strategies and computing resource availability.
But there’s another important distinction from earlier expert systems: Bunsen is not expected to simply provide “the answer.” It understands when and how to use a range of atomistic modeling techniques and then uses those methods to investigate a scientific question. That gives it an external check: the molecular simulation itself. Bunsen can pursue an approach, evaluate the results and adjust course based on what the simulation shows. MYCIN had nothing similar beyond waiting for a patient outcome.
What remains difficult for agentic AI more broadly is recognizing the limits of its own knowledge and capabilities. That’s one reason the human expert remains important – reviewing the work plan and spot-checking progress as the work proceeds.
Andrii: Over the past several years the dominant narrative in AI-driven drug discovery has shifted toward learned models (i.e. co-folding networks, generative chemistry, neural network potentials, etc) and physics-based simulation has occasionally been framed as the older approach. Schrödinger has consistently argued the two are complementary. What are the limits of a learned model's prediction abilities without grounding into physics, and what specifically does physics contribute at that point that data alone cannot?
Robert: Physics-based approaches have been the gold standard for affinity prediction in molecular discovery and remain so today. However, these simulations are computationally intensive, which has motivated consideration of many alternatives, including AI methods. These methods are faster, but they have significant shortcomings. In fact, machine learning-based affinity prediction methods are often compared against physics-based free energy calculations as a benchmark. And, in most of those comparisons, physics-based methods appear to maintain a clear advantage over purely machine learning-based approaches, especially so if the authors did a good job controlling data leakage.
The limitation of any machine learning model for property prediction is that its accuracy generally worsens as you move further away from its training data. That matters because real drug discovery frequently happens where relevant data is limited – teams may be investigating novel scaffolds, newly characterized resistance mutations or targets with few previously reported ligands. Physics-based methods can be applied in those settings even when little project-specific data exists.
That’s also where the two approaches become particularly powerful together. Physics-based methods can generate in silico training data at scale, which can then be used to train project-specific machine learning models. A typical lead-optimization project might experimentally synthesize fewer than 1,000 intentionally designed molecules over a year, while atomistic simulations can generate several thousand data points in an evening.
The resulting machine learning model can then evaluate millions or even billions of molecules, with the most promising compounds advanced into confirmatory atomistic simulations before synthesis. Together, physics and machine learning allow teams to explore chemical space much more rapidly than either technique could alone.
Andrii: In computational chemistry the danger is rarely an obvious error, I think. AI can return a plausible number and the study appears successful on the surface… until proven wrong experimentally much later. When an agent selects the methods and protocols itself, that risk scales with throughput. How is this addressed in Bunsen's design, and what allows a scientist to see the assumptions behind a result rather than only the result?
Robert: An important distinction is that the LLM powering Bunsen is not directly predicting molecular properties. Bunsen determines how and when to use computational chemistry methods including molecular dynamics, quantum mechanics, free energy calculations, protein structure refinement, docking and pKa prediction. The numerical results are generated by these extensively validated computational methods, not by the LLM itself.
Before Bunsen begins the work, it transparently communicates how it plans to use those methods and the scientist can review and refine that approach. Once the work is underway, the scientist has visibility into the intermediate results and can spot-check the data and progress throughout the process.
The scientist isn't simply presented with a final number. They can see the methodology, assumptions and intermediate work that produced the result.
Andrii: One of the stated goals is to make advanced computational methods accessible to experienced drug hunters who are not computational chemists. That changes the composition of a discovery team. Based on what you have observed so far, how does the role of the computational chemist evolve once execution is automated, and which skills become more valuable in modern days?
Robert: One of the features I’m most excited about is the ability to collaboratively share Bunsen sessions. A medicinal chemist and computational chemist can be in the same session from different locations, with the medicinal chemist driving the work while the computational chemist has full visibility into the modeling and can step in when needed.
We’ve already had medicinal chemists use Bunsen directly to perform complex compound enumerations and data analyses that previously would have required direct computational chemistry support. For more computationally intensive work, expert review of the work plan and spot-checking of progress remains important.
In some sense, everyone working in computational chemistry has been promoted to a management role. A significant share of a computational chemist’s time has historically gone toward preparing structures, setting up and monitoring jobs, and reformatting outputs. Bunsen can take on more of that execution, freeing computational chemists to spend more time on the intellectually challenging parts of the role: defining the problem, determining the best approach, interpreting the data and deciding when the methodology needs to change.
Computational chemists are still active participants in the work. They now have a virtual colleague that can help execute it.
Andrii: There is an economic asymmetry here that is rarely discussed. The cost of AI inference has fallen dramatically in recent years, while the cost of physics-based simulation is determined by the amount of sampling the physics requires and has not declined at a comparable rate. Schrödinger has indicated that Bunsen’s value will be captured through increased platform throughput rather than separate pricing. How should a research organization plan its computational budget when the system deciding how much to compute is itself an agent?
Robert: First, one important clarification: Bunsen is currently in early access, and we have not formally announced how it will be licensed.
More broadly, organizations using atomistic modeling have to consider three costs: compute costs, software licensing costs for proprietary methods, and the FTE costs of highly talented scientists applying those methods effectively.
That third cost is significant. Computational chemists in industry are often supporting three or four discovery projects simultaneously, which means there are valuable modeling activities they simply may not have time to pursue. A system like Bunsen allows those scientists to work much more efficiently and apply computational methods more broadly.
Andrii: The industry has accumulated a considerable number of acceleration claims and comparatively few independently verified outcomes, which is one reason a degree of skepticism persists. What kind of evidence would you consider sufficient to demonstrate that an agentic system improves the quality of scientific decisions rather than only their speed, and over what time horizon should the industry expect to see it?
Robert: This is a difficult question not just for agentic systems, but for predictive modeling in drug discovery more broadly. Running the same discovery programs in parallel – with and without predictive modeling – to create a controlled comparison would be extraordinarily expensive and impractical. And, one would need to run this cost prohibitive experiment for many different drug discovery projects to obtain statistically robust conclusions.
Instead, there are two important criteria to consider. First, predictive modeling needs to be sufficiently accurate to support better project decisions. Second, it needs to operate at the scale necessary to successfully identify compounds that would not otherwise have been prioritized. Our predictive modeling methods have clearly met both criteria in prospective testing reported in peer-reviewed publications, as well as through our internal and collaborative discovery programs.
For agentic systems like Bunsen, the opportunity is to allow scientists to apply those predictive methods more efficiently, thoroughly and at greater scale. To the extent one accepts the preceding evidence that predictive modeling clearly supports better project outcomes, then it naturally follows that any technology that allows predictive modeling to be applied more efficiently and at greater scale should help project teams to succeed. Ultimately, the evidence we'll want to see is whether expanded use of Agentic AI translates into better project decisions and, ultimately, better project outcomes.
Go deeper into the techbio strategic signals with BiopharmaTrend Pro Membership.
Topic: AI in Bio