Big Pharma Is Betting On LLMs and Agents. What Would Count As Success?
In the latest news, Novo Nordisk has entered a collaboration with Anthropic aimed at speeding the discovery and development of new medicines. An initial project will test Anthropic's Claude Science on specific R&D workflows where the companies expect the greatest impact, and Novo also plans to use the models to strengthen software development so it can scale AI use across the company. Financial terms were not disclosed.
"Joining forces with Anthropic will supercharge our R&D organization…," Novo’s chief executive Mike Doustdar said, restating a goal the company has repeated all year: to become the world's most AI-driven healthcare company.
It is the latest in a run of deals to bring onboard frontier models, and agentic AI processes specifically, including Novo's deal also with OpenAI in the spring to bring AI tools into every step of its operations. Bristol Myers Squibb struck its own agreement with Anthropic to let the models work across its institutional knowledge. Merck went with Google Cloud and its agentic ecosystem in April this year. Amgen was an early adopter of OpenAI’s ChatGPT Enterprise across various functions.
Go Deeper: Everyone is Launching AI Agents. What's Being Deployed? A deep dive on the next cycle of biopharma's AI buildout.
The pattern is real, and the spending is accelerating, but what gets lost in the coverage is that these deals bundle together very different propositions, and the strength of evidence behind those value propositions varies significantly.
The enterprise AI case for LLMs and agentic systems in the pharmaceutical industry
One part of what large pharma organizations are buying is functional enterprise AI, applied to a document-heavy, process-heavy business: software engineering, regulatory and medical writing, literature and internal knowledge retrieval, content creation, meeting and project orchestration, contract and submission workflows, financial planning.
One of the most fully documented examples of this in pharma is Moderna, which deployed ChatGPT Enterprise company-wide. The reported numbers from 2024 include more than 750 internal GPTs built across legal, research, manufacturing and commercial within two months, 40% of weekly active users creating their own, and an average of 120 conversations per user per week, with the legal team at full adoption. CEO Stéphane Bancel has gone further, arguing that a team of a few thousand can perform like a team of 100,000. Two years on, no updated adoption figures and no productivity assessments have followed. What Bancel has offered instead is a target: 15 new products in five years with fewer than 6,000 employees, with help of AI.
If you look at most announcements, they are mostly talking about adoption metrics, not outcome metrics. Number of assistants built, conversations held, percentage of staff logging in. Neither Moderna nor OpenAI has published a quantified productivity or financial result from the deployment, and Bancel's comparison is an assertion and planning, rather than a measurement.
There are some reports outside pharma, however, that can give a feel of what it really means to integrate LLMs/agents, with some mixed sentiment. Three randomized trials of GitHub Copilot covering 4,867 developers at Microsoft, Accenture and an anonymous Fortune 100 electronics manufacturer, published in Management Science, found a 26% increase in completed tasks, though with a wide confidence interval and gains concentrated among less experienced staff — short-tenure developers improved 27% to 39%, long-tenure developers 8% to 13%. Without access to proprietary code, the authors measured quality by proxy, finding no decline at Microsoft but falling build success rates at Accenture.
Enterprise-level returns have been harder to demonstrate, with MIT's NANDA initiative reporting that 95% of organizations are getting zero return on an estimated $30–40 billion of enterprise generative AI spending, with only 5% of integrated pilots extracting real value and the rest showing "no measurable P&L impact." That work drew on more than 300 publicly disclosed initiatives, structured interviews with 52 organizations and a survey of 153 senior leaders, and its authors caution the figures are "directionally accurate based on individual interviews rather than official company reporting."
A randomized trial by METR found that 16 experienced open-source developers took 19% longer to complete real issues in their own repositories when allowed to use AI tools, while estimating afterwards that the tools had made them 20% faster. METR has since reported that running this design is getting harder, as developers increasingly decline to work without AI and between 30% and 50% avoided submitting tasks they expected to be slow without it — biases that, by the organization's own account, push its estimate of AI's benefit downward.
Pharma-specific evaluations are even thinner on evidence. In a joint article on generative AI and clinical study reports, specialists from Eli Lilly and Novo Nordisk, writing with the AI writing-automation vendor Yseop, note that producing a CSR from receipt of final tables took eight to 15 weeks in 2013, and that a recent survey of five global pharmaceutical companies put the range at six to 15 weeks today. So, no change, essentially. But their argument is not that AI cannot help, on the contrary, they seem to be advocates for it. It is that the bottleneck is the process rather than the technology, and that automating a bloated process only generates content faster while leaving weeks of manual review in the air.
So even the operational half of pharma's AI bet is better described as promising than proven. One qualifier belongs to all of the above considerations: these studies and surveys cover late 2025 and the first half of 2026, and agentic systems are moving fast enough that a comprehensive evaluation in January 2027 could be more revealing.
The scientific AI case for agentic systems in the pharmaceutical industry
Pharma's second bet is that frontier models and agentic systems can do hard science — reason over mechanism, generate hypotheses, propose targets. The frontier AI labs have been building toward exactly that, with varying effort.
Anthropic launched Claude Science this summer as a research workbench wired to more than 60 scientific databases, with prebuilt toolkits for genomics, protein structure and chemistry, sub-agents that take delegated tasks, and a fact-checking agent that verifies citations and calculations. Early users include the Allen Institute, where a neuroscientist built a multi-agent review pipeline, and the UCSF Brain Tumor Center, which used it to speed germline analysis. In the spring, Anthropic had already paid roughly $400 million for Coefficient Bio, a stealth startup of fewer than ten people founded by alumni of Genentech's Prescient Design, whose platform drafted R&D plans, handled regulatory strategy and proposed candidates.
OpenAI has moved from a science team to a science model. GPT-Rosalind, announced in April and generally available this month through a trusted-access programme, is a reasoning model built for biology, drug discovery and translational medicine, wired to more than 50 public multi-omics databases and biology tools. OpenAI reports it outperforming GPT-5.4 on six of eleven LAB-Bench2 tasks, leading performance on the bioinformatics benchmark BixBench, and, in an evaluation with Dyno Therapeutics, scoring above the 95th percentile of human experts on an RNA prediction task. Named partners include Amgen, Moderna, Thermo Fisher, the Allen Institute, NVIDIA, Benchling and Novo.
Google has gone furthest on experimental validation. Its AI co-scientist is a multi-agent system built on Gemini 2.0, with generation, reflection, ranking, evolution, proximity and meta-review agents under a supervisor, iterating on their own hypotheses. Three applications were tested in the lab: with Houston Methodist Hospital it proposed drug repurposing candidates for acute myeloid leukemia that inhibited tumor viability at clinically relevant concentrations across multiple cell lines; with Stanford it identified epigenetic targets for liver fibrosis that showed significant anti-fibrotic activity in human hepatic organoids; and with the Fleming Initiative and Imperial College London it independently proposed that cf-PICIs interact with diverse phage tails to expand their host range — a mechanism of antimicrobial resistance that matched experimental work unpublished at the time.
How to measure success of adopting AI and agents in bio?
With the growing adoption of LLMs, agents, and more specialized AI systems by large pharma companies, there is a growing need to actually benchmark results and estimate ROIs. It is still not trivial and mostly an open question. However, there is one notable effort in this direction, a new research paper published in Cell evaluating the performance of various AI models and systems on “hardcore” biological tasks.
Specifically, an AI company Insilico Medicine, with researchers from Liquid AI, the Buck Institute for Research on Aging, Harvard Medical School and Brigham and Women's Hospital, released an open benchmark called LongevityBench, testing reasoning across clinical records, genetics, epigenetics, transcriptomics and proteomics, built to reward interpretation of new biological measurements rather than recall of training data.

Evaluation of frontier models and fine-tuned L-LLMs across 17 aging biology tasks. Image credit: Zhavoronkov A, Naumov V, Sidorenko D… An open benchmark and language models for AI in aging biology, Cell, 189, 5980-5994.e8
Eighteen frontier systems were evaluated, from OpenAI, Google, Anthropic, xAI, DeepSeek, and Moonshot AI. None led across all five domains, with performance shifting depending on how questions were phrased. Predicting biological age directly from omics measurements was the hardest task on the benchmark, and it stayed hard no matter how large the model. The team then fine-tuned five small open models on aging data — the largest just 9 billion parameters. The best of them ranked first of all 26 systems tested, ahead of Gemini-3.1-Pro, and even the 0.6-billion-parameter version placed sixth, ahead of most of the frontier field. On these tasks, what a model was trained on mattered more than its size.
This benchmark tries to move the needle in assessing technical model performance. Still, of course, it is just one part of the puzzle if we talk about being able to evaluate the strategic impact of AI adoption, ROI, staff productivity, and organization-level metrics. This is harder to achieve. For instance, ZS interviewed R&D leaders at large pharma and found ROI estimates that remain "finger-in-the-air", with organizations leaning on time saved or model accuracy and lacking baselines to attribute impact to anything. EY argues for economic metrics instead: time to candidate selection, R&D cost to IND, probability of program success, cost per program, etc. In our own BiopharmaTrend analysis of AI drug discovery pipelines, we measured internal drug candidate pipeline progression over the number of years a biotech startup is active, as a productivity proxy to assess AI platforms (although this metric would mostly be irrelevant to big pharma or large corporations with hundreds of candidates and historical research).
Anyway, beyond the value of benchmarking, this study by Insilico Medicine led to an interesting takeaway: smaller specialized models can actually be better for scientific reasoning than the largest frontier models, at least in some contexts.
That finding describes what much of the market has already built toward. Schrödinger opened early access to Bunsen this summer with positioning that leaves nothing implicit: unlike general-purpose systems that "only retrieve and summarize existing scientific information," Bunsen executes the company's validated computational methods, designing a modeling strategy, running the calculations and interpreting results across three decades of physics-based simulation. Lantern Pharma has spun Open-Medicine AI out as a separate company, arguing its agents carry operating knowledge from the clinical programs they were built inside, and is extending them downstream into trial design, patient selection and regulatory strategy. Owkin's K Pro runs on its MOSAIC multimodal patient network. And Causaly grounds every answer in proprietary biomedical knowledge graphs, each claim traceable to a source the user can open.
Pharma is buying into that side of the AI market as well. Takeda's agreement with Iambic Therapeutics, reported as potentially worth a billion dollars or more, backs specialized machine learning for molecular design. Lilly's arrangement is with Insilico Medicine; Servier licensed K Pro from Owkin bundled with MOSAIC data, joining Boehringer Ingelheim on the same platform. BenchSci, which builds on a curated biological knowledge graph rather than a general model alone, has signed argenx to deploy EMET, its agentic research environment reasoning across 38 million publications and more than 100 curated databases, adding to multi-year agreements with Merck and Sanofi. And Bristol Myers Squibb is deploying Bunsen across its research organization at scale and co-developing new functionality with Schrödinger, including an AI synthesis planning platform.
The same companies signing frontier-lab deals are also paying for grounded alternatives to have diversity of tools and R&D workflows.
The user base of agentic-driven tools is widening
To conclude the reflections on the present momentum for LLMs and agentic systems in the life sciences, there are many open questions and evolving best practices, but one thing is clear already today: the user base of life science experts using AI tools in their daily work is expanding.
Bioinformatics tools traditionally used to have quite a narrow user base. Now, with adding LLM-driven agentic layers on top of tools and platforms, more categories of scientists, managers, and business leaders can engage in cross-disciplinary work. Setting up and interpreting a free-energy calculation is specialist work, and Schrödinger built Bunsen to remove that constraint. Its launch materials say the system makes advanced computational methods accessible to "experienced drug hunters who are not computational chemists," and CEO Ramy Farid tied that to the company's economics: Bunsen is "expanding the user base for our computational platform," and "our throughput-based licensing model is ideally suited to capture the value of this expanding utilization."
Robert Abel, Schrödinger's chief scientific officer for platform, said adopting Bunsen at scale "will empower a broader group of scientists to embrace a predict-first computational approach." From the buyer's side, BMS vice president of computational sciences Stephen Johnson said Bunsen "allows our scientists to think differently about how physics-based tools can be used to navigate molecular design space" — not the computational group's scientists, the company's.
Causaly is pushing the boundary further out. Co-founder and CEO Yiannis Kiachopoulos said extending the collaboration to Microsoft Copilot "puts governed, provenance-first scientific reasoning inside the same interface used across research, development, medical, and commercial functions in biopharma."
BenchSci makes the same argument from the demand side. Its EMET agentic environment was picked by argenx's scientists after their own competitive evaluation — which CEO Liran Belenzon called earning trust "at the bench, not in the boardroom". argenx's Tim Van Acker described it as "a scientist in your pocket – always available, always reasoning across the evidence."
So, the industry is moving forward with AI, LLMs, and agentic processes on top of all that, and the questions of productivity and ROI evaluation, economics, and workflow design, will be central points of interest in 2027, I believe.
Go deeper on the rise of LLMs and agentic AI in pharma
Everyone is Launching AI Agents. What's Being Deployed? — our premium deep dive into what has actually shipped: Stanford's "Virtual Biotech" running 37,000 agents in parallel for $46 in API credits, AstraZeneca's honest account of what broke in a live multi-agent pipeline, IQVIA's 150 agents, Roche's 3,500 GPUs — and the reliability data showing why identical prompts still produce different answers 30% of the time
Topic: AI in Bio