Submit News
UVA Health logo of UVA Health Submit News

Connect

9.10.2026

Test of Large Language Models Reveals Promise, Pitfalls for Biomedical Research

New University of Virginia School of Medicine research highlights the potential of artificial-intelligence tools known as large language models to accelerate our understanding of cellular processes and advance the development of new disease treatments. But the work also underscores the limitations of existing technology and the need for careful human expertise.

Jeff Saucerman, PhD

Researchers led by UVA’s Jeff Saucerman, PhD, tested whether large language models (LLMs) including GPT, Gemini, and Claude can explain biological processes within cells and predict how cell functions are disrupted as occurs in disease. The scientists found that the LLMs were reasonably good at the former — knowing the parts catalog — but struggled far more with the latter. The models accurately predicted how cells response to disruptions less than a third of the time.

The results suggest the technology is poised to become an invaluable research tool but needs refinement and close human oversight, Saucerman says. 

“Our study shows that general-purpose LLMs can now find key aspects of the communication networks inside cells, but their knowledge has significant gaps that prevent accurate predictions a cell’s response to drugs or gene mutations. They still need help from humans,” said Saucerman, part of UVA’s Department of Biomedical Engineering, a joint program of the School of Medicine and School of Engineering. “Still, these AI models are improving quickly, and ultimately they have tremendous potential to ultimately help researchers develop a more comprehensive understanding of biological systems.”

Large Language Model Limitations

Computer models of cellular “signaling networks” — communication chains within and among cells — are already widely used by scientists to study fundamental biological processes and guide experiments to better understand and treat disease. These models, however, require a great deal of manual input based on prior scientific research performed and validated by human beings. So there is great excitement about the potential of AI to automate, enhance and accelerate this laborious work. While many academic and industry researchers are using these AI models, the quality of their predictions has been less often tested rigorously.

Saucerman’s study examined the ability of LLMs ability to model known signaling networks that control the ability of heart cells to respond to external stimuli, grow and remodel their environment. He and his team found that the LLMs were able to automatically generate up to 65% of the reactions that had been previously identified by human curation. The researchers describe the simulations as “moderately” accurate.

“While the LLMs correctly identified many of the more generic communication signals seen in most cells of your body, they were less successful in identifying relationships that are specific to certain types of heart cells,” said researcher Jeevan Tewari. “Our study demonstrates the need to rigorously test the quality of AI predictions for how cells work. The results highlight their ability to capture well-conserved pathways while also revealing limitations in their ability to model specialized biological processes.”

When it came to predicting the effects of disruptions to normal cellular processes — the drivers of disease — the LLMs were far less effective. The models predicted outcomes accurately only between 6% and 33% of the time.

That’s not to say the models don’t have value for both purposes. But they need further development and close supervision by human scientists, the UVA researchers caution. Their work identifies strengths and weaknesses in the models and aims to give researchers a “pipeline” to build better ones and use them wisely. For example, the UVA team is calling for the development of “rigorous human-created platforms” for organizing and scrutinizing the information the models produce. 

“We hope this work inspires further research into the accuracy and applications of LLMs, instead of immediately using AI in new areas without checking and refining the process first,” researcher Benjamin Dahl said. “While AI has tremendous potential research applications, it isn’t a magic wand that instantly solves any problem. AI will continue to improve and become more sophisticated, but it is a tool to be used carefully by scientists. Rigorous quality-control checks and benchmarking tests will keep these AI tools honed and accurate for future discoveries.”

Saucerman’s efforts align closely with the goals of UVA’s new Paul and Diane Manning Institute of Biotechnology, which has been launched to accelerate the development of new drugs and cures for the most complex and challenging diseases. The institute aims to speed how quickly lab discoveries can be turned into medicines to benefit patients across Virginia and beyond.

Findings Published

Saucerman and his team have published their findings in the scientific journal eLife. The article is open access, meaning it can be read for free. The research team consisted of Jeevan Tewari, Benjamin W. Dahl, B. Adam Bates, Jason A. Papin, and Saucerman

The research was supported by the National Institutes of Health, grants R01HL162925, R01HL160665, R01HL172417, T32GM156694, T32GM145443 and R01GM147257, and by a pilot grant from UVA Comprehensive Cancer Center.

To keep up with the latest medical research news from UVA and the Manning Institute, bookmark the Making of Medicine blog.

Comments (0)

Latest News