Skip to Content
SBMI Horizontal Logo

Teaching Large Language Models to Understand Biology: Evidence-Grounded Agentic Systems for Brain Cell Type Annotation

Author: Rongbin Li (2026)

Primary advisor: W. Jim Zheng, PhD

Committee members: Xiaobo Zhou, PhD and Hongfang Liu, PhD

PhD thesis, McWilliams School of Biomedical Informatics at UTHealth Houston.


ABSTRACT

The adoption of single-cell RNA sequencing has enabled large-scale discovery of brain cell types; however, functional interpretation of novel and rare cell populations remains a fundamental bottleneck. Conventional annotation approaches rely on curated reference databases which are often incomplete for newly identified marker genes. Although large language models (LLMs) offer the potential to synthesize information at scale, they are primarily trained on general-domain text and therefore lack structured biological knowledge. This limitation, coupled with hallucination and insufficient linkage to primary evidence, constrains their ability to infer gene function and cell identity.

This dissertation develops a set of integrated approaches to address core challenges in enabling LLMs to perform biologically grounded functional reasoning. I introduce BRAINCELL-AID (Brain Cell type Annotation and Integration using Distributed AI), an agentic computational framework designed to systematically infuse biological knowledge into LLMs and guide their reasoning toward evidence-supported functional interpretation. The work first fine-tunes a base LLM on more than 7,000 curated gene sets to incorporate domain-specific biological knowledge and enable the model to learn gene co-functionality. It then combines large-scale literature mining with retrieval-augmented generation to ground inferred functions in primary evidence, reducing hallucinations. These components are integrated within an agentic AI system that coordinates knowledge injection, evidence retrieval, and reasoning to support robust and evidence based biological annotation.

BRAINCELL-AID demonstrates improved performance over state-of-the-art annotation methods and enables scalable functional analysis of marker gene sets. Applied at scale, the framework synthesizes functional annotations and identifies neurological signature gene sets for more than 5,300 novel cell types across the adult mouse brain. Hence, I establish a reusable knowledge resource for the neuroscience community and propose a generalizable framework for advancing LLMs from passive text synthesis toward active, evidence-grounded biological reasoning.

Keywords: brain function, neuroscience, gene set analysis, large language models, artificial intelligence agents, data resource