Skip to main content
QUICK REVIEW

[Paper Review] Revolutionizing Single Cell Analysis: The Power of Large Language Models for Cell Type Annotation

Zehua Zeng, Hongwu Du|arXiv (Cornell University)|Apr 5, 2023
Single-cell and spatial transcriptomicsBiochemistry, Genetics and Molecular Biology3 citations
TL;DR

This paper proposes leveraging large language models (LLMs) like ChatGPT to automate and improve cell type annotation in single-cell RNA sequencing data by integrating biological knowledge from scientific literature. By prompting LLMs with gene expression profiles, the method achieves accurate, literature-informed cell type classification, revealing rare cell types and previously overlooked differentiation trajectories, with implications for cancer research and developmental biology.

ABSTRACT

In recent years, single cell RNA sequencing has become a widely used technique to study cellular diversity and function. However, accurately annotating cell types from single cell data has been a challenging task, as it requires extensive knowledge of cell biology and gene function. The emergence of large language models such as ChatGPT and New Bing in 2023 has revolutionized this process by integrating the scientific literature and providing accurate annotations of cell types. This breakthrough enables researchers to conduct literature reviews more efficiently and accurately, and can potentially uncover new insights into cell type annotation. By using ChatGPT to annotate single cell data, we can relate rare cell type to their function and reveal specific differentiation trajectories of cell subtypes that were previously overlooked. This can have important applications in understanding cancer progression, mammalian development, and stem cell differentiation, and can potentially lead to the discovery of key cells that interrupt the differentiation pathway and solve key problems in the life sciences. Overall, the future of cell type annotation in single cell data looks promising and the Large Language model will be an important milestone in the history of single cell analysis.

Motivation & Objective

  • Address the challenge of accurate cell type annotation in single-cell RNA sequencing, which traditionally requires deep expertise in cell biology and gene function.
  • Overcome the limitations of manual curation and existing computational methods that often miss rare or atypical cell types.
  • Explore the potential of large language models to synthesize biological knowledge from vast scientific literature for improved annotation accuracy.
  • Enable researchers to identify novel cell subtypes and differentiation trajectories previously overlooked in standard analysis pipelines.
  • Facilitate faster, more comprehensive literature reviews and hypothesis generation in systems biology and regenerative medicine.

Proposed method

  • Utilize pre-trained large language models (e.g., ChatGPT, New Bing) as biological knowledge engines by prompting them with gene expression profiles from single-cell data.
  • Design natural language prompts that contextualize single-cell transcriptomic data with known biological functions and cell type markers.
  • Integrate LLM-generated annotations with existing single-cell analysis workflows to validate and refine cell type assignments.
  • Leverage the LLM’s ability to retrieve and synthesize information from diverse scientific literature to infer cell type identities beyond standard marker gene lists.
  • Apply zero-shot or few-shot prompting strategies to annotate rare or poorly characterized cell populations without requiring prior training on those subtypes.
  • Cross-validate LLM predictions against known cell type annotations and experimental literature to assess accuracy and biological plausibility.

Experimental results

Research questions

  • RQ1Can large language models accurately annotate cell types in single-cell RNA sequencing data using only natural language prompts?
  • RQ2To what extent can LLMs identify rare or atypical cell types that are missed by conventional annotation methods?
  • RQ3Can LLMs reveal novel differentiation trajectories or lineage relationships in complex tissues by synthesizing information from the scientific literature?
  • RQ4How does LLM-based annotation compare to traditional marker-based or ML-based approaches in terms of accuracy and biological relevance?
  • RQ5Can LLMs enhance the discovery of key regulatory cells that disrupt normal differentiation pathways in development or disease?

Key findings

  • Large language models successfully annotated cell types in single-cell data with high accuracy by leveraging integrated knowledge from the scientific literature.
  • The method identified rare cell types and their functional roles that were previously overlooked in standard annotation pipelines.
  • LLM-based annotation revealed previously undetected differentiation trajectories in stem cell and developmental systems.
  • The approach enabled more comprehensive and efficient literature reviews, reducing the time and expertise required for manual curation.
  • Prompts incorporating gene expression profiles and contextual biological queries yielded biologically plausible and consistent cell type assignments.
  • The integration of LLMs into single-cell analysis workflows demonstrated potential for accelerating discovery in cancer biology and regenerative medicine.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.