Skip to main content
QUICK REVIEW

[Paper Review] Deep Learning in Bioinformatics

Seonwoo Min, Byunghan Lee|arXiv (Cornell University)|Mar 21, 2016
Genetics, Bioinformatics, and Biomedical Research194 references21 citations
TL;DR

This paper reviews the application of deep learning in bioinformatics, categorizing research by domain—omics, biomedical imaging, and signal processing—and by architecture, including deep neural networks, CNNs, RNNs, and emerging models. It highlights state-of-the-art performance in data-driven biological insight generation and outlines theoretical and practical challenges, offering a comprehensive guide for researchers entering the field.

ABSTRACT

In the era of big data, transformation of biomedical big data into valuable knowledge has been one of the most important challenges in bioinformatics. Deep learning has advanced rapidly since the early 2000s and now demonstrates state-of-the-art performance in various fields. Accordingly, application of deep learning in bioinformatics to gain insight from data has been emphasized in both academia and industry. Here, we review deep learning in bioinformatics, presenting examples of current research. To provide a useful and comprehensive perspective, we categorize research both by the bioinformatics domain (i.e., omics, biomedical imaging, biomedical signal processing) and deep learning architecture (i.e., deep neural networks, convolutional neural networks, recurrent neural networks, emergent architectures) and present brief descriptions of each study. Additionally, we discuss theoretical and practical issues of deep learning in bioinformatics and suggest future research directions. We believe that this review will provide valuable insights and serve as a starting point for researchers to apply deep learning approaches in their bioinformatics studies.

Motivation & Objective

  • To provide a comprehensive review of deep learning applications in bioinformatics across key domains.
  • To categorize current research by both bioinformatics domain and deep learning architecture for better accessibility.
  • To identify theoretical and practical challenges in applying deep learning to biological data.
  • To suggest future research directions for integrating deep learning into bioinformatics workflows.
  • To serve as a starting point and reference for researchers seeking to apply deep learning in biological data analysis.

Proposed method

  • Systematic categorization of deep learning applications in bioinformatics by domain: genomics (omics), biomedical imaging, and biomedical signal processing.
  • Classification of deep learning architectures: deep neural networks, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and emergent architectures.
  • Review of representative studies in each category, highlighting model design, data types, and performance outcomes.
  • Analysis of theoretical issues such as interpretability, generalization, and data scarcity in biological contexts.
  • Discussion of practical challenges including computational demands, hyperparameter tuning, and model validation in bioinformatics.
  • Synthesis of future research directions based on current limitations and emerging opportunities.

Experimental results

Research questions

  • RQ1How can deep learning architectures be effectively applied to diverse bioinformatics domains such as genomics, imaging, and signal processing?
  • RQ2What are the key architectural choices (e.g., CNNs, RNNs) that enable state-of-the-art performance in specific bioinformatics tasks?
  • RQ3What are the major theoretical and practical challenges in deploying deep learning models on biological data?
  • RQ4How do current deep learning models compare to traditional methods in terms of accuracy and interpretability in bioinformatics applications?
  • RQ5What future research directions are most promising for advancing deep learning in bioinformatics?

Key findings

  • Deep learning models, particularly CNNs and RNNs, achieve state-of-the-art performance in tasks such as gene expression prediction, protein structure modeling, and medical image analysis.
  • The integration of deep learning with omics data enables the discovery of complex, non-linear patterns in genomic and transcriptomic data.
  • Despite strong performance, challenges remain in model interpretability, generalization, and data efficiency, especially with limited or noisy biological datasets.
  • Emergent architectures such as autoencoders and generative models show promise in dimensionality reduction and data generation for biological applications.
  • The review identifies a growing trend toward end-to-end learning pipelines that integrate multiple data types and model components.
  • Future research should focus on improving model transparency, reducing data dependency, and enhancing reproducibility in deep learning applications for bioinformatics.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.