Skip to main content
QUICK REVIEW

[Paper Review] The big challenge for livestock genomics is to make sequence data pay

M. Johnsson|arXiv (Cornell University)|Feb 2, 2023
Genetic and phenotypic traits in livestockBiochemistry, Genetics and Molecular Biology3 citations
TL;DR

This paper argues that despite the theoretical advantages of whole-genome sequencing (WGS) in livestock genomics, current implementations fail to consistently improve genomic prediction accuracy over SNP chips due to noise from non-causal variants and high computational costs. The key contribution is advocating for pre-selection of functionally relevant variants from WGS data—rather than using all variants—combined with functional genomics to make sequence data economically and statistically viable for breeding programs.

ABSTRACT

This paper will argue that one of the biggest challenges for livestock genomics is to make whole-genome sequencing and functional genomics applicable to breeding practice. It discusses potential explanations for why it is so difficult to consistently improve the accuracy of genomic prediction by means of whole-genome sequence data, and three potential attacks on the problem.

Motivation & Objective

  • To address the persistent challenge of making whole-genome sequencing data deliver consistent improvements in genomic prediction accuracy over SNP chips in livestock breeding.
  • To investigate why WGS data often underperforms or even degrades prediction accuracy compared to SNP chip data.
  • To evaluate whether functional genomics and pre-selection of causal or regulatory variants can overcome the limitations of using millions of raw sequence variants.
  • To assess the economic and computational feasibility of replacing SNP chips with WGS in large-scale breeding programs.
  • To explore alternative uses of sequence and functional genomic data beyond direct genomic prediction, including microbiome and epigenomic applications.

Proposed method

  • Analyzing empirical and simulated data from livestock genomics studies comparing SNP chip-based and WGS-based genomic prediction.
  • Evaluating the 'dilution effect' where non-causal variants in WGS data reduce prediction accuracy by introducing noise.
  • Proposing a two-step approach: first, use WGS data to identify variants enriched for trait associations via genome-wide association studies (GWAS) and functional annotations; second, use only these selected variants in prediction models.
  • Leveraging functional genomics data such as open chromatin regions and eQTLs to prioritize variants with biological relevance.
  • Assessing the cost-benefit trade-off of routine WGS versus SNP chip genotyping, including storage, imputation, and analysis overhead.
  • Modeling the potential for improved persistence and cross-population accuracy using causative variants identified through WGS and functional data.

Experimental results

Research questions

  • RQ1Why does whole-genome sequencing fail to consistently improve genomic prediction accuracy compared to SNP chips in livestock?
  • RQ2To what extent does the inclusion of non-causal variants in WGS data lead to a 'dilution effect' that reduces prediction accuracy?
  • RQ3Can pre-selection of functionally relevant variants from WGS data significantly improve prediction accuracy and economic return in breeding programs?
  • RQ4What are the computational and economic barriers to replacing SNP chips with WGS in routine livestock breeding applications?
  • RQ5Can functional genomics data (e.g., eQTLs, chromatin states) enhance the utility of WGS data for genomic prediction without requiring full-scale sequencing of all animals?

Key findings

  • Empirical studies show that using all variants from whole-genome sequencing often results in no improvement or even a decrease in genomic prediction accuracy compared to SNP chips.
  • The 'dilution effect'—where non-causal variants reduce model accuracy—explains much of the inconsistent performance of WGS in genomic prediction.
  • Pre-selecting variants based on functional genomics data (e.g., eQTLs, open chromatin regions) or GWAS results significantly improves prediction accuracy compared to using all WGS variants.
  • Sequence data has the potential to detect causative variants and de novo mutations more effectively than SNP chips, but only if combined with functional annotation and selection strategies.
  • Current sample sizes for functional genomics studies (e.g., eQTL mapping) are too small to detect most associations, limiting their utility for prediction.
  • The high cost of storage, imputation, and analysis of WGS data makes it economically unjustifiable for routine use unless it delivers consistent, large gains in prediction accuracy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.