[Paper Review] Inferring genotyping error rates from genotyped trios
This paper presents a likelihood-based method to infer genotyping error rates from trio genotypes using Mendelian inheritance and allele frequency data, enabling accurate maximum likelihood estimation. Applied to 23andMe trio data, it estimates a low error rate of 8.5 × 10⁻⁵ (95% CI: 6.8–10.2 × 10⁻⁵), demonstrating high genotyping accuracy in personal genomics data.
Genotyping errors are known to influence the power of both family-based and case-control studies in the genetics of complex disease. Estimating genotyping error rate in a given dataset can be complex, but when family information is available error rates can be inferred from the patterns of Mendelian inheritance between parents and offspring. I introduce a novel likelihood-based method for calculating error rates from family data, given known allele frequencies. I apply this to an example dataset, demonstrating a low genotyping error rate in genotyping data from a personal genomics company.
Motivation & Objective
- To develop a robust statistical method for estimating genotyping error rates from family trio data, addressing limitations of existing Mendelian error rate metrics.
- To incorporate both Mendelian-inconsistent genotypes and the full distribution of genotype frequencies under Hardy-Weinberg equilibrium to improve error rate estimation.
- To provide a likelihood-based framework that accounts for both error-prone genotypes and the absence of errors across sites, increasing estimation accuracy.
- To apply the method to real-world data from a personal genomics company to assess genotyping quality in high-throughput genotyping platforms.
Proposed method
- The method models the joint likelihood of observed and true genotypes across N sites in a trio, integrating Mendelian inheritance rules and Hardy-Weinberg equilibrium genotype frequencies.
- It uses a simplified error model assuming only single errors occur (neglecting ε² and higher terms), with error probabilities (1−ε)⁶ for no errors and ε(1−ε)⁵ for one error.
- The likelihood is partitioned into terms where observed genotypes match true genotypes and terms with one genotypic error, enabling efficient computation via pre-calculation of genotype frequencies.
- Maximum likelihood estimation is performed by optimizing the likelihood function over the error rate ε, using precomputed genotype frequency terms.
- The method accounts for allele frequencies from external sources (e.g., HapMap 3 CEU data) to inform prior genotype probabilities under Hardy-Weinberg equilibrium.
- Confidence intervals for the error rate are derived using standard likelihood profile methods.
Experimental results
Research questions
- RQ1Can a likelihood-based method improve genotyping error rate estimation compared to Mendelian error rate metrics alone?
- RQ2How does incorporating both Mendelian-inconsistent genotypes and the full genotype frequency spectrum enhance error rate inference?
- RQ3What is the true genotyping error rate in high-throughput genotyping data from a personal genomics company like 23andMe?
- RQ4To what extent do rare errors in genotyping affect the power and reliability of genetic association studies?
Key findings
- The method successfully estimates a genotyping error rate of 8.5 × 10⁻⁵ for a trio genotyped by 23andMe, using 715,566 variants and HapMap 3 CEU allele frequencies.
- The 95% confidence interval for the error rate is 6.8 × 10⁻⁵ to 10.2 × 10⁻⁵, indicating high precision in the estimate.
- The low error rate suggests that 23andMe’s genotyping platform maintains high accuracy, which is critical for reliable case-control association studies.
- The method outperforms traditional Mendelian error rate metrics by incorporating information from both error-prone and error-free sites, improving estimation robustness.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.