Skip to main content
QUICK REVIEW

[Paper Review] RNA structure characterization from chemical mapping experiments

Sharon Aviran, Julius B. Lucks|arXiv (Cornell University)|Jun 24, 2011
RNA and protein synthesis mechanisms18 references4 citations
TL;DR

This paper introduces a unified maximum-likelihood model for inferring RNA reactivity from chemical mapping experiments using either capillary electrophoresis (SHAPE-CE) or high-throughput sequencing (SHAPE-Seq). The model yields closed-form estimates that improve accuracy by leveraging full-length transcript information in sequencing-based methods, demonstrating that SHAPE-Seq provides more reliable reactivity estimates than SHAPE-CE due to better signal recovery in the 5′ region of RNA molecules.

ABSTRACT

Despite great interest in solving RNA secondary structures due to their impact on function, it remains an open problem to determine structure from sequence. Among experimental approaches, a promising candidate is the "chemical modification strategy", which involves application of chemicals to RNA that are sensitive to structure and that result in modifications that can be assayed via sequencing technologies. One approach that can reveal paired nucleotides via chemical modification followed by sequencing is SHAPE, and it has been used in conjunction with capillary electrophoresis (SHAPE-CE) and high-throughput sequencing (SHAPE-Seq). The solution of mathematical inverse problems is needed to relate the sequence data to the modified sites, and a number of approaches have been previously suggested for SHAPE-CE, and separately for SHAPE-Seq analysis. Here we introduce a new model for inference of chemical modification experiments, whose formulation results in closed-form maximum likelihood estimates that can be easily applied to data. The model can be specialized to both SHAPE-CE and SHAPE-Seq, and therefore allows for a direct comparison of the two technologies. We then show that the extra information obtained with SHAPE-Seq but not with SHAPE-CE is valuable with respect to ML estimation.

Motivation & Objective

  • To develop a general statistical framework for inferring RNA reactivity from chemical mapping data that applies to both SHAPE-CE and SHAPE-Seq.
  • To address the limitation in SHAPE-CE where only the first modification in a transcript is detected, leading to incomplete signal recovery.
  • To compare the information content and estimation accuracy of SHAPE-CE versus SHAPE-Seq using a common likelihood-based inference model.
  • To demonstrate that the additional full-length transcript data in SHAPE-Seq improves maximum-likelihood reactivity estimation, especially in the 5′ region of RNA molecules.
  • To provide a closed-form solution for maximum-likelihood estimation of reactivity and reverse transcriptase dropoff probabilities, enabling fast and robust analysis.

Proposed method

  • Formulates a joint likelihood model for chemical modification reactivity (βk) and reverse transcriptase dropoff (γk) at each nucleotide position k.
  • Derives closed-form maximum-likelihood estimates for βk and γk by optimizing the log-likelihood function under constraints 0 ≤ βk ≤ 1 and 0 ≤ γk ≤ 1.
  • Applies the model to both SHAPE-CE and SHAPE-Seq data by modeling the observed fragment counts from cDNA synthesis and reverse transcription.
  • Uses a constrained optimization approach to ensure estimates remain within the feasible parameter space, with theoretical justification via compact set and stationary point analysis.
  • Specializes the general model to both SHAPE-CE and SHAPE-Seq by adjusting for differences in data representation and signal detection.
  • Validates the model by comparing ML estimates from CE and NGS data on real biological RNAs, such as pT181 and RNase P.

Experimental results

Research questions

  • RQ1Does the inclusion of full-length transcript information in SHAPE-Seq improve the accuracy of reactivity estimation compared to SHAPE-CE?
  • RQ2Can a unified statistical model be formulated for both SHAPE-CE and SHAPE-Seq that yields closed-form maximum-likelihood estimates?
  • RQ3How does the absence of full-length signal in CE-based methods affect the reliability of reactivity estimates, particularly in the 5′ region of RNA molecules?
  • RQ4To what extent does the extra information in SHAPE-Seq enhance the precision of reactivity inference compared to CE-based methods?
  • RQ5Is the likelihood-based estimation framework more accurate than visual correction methods used in prior work, and what are its limitations?

Key findings

  • The proposed model yields closed-form maximum-likelihood estimates for reactivity (βk) and reverse transcriptase dropoff (γk), enabling fast and robust inference without iterative optimization.
  • SHAPE-Seq provides more accurate reactivity estimates than SHAPE-CE because it captures full-length transcript signals, which are missing in CE-based detection.
  • In the 5′ region of RNA molecules, SHAPE-CE leads to over-estimation of reactivity due to incomplete signal recovery, while SHAPE-Seq estimates remain more accurate.
  • For the Staphylococcus aureus pT181 RNA, the fraction of modified molecules was estimated as 0.9 using CE and 0.86 using sequencing, indicating minor divergence.
  • For the Bacillus subtilis RNase P RNA, the divergence was more pronounced, with CE estimating 0.62 and sequencing estimating 0.52 for the fraction of modified molecules.
  • The likelihood-based framework outperforms visual correction methods in reproducibility but reveals a systematic bias in CE data due to missing full-length signals, which the model explicitly accounts for in NGS-based inference.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.