Skip to main content
QUICK REVIEW

[논문 리뷰] Fast reconstruction of compact context-specific metabolic networks via integration of microarray data

Maria Pires Pacheco, Thomas Sauter|arXiv (Cornell University)|2014. 07. 24.
Microbial Metabolic Engineering and Bioproduction참고 문헌 24인용 수 4
한 줄 요약

이 논문은 마이크로어레이 데이터를 유전체 수준 대사 모델(GEMs)과 통합하여 몇 분 내에 컴act하고 맥락 특화된 대사 네트워크를 재구성하는 빠르고 강력한 워크플로우인 FASTCORMICS를 소개한다. 임계값이 없는 유전자 발현 이산화를 위한 Barcode 알고리즘을 사용하고, 이를 FASTCORE의 효율적인 선형 프로그래밍 기반 접근 방식과 결합함으로써, 기존 방법보다 정확도와 강력성 면에서 뛰어난 예측 성능을 보이는 핵심 유전자를 위한 고품질의 모델을 생성한다.

ABSTRACT

Recently we proposed an algorithm for the fast reconstruction of compact context-specific metabolic networks (FASTCORE) that allowed dropping the reconstruction time to the time order of seconds (Vlassis et al.,2014). This extremely low computational demand opens new possibilities for improving the quality of the models. Several rounds of model reconstruction, testing of the model's predictions against real experimental data, curation steps of the input model and the set of core reactions as well as cross-validations assays are required to reconstruct high-quality models. These semi-automated model curations steps are in such extend not possible with competing algorithms due to their high computational demands. To adapt FASTCORE for the integration of microarray data, we therefore propose a new workflow: FASTCORMICS. FASTCORMICS requires as input microarray data and a Genome-scale reconstruction. FASTCORMICS is devoid of heuristic parameter settings and has a low computational demand with overall building times in the order of a few minutes. FASTCORMICS preprocesses the microarrays data with the discretization tool Barcode (Zillox et al, 2007). Barcode uses prior knowledge on the intensity distribution of each probe set for a given microarray platform to segregate between expressed genes and non-expressed genes. This preprocessing step allows circumventing the need of setting a heuristic expression threshold, which is critical for the output models as in response to this threshold alternative pathways or subsystems might be included or excluded, thereby heavily changing the functionalities of the model. In general, FASTCORMICS outperforms competing algorithms and allows obtaining high-quality, robust models in a high-throughput manner. This will allow the use of metabolic modelling as routine process for the analysis of microarray data e.g. in the field of personalized medicine.

연구 동기 및 목표

  • 마이크로어레이 데이터로부터 고품질의 컴 pact한 맥락 특화 대사 모델을 생성하는 데 도전한다.
  • 기존 방법에서 히وري스틱 발현 임계값의 한계를 극복하여 모델 구조와 예측 성능에 대한 편향을 줄인다.
  • 모델 재구성에 소요되는 계산 시간을 몇 분으로 줄여 반복적 코딩과 고속 스트림라인 분석을 가능하게 한다.
  • 기본 GEMs와 경쟁 알고리즘에 비해 암 모델에서 핵심 유전자 식별의 정확도를 향상시킨다.
  • 감도가 낮고, 파라미터가 없는 워크플로우를 개발하여 마이크로어레이 데이터의 배치 효과와 플랫폼 특화 노이즈에 대한 민감도를 최소화한다.

제안 방법

  • 프로브 강도 분포의 사전 지식을 활용하여 발현 여부를 분류하는 이산화 도구인 Barcode를 사용하여 마이크로어레이 데이터를 전처리함으로써, 임의의 발현 임계값을 피한다.
  • Barcode로 식별된 발현 유전자들을 유전자-단백질-반응(GPR) 규칙를 통해 반응에 매핑하여 모델 재구성의 핵심 반응 집합을 정의한다.
  • FASTCORE라는 빠른 선형 프로그래밍 기반 알고리즘을 적용하여, 핵심 반응 전부를 통해 유량이 가능하도록 하기 위해 필요한 최소 비핵심 반응 집합을 식별한다.
  • 비핵심 반응의 포함을 제재하여 복잡성이 최소화되고 생물학적으로 관련성이 있는 컴 pact한 모델을 확보한다.
  • 선택적으로 배지 조성과 생체량 함수를 제약 조건으로 설정하여 생리적 조건을 반영한다.
  • shRNA 내림줄이기 데이터를 활용한 핵심성 실험을 통해 모델을 검증하고, DisGeNET에서 유래한 종양형질 관련 유전자들의 과잉 표현을 초기 기하확률 검정을 통해 평가한다.

실험 결과

연구 질문

  • RQ1임계값이 없는 방식으로 마이크로어레이 데이터를 유전체 수준 대사 모델에 통합할 수 있는가?
  • RQ2FASTCORMICS 워크플로우는 일반 GEMs나 경쟁 알고리즘에 비해 핵심 유전자에 대한 예측 정확도가 더 높은 맥락 특화 모델을 생성하는가?
  • RQ3Barcode를 사용한 발현 이산화가 임계값 기반 방법에 비해 모델의 강력성과 배치 효과에 대한 민감도에 어떤 영향을 미치는가?
  • RQ4MBA, GIMME, 또는 IMAT와 같은 기존 알고리즘에 비해 FASTCORMICS는 계산 효율성과 예측 성능 면에서 어느 정도 뛰어나게 성능을 발휘하는가?
  • RQ5이 워크플로우는 개인 맞춤 의료와 같은 일상적인 시스템 생물학 응용 분야를 지원하기 위해 고속 스트림라인으로 적용될 수 있는가?

주요 결과

  • FASTCORMICS는 몇 분 내에 맥락 특화 대사 모델을 재구성하여 기존 방법에 비해 계산 시간을 크게 단축시킨다.
  • Recon 1에서 유도된 cancer1 모델은 핵심성 테스트에서 KS-스코어 p-값 8.1e-07을 기록하여 MBA 알고리즘의 p-값 0.0044를 초월한다.
  • Recon 2에서 유도된 cancer2 모델은 KS-스코어 p-값 2.9e-04를 기록하여 핵심 유전자 예측에서 통계적으로 유의미한 결과를 보였다.
  • 초기 기하확률 검정을 통해 모든 모델에서 DisGeNET의 종양형질 관련 유전자들이 예측된 핵심 유전자 집합에서 유의미하게 과잉 표현됨을 확인하여 생물학적 관련성을 검증하였다.
  • cancer1 모델은 188개의 핵심 반응과 90개의 비핵심 반응을 포함하며, 총 1168개의 반응과 377개의 유량 반응을 기록하여 컴 pact함과 기능적 일관성을 입증하였다.
  • Barcode에 의존함으로써 히وري스틱 임계값이 필요 없어졌고, 이는 모델 편향을 감소시키고 데이터 변동성에 대한 강력성을 향상시켰다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.