早稲田大学 · 免疫学・微生物学
伊井寿教授の研究室は、高スループットシーケンシングで急速に蓄積されるゲノム・トランスクリプトーム・プロテオームデータの解析に、自然言語処理(NLP)のアプローチを応用する研究を展開しています。特に、生物学的配列を「文章」とみなし、k-mersを「単語」として扱う表現学習技術の開発が柱であり、時間的変化を示すオミックスデータからの周期的パターン同定や、ウイルスと宿主の相互作用予測にも応用しています。また、単細胞RNA-Seqを用いた細胞動態の可視化や、偽時刻(pseudotime)に伴う周期的発現の同定にも注力しています。
Figures are computed from collected data and may differ slightly.
Although remarkable advances have been reported in high-throughput sequencing, the ability to aptly analyze a substantial amount of rapidly generated biological (DNA/RNA/protein) sequencing data remains a critical hurdle. To tackle this issue, the application of natural language processing (NLP) to biological sequence analysis has received increased attention. In this method, biological sequences are regarded as sentences while the single nucleic acids/amino acids or k-mers in these sequences re
The coronavirus disease-2019 (COVID-19) pandemic has elucidated major limitations in the capacity of medical and research institutions to appropriately manage emerging infectious diseases. We can improve our understanding of infectious diseases by unveiling virus-host interactions through host range prediction and protein-protein interaction prediction. Although many algorithms have been developed to predict virus-host interactions, numerous issues remain to be solved, and the entire network rem
In this paper, we presented MICOP, which is an MIC-based algorithm, for predicting periodic patterns in large-scale time-resolved protein expression profiles. The performance test using artificially generated simulation data revealed that the performance of MICOP for decaying data was superior to that of the existing widely used methods. It can reveal novel findings from time-series data and may contribute to biologically significant results. This study suggests that MICOP is an ideal approach f
ABSTRACT Remarkable advances in high-throughput sequencing have resulted in rapid data accumulation, and analyzing biological (DNA/RNA/protein) sequences to discover new insights in biology has become more critical and challenging. To tackle this issue, the application of natural language processing (NLP) to biological sequence analysis has received increased attention, because biological sequences are regarded as sentences and k-mers in these sequences as words. Embedding is an essential step i
Time-course experiments using parallel sequencers have the potential to uncover gradual changes in cells over time that cannot be observed in a two-point comparison. An essential step in time-series data analysis is the identification of temporal differentially expressed genes (TEGs) under two conditions (e.g. control versus case). Model-based approaches, which are typical TEG detection methods, often set one parameter (e.g. degree or degree of freedom) for one dataset. This approach risks model
Abstract Motivation In recent years, single-cell RNA sequencing (scRNA-seq) has provided high-resolution snapshots of biological processes and has contributed to the understanding of cell dynamics. Trajectory inference has the potential to provide a quantitative representation of cell dynamics, and several trajectory inference algorithms have been developed. However, the downstream analysis of trajectory inference, such as the analysis of differentially expressed genes (DEG), remains challenging
Open papers in the app to read, cite, and organize with AI.