Skip to main content
QUICK REVIEW

[论文解读] Set-based complexity and biological information

David J. Galas, Matti Nykter|ArXiv.org|Jan 25, 2008
Computability, Logic, AI Algorithms参考文献 39被引用 8
一句话总结

本文提出了一种基于集合的复杂性度量方法,其基础为柯尔莫哥洛夫复杂性,通过剔除随机性和冗余性来量化具有生物学意义的信息。该方法以通用信息距离为基础,构建了一种天然捕捉上下文信息的度量,无需预设状态空间,其最大化可揭示基因组和调控序列中的生物学相关复杂性模式。

ABSTRACT

It is not obvious what fraction of all the potential information residing in the molecules and structures of living systems is significant or meaningful to the system. Sets of random sequences or identically repeated sequences, for example, would be expected to contribute little or no useful information to a cell. This issue of quantitation of information is important since the ebb and flow of biologically significant information is essential to our quantitative understanding of biological function and evolution. Motivated specifically by these problems of biological information, we propose here a class of measures to quantify the contextual nature of the information in sets of objects, based on Kolmogorov's intrinsic complexity. Such measures discount both random and redundant information and are inherent in that they do not require a defined state space to quantify the information. The maximization of this new measure, which can be formulated in terms of the universal information distance, appears to have several useful and interesting properties, some of which we illustrate with examples.

研究动机与目标

  • 为解决在分子系统中量化具有生物学意义的信息的挑战,其中随机或重复序列的功能价值极低。
  • 开发一种复杂性度量方法,天然地考虑上下文信息,而无需依赖外部预定义的状态空间。
  • 形式化一种度量方法,以区分序列或结构集合中的有意义生物信息与噪声或冗余信息。
  • 探讨该度量在生物学功能和进化背景下的特性,特别是在基因组学和调控网络中的应用。

提出的方法

  • 将柯尔莫哥洛夫的内在复杂性适应于对象集合,聚焦于整个集合的算法复杂性,而非单个元素。
  • 以通用信息距离为基础,定义一种捕捉结构和上下文关系的基于集合的复杂性度量。
  • 设计一种度量方法,天然地剔除随机序列和相同重复序列,因为它们不提供有意义的信息。
  • 将该度量应用于基因组序列和调控基序等生物数据集,以实现上下文感知的复杂性评估。
  • 通过最大化复杂性度量,识别出信息含量高且具有功能相关性的配置。
  • 采用信息论原则,确保该度量在某些变换下保持不变,并对噪声具有鲁棒性。

实验结果

研究问题

  • RQ1在分子系统中,如何区分具有生物学意义的信息与随机或冗余序列?
  • RQ2对于不依赖预定义状态空间的生物对象集合,其复杂性度量的合适形式是什么,以体现上下文关系?
  • RQ3所提出的基于集合的复杂性度量与已知的生物学功能和进化动态有何关联?
  • RQ4从通用信息距离导出的复杂性度量在生物学背景下的数学和计算特性是什么?
  • RQ5该复杂性度量的最大化能否揭示基因组和调控序列中的生物学显著模式?

主要发现

  • 所提出的基于集合的复杂性度量能有效过滤随机和冗余序列,仅聚焦于结构和上下文上有意义的信息。
  • 该度量天然定义,无需预先指定状态空间,因此适用于多种类型的生物数据。
  • 复杂性度量的最大化揭示了具有高度功能相关性的配置,暗示其与生物学调控和进化存在关联。
  • 该度量在存在噪声的情况下,仍能稳健识别基因组序列和调控网络中的有意义模式。
  • 理论分析表明,该度量与算法信息论的已知原理一致,并为量化生物信息提供了原则性框架。
  • 实证示例表明,该度量在捕捉生物系统功能复杂性方面,优于传统的熵或信息度量。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。