[论文解读] Machine Learning Analysis of Complex Networks in Hyperspherical Space
本文提出了一种新颖的几何机器学习方法,通过基于信息流测量的连通性函数,将复杂网络嵌入超球面空间。通过使用非度量多维尺度分析(nonmetric MDS)和K均值聚类分析节点位置向量之间的夹角,揭示了网络中的结构社区——在引文网络和基因互作网络上进行了验证,识别出与多种疾病和癌症相关的生物学显著聚类,结果与近期生物学文献一致。
A complex network is a condensed representation of the relational topological framework of a complex system. A main reason for the existence of such networks is the transmission of items through the entities of these complex systems. Here, we consider a communicability function that accounts for the routes through which items flow on networks. Such a function induces a natural embedding of a network in a Euclidean high-dimensional sphere. We use one of the geometric parameters of this embedding, namely the angle between the position vectors of the nodes in the hyperspheres, to extract structural information from networks. Such information is extracted by using machine learning techniques, such as nonmetric multidimensional scaling and K-means clustering algorithms. The first allows us to reduce the dimensionality of the communicability hyperspheres to 3-dimensional ones that allow network visualization. The second permits to cluster the nodes of the networks based on their similarities in terms of their capacity to successfully deliver information through the network. After testing these approaches in benchmark networks and compare them with the most used clustering methods in networks we analyze two real-world examples. In the first, consisting of a citation network, we discover citation groups that reflect the level of mathematics used in their publications. In the second, we discover groups of genes that coparticipate in human diseases, reporting a few genes that coparticipate in cancer and other diseases. Both examples emphasize the potential of the current methodology for the discovery of new patterns in relational data.
研究动机与目标
- 开发一种基于自然信息流的新型几何框架,用于分析复杂网络,通过将网络嵌入超球面实现。
- 克服传统聚类方法仅依赖边密度的局限性,引入基于信息流的相似性度量。
- 利用超球面嵌入的几何特性,实现网络的可视化与无监督聚类。
- 在真实世界网络中发现具有生物学意义的社区,如引文群体和与疾病相关的基因聚类。
- 通过与已知的疾病-基因关联及近期文献对比,验证所发现聚类的生物学相关性。
提出的方法
- 使用连通性函数将网络嵌入(n−1)维欧几里得超球面,该函数基于网络中所有路径量化节点间的有效连通性。
- 计算超球面中节点位置向量之间的夹角,作为反映其信息传输能力的几何相似性度量。
- 应用非度量多维尺度分析(NMDS)将超球面嵌入降维至3D以实现可视化,同时保持相似性的等级顺序。
- 在角度相似性矩阵上应用K均值聚类,根据节点在信息传输中的拓扑角色对节点进行分组。
- 通过与已知的生物注释及基因-疾病关联的近期文献对比,验证聚类结果。
- 利用连通性嵌入的自然几何结构,避免其他几何学习方法中常见的任意或人为设定的嵌入。
实验结果
研究问题
- RQ1基于网络中信息流的几何嵌入是否能揭示超越边密度定义的有意义的结构社区?
- RQ2超球面空间中节点向量夹角作为网络聚类的相似性度量有多有效?
- RQ3该方法能否在人类疾病网络中发现标准聚类方法无法检测到的生物相关基因聚类?
- RQ4基于连通性相似性的聚类中,被分在同一簇的基因是否表现出在多种疾病(包括癌症)中的共同参与?
- RQ5所提出的方法能否揭示文献中尚未报道的新疾病-基因关联?
主要发现
- 该方法成功利用非度量多维尺度分析在3D中可视化复杂网络,同时保留了超球面嵌入中节点的相对结构关系。
- 在引文网络中,算法识别出反映论文中数学复杂程度的显著聚类,验证了其检测主题分组的能力。
- 在人类基因互作网络中,聚类1中37%的基因被发现与癌症相关,尽管此前未被归类为相关基因,表明其功能相关性。
- 聚类5包含107个基因,其中26个属于“灰色”类别(涉及多种疾病),且其中14个在近期研究中被新发现与癌症相关,支持了在疾病网络中存在共享拓扑角色的假设。
- 聚类5中14个此前未报告的“灰色”基因中有10个在近期文献中被证实与多种癌症相关,例如ABCA1与前列腺癌相关,ESR1与激素抵抗性乳腺癌相关。
- 结果表明,基于连通性和超球面嵌入的几何聚类方法可在无需先验标签的情况下,发现具有生物学显著性的高质量网络社区。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。