[论文解读] Who is Who in Phylogenetic Networks: Articles, Authors and Programs
本文介绍了“系统发育网络中的谁是谁”(Who is Who in Phylogenetic Networks)数据库,这是一个综合性、开放获取的数据库,整合了600多篇文献、500多位作者以及200多个关键词,涵盖系统发育网络研究领域。该数据库支持使用社会网络指标(例如,介数中心性、特征向量中心性)对合作者网络进行可视化与分析,并通过文本分析揭示研究趋势的演变,例如2005年至2015年间,研究重点从重组与水平基因转移(HGT)转向分歧分析与基于三元网络(trinet-based)的方法。
The phylogenetic network emerged in the 1990s as a new model to represent the evolution of species in the case where coexisting species transfer genetic information through hybridization, recombination, lateral gene transfer, etc. As is true for many rapidly evolving fields, there is considerable fragmentation and diversity in methodologies, standards and vocabulary in phylogenetic network research, thus creating the need for an integrated database of articles, authors, techniques, keywords and software. We describe such a database, "Who is Who in Phylogenetic Networks", available at http://phylnet.univ-mlv.fr. "Who is Who in Phylogenetic Networks" comprises more than 600 publications and 500 authors interlinked with a rich set of more than 200 keywords related to phylogenetic networks. The database is integrated with web-based tools to visualize authorship and collaboration networks and analyze these networks using common graph and social network metrics such as centrality (betweenness, eigenvector, degree and closeness) and clustering. We provide downloads of raw information about entries in the database, and a facility to suggest modifications and contribute new information to the database. We also present in this article common use cases of the database and identify trends in the research on phylogenetic networks using the information in the database and textual analysis.
研究动机与目标
- 将分散且快速演变的系统发育网络研究领域整合为一个统一、可导航的数据库。
- 为研究人员和学生提供一个易于访问的百科全书式资源,用于识别专家、方法、软件和关键文献。
- 通过社会网络指标与文献摘要的文本分析,实现对研究趋势和合作网络的分析。
- 通过提供可视化作者合作网络的工具并追踪研究主题随时间的演变,支持社区发展。
- 通过结构化的关键词标记与特定程序的元数据,促进软件与方法的发现。
提出的方法
- 该数据库整合了600多篇文献,人工校对的关键词涵盖问题类型、算法方法、输入数据、网络子类及软件可用性。
- 使用Gephi可视化合作者网络,并计算并显示社会网络指标(度数、介数中心性、特征向量中心性、接近度中心性、聚类系数、离心率)。
- 动态与预计算的网络可视化支持用户通过时间范围、文献数量阈值及中心性度量进行筛选,并可自定义颜色渐变。
- 对305篇摘要(2005–2015年)进行文本分析,采用Lexico 3进行因子分析,以识别随时间演变的研究主题与词汇变化。
- 使用DOI链接从Scopus和Web of Science获取摘要,以确保语料库的一致性并支持趋势分析。
- 该数据库在线托管,开源代码托管于GitHub,支持公众贡献与数据下载。
实验结果
研究问题
- RQ11990年至2015年间,系统发育网络研究领域的合作模式与关键贡献者如何演变?
- RQ22005年至2015年间,系统发育网络研究中的主导研究主题与方法论趋势是什么?
- RQ3介数中心性与特征向量中心性等社会网络指标在识别该领域有影响力的研究者方面有何作用?
- RQ4随着时间推移,文献摘要中的词汇与关注点发生了哪些变化,特别是关于网络类型与进化过程方面?
- RQ5合作者网络在多大程度上反映了区域研究集群,特别是欧洲的研究集群?
主要发现
- 合作者网络从1990–2000年的小型孤立组件(规模≤10)演变为2010–2015年更大、更紧密连接的结构(规模12–15),表明合作日益增强。
- 2005–2009年间,“recombination”(重组)、“hgt”(水平基因转移)、“consensus”(共识)等术语被过度使用,反映出早期研究聚焦于网络重建与基因流动。
- 2012–2015年间,“reconciliation”(分歧分析)、“trinets”(三元网络)、“duplication”(重复)、“loss”(丢失)等术语变得突出,表明研究重点转向对复杂进化事件的建模。
- 因子分析揭示出明显的二分结构:因子图右侧代表算法与数学类论文(如“networks”(网络)、“algorithm”(算法)、“rooted”(有根)),左侧则反映生物学与方法学研究(如“gene”(基因)、“species”(物种)、“inference”(推断))。
- 第二因子轴区分了显式网络重建(如“split”(分裂)、“quartet”(四元组)、“trinets”(三元网络))与分歧建模(如“duplication”(重复)、“model”(模型)、“lineage”(谱系))。
- 一篇关于公式推导的高阶数学论文在因子图右下角孤立存在,确认其在文献中的独特地位。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。