Skip to main content
QUICK REVIEW

[论文解读] Deep Multi-Species Embedding

Di Chen, Yexiang Xue|arXiv (Cornell University)|Sep 28, 2016
Species Distribution and Climate Change参考文献 20被引用 5
一句话总结

本文提出深度多物種嵌入(DMSE),一种深度神经网络框架,可将多个鸟类物種及環境共變量聯合嵌入至共享的高維空間,以模擬物種共現與物種間相關性。透過利用深度特徵提取所學得的表示,DMSE 在預測物種分佈方面顯著優於單物種模型與多標籤基線模型,尤其在物種數量增加時表現更為突出。

ABSTRACT

Understanding how species are distributed across landscapes over time is a fundamental question in biodiversity research. Unfortunately, most species distribution models only target a single species at a time, despite strong ecological evidence that species are not independently distributed. We propose Deep Multi-Species Embedding (DMSE), which jointly embeds vectors corresponding to multiple species as well as vectors representing environmental covariates into a common high-dimensional feature space via a deep neural network. Applied to bird observational data from the citizen science project extit{eBird}, we demonstrate how the DMSE model discovers inter-species relationships to outperform single-species distribution models (random forests and SVMs) as well as competing multi-label models. Additionally, we demonstrate the benefit of using a deep neural network to extract features within the embedding and show how they improve the predictive performance of species distribution modelling. An important domain contribution of the DMSE model is the ability to discover and describe species interactions while simultaneously learning the shared habitat preferences among species. As an additional contribution, we provide a graphical embedding of hundreds of bird species in the Northeast US.

研究动机与目标

  • 解決單物種物種分佈模型忽略生態互動與共現模式的限制。
  • 發展一種可擴展的聯合建模方法,適用於數百種物種,以捕捉共享的環境偏好與物種間相關性。
  • 透過整合深度神經網絡從環境共變量中提取特徵,提升物種分佈建模的預測表現。
  • 透過學習到的向量嵌入,實現物種關係與棲息地偏好的可解釋可視化。
  • 以資料驅動、可擴展的方式量化物種互動,超越生態學者以往的定性描述。

提出的方法

  • 使用深度神經網絡將物種身分與環境共變量映射至共享的高維嵌入空間。
  • 使用 probit 鏈接函數建模物種存在/不存在,其中偵測機率取決於潛在變數是否超過零。
  • 利用嵌入向量之間的內積來捕捉語義關係,例如環境偏好相似性與物種共現相關性。
  • 使用 eBird 觀測資料,透過對數似然優化進行端到端訓練,物種間共享參數。
  • 整合投影矩陣與深度非線性轉換,從環境共變量中提取分層特徵。
  • 使用降維技術(例如 t-SNE)可視化嵌入,以揭示美國東北部地區物種共現與棲息地偏好之空間模式。

实验结果

研究问题

  • RQ1聯合嵌入模型是否能有效捕捉物種共現中的物種間相關性,同時模擬環境偏好?
  • RQ2與線性或淺層方法相比,整合深度神經網絡進行特徵提取是否能提升多物種分佈建模的預測表現?
  • RQ3建模物種相關性對預測準確率有何影響,特別是在物種數量增加時?
  • RQ4學習到的嵌入在多大程度上能揭示可解釋的生態關係與棲息地偏好?
  • RQ5該模型能否以定量方式衡量以往僅由生態學者定性描述的物種互動?

主要发现

  • DMSE 模型在物種分佈預測方面顯著優於單物種模型(隨機森林與支援向量機)與競爭性多標籤模型,且隨著物種數量增加,性能提升更為顯著。
  • 引入深度神經網絡進行特徵提取顯著提升了預測能力,如 AUC 分數較非深度基線模型顯著提高。
  • 多物種 DMSE 在所有測試物種對中均優於單物種版本的 DMSE,顯示建模物種間相關性的關鍵作用。
  • 高相關性的物種對(例如紅眼小鷚與東方鷚,r = 0.607)展現出強烈的共現模式,且模型成功捕捉與量化了這些模式。
  • 如圖 7 所示,隨著物種數量增加,模型的表現差距進一步擴大,DMSE 持續優於單物種 DMSE 與集成分類鏈模型。
  • 在美國東北部地區對數百種鳥類的圖形化嵌入,提供了物種環境偏好與共現模式的可解釋、直覺化可視化。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。