Skip to main content
QUICK REVIEW

[論文レビュー] Estimating the number of species to attain sufficient representation in a random sample

Chao Deng, Timothy Daley|arXiv (Cornell University)|Jul 11, 2016
Census and Population Estimation参考文献 13被引用数 6
ひとこと要約

本稿では、初期のサンプル頻度に基づいて、将来のサンプルサイズ t において少なくとも r 回現れる種の期待数を予測する非パrametric推定量 Ψ_{r,m}(t) を提案する。r 階微分の種の発見率の有理関数近似を用いることで、大規模な r および高次元データセットにおいても、正確で安定した長距離外挿が可能となり、シミュレーションおよびゲノムやソーシャルネットワークを含む実世界の応用において、既存手法を上回る性能を発揮する。

ABSTRACT

The statistical problem of using an initial sample to estimate the number of species in a larger sample has found important applications in fields far removed from ecology. Here we address the general problem of estimating the number of species that will be represented by at least a number r of observations in a future sample. The number r indicates species with sufficient observations, which are commonly used as a necessary condition for any robust statistical inference. We derive a procedure to construct consistent estimators that apply universally for a given population: once constructed, they can be evaluated as a simple function of r. Our approach is based on a relation between the number of species represented at least r times and the higher derivatives of the expected number of species discovered per unit of time. Combining this relation with a rational function approximation, we propose nonparametric estimators that are accurate for both large values of r and long-range extrapolations. We further show that our estimators retain asymptotic behaviors that are essential for applications on large-scale datasets. We evaluate the performance of this approach by both simulation and real data applications for inferences of the vocabulary of Shakespeare and Dickens, the topology of a Twitter social network, and molecular diversity in DNA sequencing data.

研究の動機と目的

  • 将来のサンプルにおいて少なくとも r 回観測される種の数を推定する統計的課題に取り組むこと。r > 1 であると、頑健な推論に十分な代表度が得られることを意味する。
  • 種の豊度分布のパラメトリックな仮定を必要とせず、ある母集団に対して異なる r 値に普遍的に適用可能な非パラメトリック推定量を開発すること。
  • 特に大規模な r 値および大規模データセット(例:DNAシークエンシングやソーシャルメディアデータ)に対して、r=1 を超える種の蓄積曲線の長距離外挿の正確性を向上させること。
  • 高次モーメントを活用することで推定の信頼性を高める、安定的かつ理論的根拠を持つ方法を提供すること。

提案手法

  • 本手法は、r 種の蓄積曲線 E[S_r(t)] と平均発見率 E[S_1(t)]/t の (r−1) 階微分との間の理論的関係を導出する。
  • 種の頻度分布を潜在的強度分布 G(λ) を持つポアソン過程の混合でモデル化し、G(λ) の直接推定を避ける。
  • 発見率の微分を近似するための有理関数近似(RFA)を用い、形式 P_{m-1}(t)/Q_m(t) を採用することで、安定的かつ正確な推定を可能にする。
  • 推定量 Ψ_{r,m}(t) は、初期サンプルの頻度度数 N_j の関数として構築され、特にパデ近似を用いて高次モーメントを活用する。
  • 個々の E[N_j(t)] 推定値の和を取るのを避け、微分に基づく関係を通じて直接的に E[S_r(t)] をモデル化する。
  • シミュレーションおよび実データ応用(シェイクスピアの語彙、Twitterユーザーの活動、DNAシークエンシングデータなど)を通じて、手法の妥当性を検証する。

実験結果

リサーチクエスチョン

  • RQ1大規模な r 値および長距離外挿において、将来のサンプルで少なくとも r 回現れる種の数を正確に推定する方法は何か?
  • RQ2種の豊度に特定のパラメトリックな形を仮定しない非パラメトリック推定量を構築できるか? また、その推定量は異なる r 値においても一貫性と安定性を保つのか?
  • RQ3種の発見率の高次微分を正確に近似する最適な方法は何か? これにより、種の数の予測精度が向上する。
  • RQ4提案された推定量 Ψ_{r,m}(t) は、ZTNB推定量など既存手法と比較して、多様なデータセットにおいて精度と頑健性に優れているか?

主な発見

  • DNAシークエンシングデータにおいて、初期サンプルサイズの100倍まで外挿する場合、推定量 Ψ_{r,m}(t) は相対誤差が5%未満に抑えられる。r > 1 であっても同様に高い精度を示す。
  • 同じDNAシークエンシングデータにおいて、Ψ_{r,m}(t) は複数の r 値において真の期待値をよく追跡するが、ZTNB推定量は E[S_1(t)] を過剰推定し、r > 1 では逆に過小推定する。
  • シェイクスピアの語彙、ディケンズの作品、Twitterソーシャルネットワーク、ゲノムシークエンシングなど、多様なデータセットにおいても、本手法は高い正確性を維持し、ZTNBをほぼすべてのケースで上回る。
  • 推定量は良好な漸近的挙動を示し、大規模応用において安定性と一貫性を保証する。
  • パデ近似を用いた有理関数近似(RFA)により、発見率の高次微分の正確なモデル化が可能となり、長距離予測にとって不可欠である。
  • 本手法はパラメトリックな仮定を必要とせず、種の頻度度数の高次モーメントを効果的に活用でき、分野を問わず広く適用可能である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。