[論文レビュー] A generative, predictive model for menstrual cycle lengths that accounts for potential self-tracking artifacts in mobile health data
本論文は、モバイルヘルス(mHealth)データにおけるユーザーが報告するトラッキングアーティファクト(特にサイクルスキップ)を明示的に扱う階層的で生成的確率モデルを提案する。サイクル長とスキップ行動を同時にモデル化することで、186,000人のユーザーから得られた200万件以上のサイクルデータにおいて、ニューラルネットワークや要約統計量ベースラインを上回る最先端の予測精度を達成するとともに、サイクルの進行に伴いオンラインで解釈可能な更新が可能となる。
Mobile health (mHealth) apps such as menstrual trackers provide a rich source of self-tracked health observations that can be leveraged for health-relevant research. However, such data streams have questionable reliability since they hinge on user adherence to the app. Therefore, it is crucial for researchers to separate true behavior from self-tracking artifacts. By taking a machine learning approach to modeling self-tracked cycle lengths, we can both make more informed predictions and learn the underlying structure of the observed data. In this work, we propose and evaluate a hierarchical, generative model for predicting next cycle length based on previously-tracked cycle lengths that accounts explicitly for the possibility of users skipping tracking their period. Our model offers several advantages: 1) accounting explicitly for self-tracking artifacts yields better prediction accuracy as likelihood of skipping increases; 2) because it is a generative model, predictions can be updated online as a given cycle evolves, and we can gain interpretable insight into how these predictions change over time; and 3) its hierarchical nature enables modeling of an individual's cycle length history while incorporating population-level information. Our experiments using mHealth cycle length data encompassing over 186,000 menstruators with over 2 million natural menstrual cycles show that our method yields state-of-the-art performance against neural network-based and summary statistic-based baselines, while providing insights on disentangling menstrual patterns from self-tracking artifacts. This work can benefit users, mHealth app developers, and researchers in better understanding cycle patterns and user adherence.
研究の動機と目的
- 不規則なユーザーのトラッキング、特にサイクルスキップに起因する信頼性の低いmHealthデータの課題に対処すること。
- モバイルヘルスデータにおける自己トラッキングアーティファクトと真の月経周期パターンを区別できる予測モデルを開発すること。
- サイクルの進行に伴い、現在の日付の情報を組み込みながら、リアルタイムでオンライン予測更新を可能にすること。
- 階層ベイジアンフレームワークを用いて、個々のサイクル履歴と集団レベルの事前分布を統合すること。
- 予測精度の向上と、ユーザー行動および周期の規則性に関する解釈可能な洞察の提供
提案手法
- モデルは、個人の履歴的トラッキングデータに基づき、サイクル長とスキップ確率を同時に推定する階層ベイジアン生成フレームワークを用いる。
- 現在のサイクル内の日付と学習済みハイパーパrameterを条件として、次回のサイクル長の事後予測分布を計算する。
- 予測に不確実性を組み込むために、$ p(d^{*} | \hat{u}, d_i, d^{*} > d_{\text{current}}) $ を用いて、可能性のあるスキップ状態の積分を行う。
- 現在のサイクルの日付に条件づけて予測を行うことで、動的かつリアルタイムに予測分布を更新できる。
- サイクル長とスキップ行動をそれぞれ正規分布とベルヌーイ分布の混合で表現する。
- ハイパーパrameterは集団レベルのデータから学習され、トランスファーラーニングを可能にし、個々のユーザーにわたる一般化性能を向上させる。
実験結果
リサーチクエスチョン
- RQ1ユーザーのサイクルスキップを考慮することで、mHealthデータにおける月経周期長の予測精度はどの程度向上するか?
- RQ2生成的モデルは、現実のmHealthサイクルデータにおいて、ニューラルネットワークや要約統計量ベースラインを上回る予測性能を示せるか?
- RQ3オンラインで日次に更新される予測は、予測の解釈可能性と信頼性をどの程度向上させるか?
- RQ4モデルの階層的構造は、限られた個人データでも効果的なパーソナライズを可能にし、信頼性を維持するのにどの程度寄与するか?
- RQ5周期長の予測精度と、ユーザーの追跡アーティファクト(例:スキップ)の確率との間にどのような関係があるか?
主な発見
- 提案されたモデルは、同じデータセット上でニューラルネットワークベースのモデル(例:CNN、RNN、LSTM)および要約統計量ベースライン(平均、中央値)を著しく上回る最先端の予測精度を達成した。
- サイクルスキップの考慮により、特にスキップ率の高いユーザーにおいても予測精度が向上し、自己トラッキングアーティファクトに対してモデルのロバストネスが確認された。
- モデルは、サイクルの進行に伴い進化する動的で解釈可能な予測を提供するリアルタイムのオンライン予測更新を可能にした。
- 階層的構造により、個人の履歴データを活用した効果的なパーソナライズが可能であり、同時に集団レベルの事前分布を活用することで、データが少ないユーザーの性能も向上した。
- 予測精度の向上と多様なユーザー群における一貫性あるパフォーマンスから、モデルが真のサイクル長パターンとトラッキングアーティファクトを効果的に分離できていることが裏付けられた。
- RMSE、MAE、中央周期長差(median CLD)の評価により、モデルが高い精度を維持し、ベースラインよりも周期の規則性をより効果的に捉えていることが確認された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。