Skip to main content
QUICK REVIEW

[论文解读] Integrating Random Forests and Generalized Linear Models for Improved Accuracy and Interpretability

Abhineet Agarwal, Ana Kenney|arXiv (Cornell University)|Jul 4, 2023
Gene expression and cancer classification被引用 8
一句话总结

所提供文本是一个带占位内容的模板;未描述具体的方法、结果或研究问题。

ABSTRACT

Random forests (RFs) are among the most popular supervised learning algorithms due to their nonlinear flexibility and ease-of-use. However, as black box models, they can only be interpreted via algorithmically-defined feature importance methods, such as Mean Decrease in Impurity (MDI), which have been observed to be highly unstable and have ambiguous scientific meaning. Furthermore, they can perform poorly in the presence of smooth or additive structure. To address this, we reinterpret decision trees and MDI as linear regression and $R^2$ values, respectively, with respect to engineered features associated with the tree's decision splits. This allows us to combine the respective strengths of RFs and generalized linear models in a framework called RF+, which also yields an improved feature importance method we call MDI+. Through extensive data-inspired simulations and real-world datasets, we show that RF+ improves prediction accuracy over RFs and that MDI+ outperforms popular feature importance measures in identifying signal features, often yielding more than a 10% improvement over its closest competitor. In case studies on drug response prediction and breast cancer subtyping, we further show that MDI+ extracts well-established genes with significantly greater stability compared to existing feature importance measures.

研究动机与目标

  • 该文档似乎是一个格式化/模板样本,而非实际研究目标的报告。
  • 在提供的文本中没有描述具体的动机或科学目标。
  • 摘要、方法和验证部分是占位符,内容不真实。

提出的方法

  • 文本中没有提供实际的方法学内容;各节作为占位符标注。
  • 本文档演示排版指南(如边距、图形、表格等),而非提出新的方法学贡献。
  • 在所提供的材料中未描述关键技术或方程。
Figure 1: Consistency comparison in fitting surrogate model in the tidal power example.
Figure 1: Consistency comparison in fitting surrogate model in the tidal power example.

实验结果

研究问题

  • RQ1在所提供的文本中没有明确的研究问题。

主要发现

  • 在提供的内容中未报告任何经验结果或定量发现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。