Skip to main content
QUICK REVIEW

[论文解读] Floating Forests: Quantitative Validation of Citizen Science Data Generated From Consensus Classifications

Isaac S. Rosenthal, Jarrett E. K. Byrnes|arXiv (Cornell University)|Jan 25, 2018
Mobile Crowdsensing and Crowdsourcing参考文献 9被引用 13
一句话总结

本研究验证了浮游森林项目中公民科学数据的可靠性,表明非专家通过共识分类(每幅图像4.2名用户)可实现高精度(Landsat 5/7的MCC = 0.400,Landsat 8的MCC = 0.639),与专家分类结果相当。该方法利用聚合的用户勾绘结果生成可靠的海带森林地图,支持通过可扩展、低成本的数据采集实现大规模生态监测。

ABSTRACT

Large-scale research endeavors can be hindered by logistical constraints limiting the amount of available data. For example, global ecological questions require a global dataset, and traditional sampling protocols are often too inefficient for a small research team to collect an adequate amount of data. Citizen science offers an alternative by crowdsourcing data collection. Despite growing popularity, the community has been slow to embrace it largely due to concerns about quality of data collected by citizen scientists. Using the citizen science project Floating Forests (http://floatingforests.org), we show that consensus classifications made by citizen scientists produce data that is of comparable quality to expert generated classifications. Floating Forests is a web-based project in which citizen scientists view satellite photographs of coastlines and trace the borders of kelp patches. Since launch in 2014, over 7,000 citizen scientists have classified over 750,000 images of kelp forests largely in California and Tasmania. Images are classified by 15 users. We generated consensus classifications by overlaying all citizen classifications and assessed accuracy by comparing to expert classifications. Matthews correlation coefficient (MCC) was calculated for each threshold (1-15), and the threshold with the highest MCC was considered optimal. We showed that optimal user threshold was 4.2 with an MCC of 0.400 (0.023 SE) for Landsats 5 and 7, and a MCC of 0.639 (0.246 SE) for Landsat 8. These results suggest that citizen science data derived from consensus classifications are of comparable accuracy to expert classifications. Citizen science projects should implement methods such as consensus classification in conjunction with a quantitative comparison to expert generated classifications to avoid concerns about data quality.

研究动机与目标

  • 评估浮游森林项目中通过共识分类生成的公民科学数据的准确性。
  • 确定能最大化分类准确性的用户分类数量(每幅图像)的最佳值。
  • 将共识生成的数据与专家生成的分类结果进行比较,以验证数据质量。
  • 为在大规模生态监测中使用公民科学提供实证支持。

提出的方法

  • 公民科学家自2014年起对75万幅海带森林卫星图像进行了分类,勾绘海带斑块边界。
  • 每幅图像由15名独立用户进行分类,通过叠加所有勾绘结果生成共识分类。
  • 针对每个用户阈值(1至15)计算 Matthews 相关系数(MCC),以评估分类准确性。
  • 最优阈值被确定为在准确性和数据效率之间达到最佳平衡时的值。
  • 使用专家分类作为真实值,对共识输出进行定量验证。
  • 统计分析比较了Landsat 5/7与Landsat 8影像的MCC值,以评估传感器特异性表现。

实验结果

研究问题

  • RQ1在基于共识的公民科学数据中,每幅图像的最优用户分类数量是多少,才能实现最大准确性?
  • RQ2共识分类在海带森林制图中的准确性与专家生成的分类结果相比如何?
  • RQ3共识数据的准确性是否因不同卫星传感器(Landsat 5/7 与 Landsat 8)而异?
  • RQ4非专家志愿者的共识分类能否生成可靠的数据,用于大规模生态监测?
  • RQ5哪种统计指标最能量化遥感应用中共识分类的可靠性?

主要发现

  • 共识分类的最优用户阈值为4.2,对应Landsat 5和7影像的Matthews相关系数(MCC)为0.400(±0.023标准误)。
  • 对于Landsat 8影像,最优阈值产生的MCC为0.639(±0.246标准误),表明准确率显著更高。
  • 共识分类达到了与专家生成分类相当的准确度,验证了其在生态学研究中的适用性。
  • 该方法在不同卫星传感器上均表现出稳健性能,Landsat 8的准确度高于Landsat 5/7。
  • 本研究证实,共识分类是一种可靠且可扩展的方法,适用于处理大规模遥感数据。
  • 使用MCC进行的定量验证提供了一个可重复的框架,用于评估公民科学项目中的数据质量。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。