[论文解读] Crowdsourcing Dermatology Images with Google Search Ads: Creating a Real-World Skin Condition Dataset
本文提出了一种利用谷歌搜索广告进行可扩展众包的方法,从公众处收集真实世界中的皮肤病图像,从而创建了包含10,408张图像、来自5,033名贡献者的SCIN数据集。该数据集包含皮肤病专家验证的标签、估算的弗吉尼亚皮肤类型(eFST)以及蒙克皮肤色调(eMST),广泛涵盖了常见、短期病程的皮肤状况及多样的肤色,其中97.5%的图像为真实的临床病例。
Background: Health datasets from clinical sources do not reflect the breadth and diversity of disease in the real world, impacting research, medical education, and artificial intelligence (AI) tool development. Dermatology is a suitable area to develop and test a new and scalable method to create representative health datasets. Methods: We used Google Search advertisements to invite contributions to an open access dataset of images of dermatology conditions, demographic and symptom information. With informed contributor consent, we describe and release this dataset containing 10,408 images from 5,033 contributions from internet users in the United States over 8 months starting March 2023. The dataset includes dermatologist condition labels as well as estimated Fitzpatrick Skin Type (eFST) and Monk Skin Tone (eMST) labels for the images. Results: We received a median of 22 submissions/day (IQR 14-30). Female (66.72%) and younger (52% < age 40) contributors had a higher representation in the dataset compared to the US population, and 32.6% of contributors reported a non-White racial or ethnic identity. Over 97.5% of contributions were genuine images of skin conditions. Dermatologist confidence in assigning a differential diagnosis increased with the number of available variables, and showed a weaker correlation with image sharpness (Spearman's P values <0.001 and 0.01 respectively). Most contributions were short-duration (54% with onset < 7 days ago ) and 89% were allergic, infectious, or inflammatory conditions. eFST and eMST distributions reflected the geographical origin of the dataset. The dataset is available at github.com/google-research-datasets/scin . Conclusion: Search ads are effective at crowdsourcing images of health conditions. The SCIN dataset bridges important gaps in the availability of representative images of common skin conditions.
研究动机与目标
- 解决缺乏代表性、多样化且真实世界中的皮肤病数据集的问题,这些数据集能反映常见、短期病程且非恶性皮肤状况。
- 克服临床数据集中对深色皮肤人群和非白人人群代表性不足的偏见。
- 开发一种可扩展、低成本的方法,无需针对特定人口统计特征采样,即可收集患者自拍的皮肤病图像。
- 创建一个公开可用、开放获取的数据集,包含临床级别标签和肤色估计值,以支持公平的人工智能开发与医学教育。
- 证明利用网络搜索广告招募正在主动寻找健康信息的个体用于健康数据收集的可行性。
提出的方法
- 使用与常见皮肤症状(如皮疹、痤疮、感染)相关的关键词投放有针对性的谷歌搜索广告,以触达正在主动搜索皮肤病信息的用户。
- 引导用户通过已获同意的网络表单提交自拍图像,同时提供自报的症状、病程时长及人口统计信息。
- 对收集到的图像进行皮肤病专家确认的诊断标签、估算的弗吉尼亚皮肤类型(eFST)以及蒙克皮肤色调(eMST)的标注。
- 采用知情同意程序确保伦理数据收集与隐私保护,数据托管于GitHub以供公众访问。
- 通过统计分析评估标签质量、与图像质量的相关性,以及肤色和病症类型的分布情况。
- 将eFST与eMST的分布与已知的人口统计数据进行对比,以评估数据集的代表性。

实验结果
研究问题
- RQ1谷歌搜索广告能否在不进行人口统计目标定位的情况下,有效招募到多样化、真实世界背景的群体参与皮肤病图像贡献?
- RQ2SCIN数据集在多大程度上反映了常见皮肤状况(包括过敏性、感染性和炎症性疾病)的真实分布?
- RQ3SCIN数据集中估算的肤色(eFST与eMST)分布与美国已知的人口分布相比如何?
- RQ4在基于图像的鉴别诊断中,是否包含自报症状和图像质量信息能提升皮肤病专家的诊断信心?
- RQ5与临床来源的数据集相比,众包生成的患者自拍图像是否能支持更公平、更具泛化能力的人工智能模型训练?
主要发现
- SCIN数据集在8个月内共收集到10,408张图像,来自5,033名贡献者,日均提交量中位数为22张(四分位距14–30)。
- 66.72%的贡献者为女性,52%年龄在40岁以下,32.6%自报非白人种族或族裔身份,表明相较于美国总人口,本数据集在年轻群体和女性中的代表性更高。
- 超过97.5%的图像贡献被验证为真实的皮肤病状况,其中89%被归类为过敏性、感染性或炎症性。
- 皮肤病专家的诊断信心随可利用变量(症状、图像质量)数量的增加而提升,且与图像清晰度存在统计学上显著的弱相关性(Spearman相关系数ρ < 0.001 和 0.01,分别对应不同变量)。
- 数据集中eFST与eMST的分布与贡献者的地理来源相关,其中MST 8+(最高肤色等级)的贡献者相对较少,表明最高肤色类别存在代表性不足。
- 该数据集为美国FST与MST的分布提供了基准,支持未来在皮肤病学中关于肤色与皮肤类型的研究,包括人工智能模型的公平性评估。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。