[论文解读] Harvesting Collective Intelligence: Temporal Behavior in Yahoo Answers
本文研究了Yahoo Answers用户在获取信息时如何权衡速度与准确性,发现其在答案数量上表现出递减的边际收益,且每个问题的答案数量服从逆高斯分布——表明存在基于阈值的决策过程。该研究结合真实世界数据的实证分析与理论建模,揭示了集体智能获取中的行为模式,对优化问答平台设计具有启示意义。
When harvesting collective intelligence, a user wishes to maximize the accuracy and value of the acquired information without spending too much time collecting it. We empirically study how people behave when facing these conflicting objectives using data from Yahoo Answers, a community driven question-and-answer site. We take two complementary approaches. We first study how users behave when trying to maximize the amount of the acquired information, while minimizing the waiting time. We identify and quantify how question authors at Yahoo Answers trade off the number of answers they receive and the cost of waiting. We find that users are willing to wait more to obtain an additional answer when they have only received a small number of answers; this implies decreasing marginal returns in the amount of collected information. We also estimate the user's utility function from the data. Our second approach focuses on how users assess the qualities of the individual answers without explicitly considering the cost of waiting. We assume that users make a sequence of decisions, deciding to wait for an additional answer as long as the quality of the current answer exceeds some threshold. Under this model, the probability distribution for the number of answers that a question gets is an inverse Gaussian, which is a Zipf-like distribution. We use the data to validate this conclusion.
研究动机与目标
- 理解用户在社区驱动的问答平台中如何平衡信息获取速度与答案准确性。
- 建模问题发布者在收到足够数量和质量的答案后决定何时关闭问题的停止行为。
- 检验每个问题的答案数量是否符合齐普夫(Zipf)类分布,以验证是否存在基于阈值的决策过程。
- 从答案数量和等待时间的角度估计用户效用函数,揭示递减的边际收益。
- 通过识别最优问题置顶策略以最大化社会总 surplus,为平台设计提供建议。
提出的方法
- 基于Yahoo Answers数据进行实证分析,建模答案数量与等待时间之间的权衡,假设用户在时间与信息约束下最大化效用。
- 从数据中估计一个凹效用函数,显示答案数量的边际收益递减。
- 将用户决策建模为基于阈值的停止规则:只要当前答案的价值超过阈值,用户就继续等待。
- 推导出在阈值模型下,答案数量服从逆高斯分布,其尾部行为具有齐普夫(Zipf)类特征。
- 使用最大似然估计将逆高斯分布拟合到数据,参数为 μ = 6.1 和 λ = 5.8。
- 通过经验累积分布函数和频率分布的双对数图验证模型。
实验结果
研究问题
- RQ1用户在Yahoo Answers中如何权衡所获答案数量与等待时间?
- RQ2用户效用函数在答案数量与等待成本方面的形状如何?
- RQ3每个问题的答案数量是否符合基于阈值的决策模型所预测的分布?
- RQ4逆高斯分布是否能很好地拟合答案数量的经验分布?
- RQ5所观察到的行为是否可用具有吸收屏障的阈值的随机游走模型来解释?
主要发现
- 用户在答案数量上表现出递减的边际收益,表明其效用函数关于答案数量呈凹性。
- 从数据中估计出的效用函数显示,用户为前几个答案愿意等待更长时间,而对后续答案的等待意愿显著下降。
- 每个问题的答案数量服从逆高斯分布,均值 μ = 6.1,尺度参数 λ = 5.8,经经验累积分布函数和双对数频率图验证。
- 在双对数尺度下,分布的尾部斜率约为 -3/2,证实了逆高斯模型所预测的齐普夫(Zipf)类行为。
- 基于阈值的决策模型(即用户持续等待,直到当前答案价值超过某一阈值)能够解释观察到的答案数量分布。
- 结果表明,问答平台可通过利用估计的效用函数和答案分布模型,优化问题的可见性与回答速率。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。