Skip to main content
QUICK REVIEW

[论文解读] Quantifying the Impact of User Attention on Fair Group Representation in Ranked Lists

Piotr Sapieżyński, Wesley Zeng|arXiv (Cornell University)|Jan 29, 2019
Electoral Systems and Political Participation被引用 8
一句话总结

本文提出了 Viable-Λ 檢測,這是一種新型度量指標,用於透過參數化分佈建模使用者注意力,而非假設固定折舊(例如對數或幾何分佈),來審計排序清單中的群體公平性。研究顯示,同一份排序清單在不同注意力模型下可能表現為公平或偏頗,強調公平性評估必須考量使用者行為與服務情境,以避免對演算法偏見得出誤導性結論。

ABSTRACT

In this work, we introduce a novel metric for auditing group fairness in ranked lists. Our approach offers two benefits compared to the state of the art. First, we offer a blueprint for modeling of user attention. Rather than assuming a logarithmic loss in importance as a function of the rank, we can account for varying user behaviors through parametrization. For example, we expect a user to see more items during a viewing of a social media feed than when they inspect the results list of a single web search query. Second, we allow non-binary protected attributes to enable investigating inherently continuous attributes (\\eg political alignment on the liberal to conservative spectrum) as well as to facilitate measurements across aggregated sets of search results, rather than separately for each result list. By combining these two elements into our metric, we are able to better address the human factors inherent in this problem. We measure the whole sociotechnical system, consisting of a ranking algorithm and individuals using it, instead of exclusively focusing on the ranking algorithm. Finally, we use our metric to perform three simulated fairness audits. We show that determining fairness of a ranked output necessitates knowledge (or a model) of the end-users of the particular service. Depending on their attention distribution function, a fixed ranking of results can appear biased both in favor and against a protected group.

研究动机与目标

  • 彌補公平排序研究中的缺口,透過比固定折舊函數更真實的方式建模使用者注意力。
  • 實現對排序清單的公平性審計,並考量不同服務(例如社群媒體與搜尋引擎)中使用者行為的差異。
  • 在公平性評估中支援非二元與連續受保護屬性(例如政治立場),超越二元群體分類。
  • 將關注焦點從排序演算法本身,轉向排序 + 使用者行為的社會技術系統來進行公平性評估。
  • 示範公平性結論對假設的注意力分佈極其敏感,因此必須進行情境感知的審計。

提出的方法

  • 提出 Viable-Λ 檢測,一種度量指標,用於判斷是否存在某種使用者注意力分佈 P(Λ),使得排序清單達成群體公平。
  • 使用參數化注意力模型(例如截斷幾何分佈)來表示不同使用者行為,其中參數 λ₁,…,λₘ 控制注意力衰減。
  • 提出由觀察到的前 k 筆結果推導出的群體估計器 p̂,以估計完整結果集中受保護群體成員的真實比例。
  • 將該度量指標應用於使用真實搜尋結果(例如「obamacare continue」、「medicare reform」)進行的模擬審計,涵蓋不同注意力模型。
  • 採用 SEME(社會技術系統模型公平性評估)框架,量化在不同注意力假設下的公平性。
  • 使用該度量指標評估某排序是否在任何合理的注意力模型下皆可視為公平,而非假設單一固定模型。

实验结果

研究问题

  • RQ1在任何合理的使用者注意力分佈下,排序清單是否可被視為群體公平?還是公平性高度依賴於所假設的注意力模型?
  • RQ2注意力分佈的形狀(例如陡峭 vs. 平坦)如何影響已知結果之排序清單的公平性感知?
  • RQ3連續或非二元屬性(例如政治立場)在多大程度上可被有意義地納入排序清單的公平性度量中?
  • RQ4注意力模型的選擇如何影響真實搜尋結果中偏見的可偵測性?
  • RQ5若注意力模型未能反映特定服務情境下的實際使用者行為,公平性結論是否可能具有誤導性?

主要发现

  • 同一份排序清單在不同假設的注意力分佈下,可能對受保護群體表現為偏頗或公平,顯示公平性並非絕對,而是取決於情境。
  • 對於「obamacare continue」,在陡峭注意力模型下(偏好頂端結果),清單顯得偏向自由派;但在較平坦的注意力分佈下,則可能顯得更為平衡。
  • 對於「medicare reform」,頂端結果若偏向保守派,則在陡峭注意力模型下會使清單顯得偏向保守派,即使多數結果實際上偏向自由派。
  • Viable-Λ 檢測成功識別出某排序是否在任何合理的注意力模型下皆可達成公平,揭示部分排序本質上即不公平,無視注意力形狀。
  • 研究顯示,假設固定注意力模型(例如對數或幾何模型)可能導致錯誤的公平性評估,特別是在使用者行為因介面或任務而異時。
  • 該度量指標顯示,公平性審計必須考量服務特定的使用者行為,因為相同排序在不同注意力模型下可能被視為公平或不公平。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。