Skip to main content
QUICK REVIEW

[논문 리뷰] Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI

Alon Jacovi, Ana Marasović|arXiv (Cornell University)|2020. 10. 15.
Explainable Artificial Intelligence (XAI)참고 문헌 71인용 수 66
한 줄 요약

이 논문은 대인 관계 신뢰에서 영감을 받아 인간-인공지능 신뢰를 형식화하고, 계약적 신뢰와 신뢰성의 개념을 도입하며, 정당한 신뢰와 부당한 신뢰를 구분하고, 신뢰를 XAI 및 평가와 연결한다.

ABSTRACT

Trust is a central component of the interaction between people and AI, in that 'incorrect' levels of trust may cause misuse, abuse or disuse of the technology. But what, precisely, is the nature of trust in AI? What are the prerequisites and goals of the cognitive mechanism of trust, and how can we promote them, or assess whether they are being satisfied in a given interaction? This work aims to answer these questions. We discuss a model of trust inspired by, but not identical to, sociology's interpersonal trust (i.e., trust between people). This model rests on two key properties of the vulnerability of the user and the ability to anticipate the impact of the AI model's decisions. We incorporate a formalization of 'contractual trust', such that trust between a user and an AI is trust that some implicit or explicit contract will hold, and a formalization of 'trustworthiness' (which detaches from the notion of trustworthiness in sociology), and with it concepts of 'warranted' and 'unwarranted' trust. We then present the possible causes of warranted trust as intrinsic reasoning and extrinsic behavior, and discuss how to design trustworthy AI, how to evaluate whether trust has manifested, and whether it is warranted. Finally, we elucidate the connection between trust and XAI using our formalization.

연구 동기 및 목표

  • 사회학적 대인 관계 신뢰에서 영감을 받은 인간과 AI 모델 간의 신뢰를 정의한다.
  • 계약적 신뢰를 도입하고 신뢰성의 차이를 구분한다.
  • 정당한 신뢰와 부당한 신뢰를 구분하고 그 시사점을 논의한다.
  • 신뢰가 XAI 및 모델 평가와 어떻게 연결되는지 설명한다.
  • 현실 세계 상호작용에서 신뢰할 수 있는 AI를 설계하고 평가하기 위한 프레임워크를 제안한다.

제안 방법

  • Adopt two core properties of trust: 사용자 취약성과 AI의 영향 예측 능력을 채택한다.
  • Formalize contractual trust to specify what the AI is trusted to do.
  • Differentiate trustworthiness from trust and define warranted vs unwarranted trust.
  • Characterize intrinsic trust (aligned internal reasoning) and extrinsic trust (credible external behavior) as trust-promoting mechanisms.
  • Relate the framework to European guidelines and standard documentations (e.g., datasheets, model cards) to specify contracts.
  • Discuss evaluation methodologies (proxy explanations, post-deployment data, test sets) to justify extrinsic trust.

실험 결과

연구 질문

  • RQ1What are the prerequisites for Human-AI trust?
  • RQ2How can contractual trust and trustworthiness be formally defined for Human-AI interactions?
  • RQ3What differentiates warranted from unwarranted trust in AI?
  • RQ4How can trust be promoted and evaluated in practice, including the role of XAI?
  • RQ5How does the proposed formalization connect to existing guidelines and documentation practices?

주요 결과

  • Trust between a human and AI is a directional transaction requiring vulnerability and anticipation of impact.
  • Contractual trust and trustworthiness can be formalized to distinguish when trust is warranted versus unwarranted.
  • Trust can be context-dependent, with contracts conditioned on interaction context.
  • Trust mechanisms include intrinsic trust (explainable reasoning aligned with user priors) and extrinsic trust (trust in evaluation methods and data).
  • Evaluation schemes (proxy judgments, post-deployment data, and test sets) are essential to establishing extrinsic trust and to assess whether a contract can be maintained.
  • Explainability in AI should align with user priors and contract-specific goals to foster genuine trust rather than mere perception.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.