[論文レビュー] Trust in AI: Interpretability is not necessary or sufficient, while black-box interaction is necessary and sufficient
本論文は、解釈可能性がAIに対する人間の信頼において必要でも十分でもないと主張している。代わりに、ブラックボックス相互作用――具体的には、自由にモデルを実行・テストできる能力――が信頼のためには必要かつ十分である。著者らは、分布外およびタスク外の性能に関する実験的・理論的証拠を統合する行動証明書フレームワークを提案し、モデルの内部構造の理解からモデル行動の理解へと焦点を移している。
The problem of human trust in artificial intelligence is one of the most fundamental problems in applied machine learning. Our processes for evaluating AI trustworthiness have substantial ramifications for ML's impact on science, health, and humanity, yet confusion surrounds foundational concepts. What does it mean to trust an AI, and how do humans assess AI trustworthiness? What are the mechanisms for building trustworthy AI? And what is the role of interpretable ML in trust? Here, we draw from statistical learning theory and sociological lenses on human-automation trust to motivate an AI-as-tool framework, which distinguishes human-AI trust from human-AI-human trust. Evaluating an AI's contractual trustworthiness involves predicting future model behavior using behavior certificates (BCs) that aggregate behavioral evidence from diverse sources including empirical out-of-distribution and out-of-task evaluation and theoretical proofs linking model architecture to behavior. We clarify the role of interpretability in trust with a ladder of model access. Interpretability (level 3) is not necessary or even sufficient for trust, while the ability to run a black-box model at-will (level 2) is necessary and sufficient. While interpretability can offer benefits for trust, it can also incur costs. We clarify ways interpretability can contribute to trust, while questioning the perceived centrality of interpretability to trust in popular discourse. How can we empower people with tools to evaluate trust? Instead of trying to understand how a model works, we argue for understanding how a model behaves. Instead of opening up black boxes, we should create more behavior certificates that are more correct, relevant, and understandable. We discuss how to build trusted and trustworthy AI responsibly.
研究の動機と目的
- AIシステムに対する人間の信頼における解釈可能性の基礎的役割を明確化すること。
- 信頼できるAIを構築するために解釈可能性が不可欠であるという広く共有された仮定に挑戦すること。
- 行動証明書によるモデル行動評価へのシフトを提案すること。
- ブラックボックス相互作用を、AIにおける信頼の評価および実現の根幹的メカニズムとして確立すること。
- 信頼最大化よりも信頼のキャリブレーションを推進し、契約意識のあるモデル設計、レジリエンステスト、信頼性の高いソフトウェア工学的手法を用いた信頼性評価を提言すること。
提案手法
- 人間-AI間の信頼と人間-AI-人間間の信頼を区別するAI-as-toolフレームワークを導入する。
- 分布外およびタスク外の評価を実験的に統合し、アーキテクチャと行動を結びつける理論的証明を含む行動証明書(BCs)を提案する。
- モデルアクセスの段階を定義する:段階2(ブラックボックス相互作用)が信頼に必要かつ十分であり、段階3(解釈可能性)はどちらでもない。
- 感度分析、敵対的テスト、ハイパーパramータ最適化などのブラックボックスアクセスのみを用いたモデルデバッグおよび科学的発見を提唱する。
- 解釈可能性に依存するアプローチを、より正確で関連性があり、理解しやすい行動証明書に置き換えるよう勧告する。
- 信頼最大化ではなく信頼のキャリブレーションを推進し、モデルカード、適合性宣言、ソフトウェア工学から取り入れたレジリエンステストの原則を活用する。
実験結果
リサーチクエスチョン
- RQ1AIシステムに対する人間の信頼に、実際に必要かつ十分なメカニズムは何か?
- RQ2解釈可能性とブラックボックス相互作用の両方が信頼を促進する上で、どのように比較されるか?
- RQ3内部モデルメカニズムの理解なしに、モデルの信頼性を評価できるか?
- RQ4行動証明書が信頼できるAIを確立するために果たす役割は何か?
- RQ5解釈可能性がなければ、科学的発見やモデルデバッグはどのように進められるか?
主な発見
- 解釈可能性は、広く信じられているようにAIに対する人間の信頼に必要でも十分でもない。
- ブラックボックス相互作用――具体的には、自由にモデルを実行・テストできる能力――が信頼に必要かつ十分である。
- 実験的および理論的証拠を統合する行動証明書は、解釈可能性よりも信頼性の高い指標である。
- 感度分析やハイパーパramータ最適化などのブラックボックスアクセスのみを用いて、モデルデバッグや科学的発見を効果的に行える。
- 解釈可能性は、ショートカット学習に傾倒するモデルでは、誤解を招く洞察をもたらすことがあるため、信頼を損なうことがある。
- 実世界の応用において、解釈可能性の手法よりも、レジリエンステスト、ホールドアウトデータ評価、モデルカードが信頼性に与える影響が大きい。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。