Skip to main content
QUICK REVIEW

[논문 리뷰] Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep Learning

Hamza Harkous, Kassem Fawaz|arXiv (Cornell University)|2018. 02. 07.
Privacy-Preserving Technologies in Data인용 수 174
한 줄 요약

Polisis는 구조화된 쿼리와 자유 형식 QA를 가능하게 하는 프라이버시 정책 분석 프레임워크를 도입하고, 130K 정책으로 학습된 계층적 다중레이블 CNN 분류기와 자유 형식 QA 시스템 PriBot을 통해 아이콘 및 QA 정확도를 높게 달성합니다.

ABSTRACT

Privacy policies are the primary channel through which companies inform users about their data collection and sharing practices. These policies are often long and difficult to comprehend. Short notices based on information extracted from privacy policies have been shown to be useful but face a significant scalability hurdle, given the number of policies and their evolution over time. Companies, users, researchers, and regulators still lack usable and scalable tools to cope with the breadth and depth of privacy policies. To address these hurdles, we propose an automated framework for privacy policy analysis (Polisis). It enables scalable, dynamic, and multi-dimensional queries on natural language privacy policies. At the core of Polisis is a privacy-centric language model, built with 130K privacy policies, and a novel hierarchy of neural-network classifiers that accounts for both high-level aspects and fine-grained details of privacy practices. We demonstrate Polisis' modularity and utility with two applications supporting structured and free-form querying. The structured querying application is the automated assignment of privacy icons from privacy policies. With Polisis, we can achieve an accuracy of 88.4% on this task. The second application, PriBot, is the first freeform question-answering system for privacy policies. We show that PriBot can produce a correct answer among its top-3 results for 82% of the test questions. Using an MTurk user study with 700 participants, we show that at least one of PriBot's top-3 answers is relevant to users for 89% of the test questions.

연구 동기 및 목표

  • 프라이버시 정책의 세밀한 주석 자동화를 통해 확장 가능한 다차원 쿼리를 가능하게 한다.
  • 프라이버시 전용 언어 모델과 신경망 분류기를 활용하여 세그먼트를 고수준 및 세부 프라이버시 클래스에 매핑한다.
  • 구조화된 쿼리(프라이버시 아이콘)와 자유 형식 QA(PriBot)에 대한 응용 사례를 보여준다.
  • 전문가 주석 및 사용자 연구에 대한 정확도를 평가한다.
  • 정책 개요, QA 및 프라이버시 라벨 인터페이스를 보여주는 공개적으로 접근 가능한 웹 서비스를 제공한다.

제안 방법

  • 130K 프라이버시 정책에서 subword 정보를 사용하는 fastText를 이용한 CorPus 기반 프라이버시 특화 단어 임베딩(Policies Embeddings)을 생성한다.
  • 정책 세그먼트당 10개의 고수준 범주와 122개의 세부 속성 값을 예측하도록 계층적 다중 레이블 CNN 분류기를 훈련시킨다.
  • HTML 기반 세분화 및 도메인 특화 임베딩과 함께 GraphSeg를 사용하여 정책을 의미적으로 일관된 조각으로 분할한다.
  • 카테고리 및 속성 수준 전반의 분류기를 감독하기 위해 OPP-115 데이터셋을 사용하여 정책에 주석을 달다.
  • 세그먼트와 예측 클래스에 대해 구조화된(술어 기반) 쿼리와 자유 형식(자연어) 쿼리를 가능하게 하는 애플리케이션 계층을 구현한다.
  • 실제 질문 및 MTurk 사용자를 대상으로 평가된, 사용자의 질문에 관련된 정책 세그먼트를 순위화하고 반환하는 QA 시스템 PriBot을 개발한다.

실험 결과

연구 질문

  • RQ1Polisis가 정책 세그먼트에 대해 고수준 프라이버시 범주와 세밀한 속성을 정확하게 할당할 수 있는가?
  • RQ2Polisis가 정책에 대해 구조화된 쿼리(예: 프라이버시 아이콘)를 얼마나 효과적으로 지원할 수 있는가?
  • RQ3PriBot가 프라이버시 관행에 대한 자유 형식의 사용자 질문에 대해 관련 있는 답변을 제공할 수 있는가?
  • RQ4시스템은 대규모 정책 코퍼스에서 확장 가능하고 견고한가?
  • RQ5자동 아이콘 할당은 전문가 주석 및 기존 인증 체계와 어떻게 비교되는가?

주요 결과

  • 전문가 주석과 비교했을 때 아이콘 전반에 걸친 자동 프라이버시 아이콘 할당의 평균 정확도는 88.4%를 달성했다.
  • 사용자 질문에 대한 범주 수준 쿼리는 높은 정밀도와 재현율을 보이며 매크로 평균 정밀도 0.87, 재현율 0.83, F1 0.84, 상위 1개 정밀도 0.84를 나타낸다.
  • PriBot은 테스트 질문의 상위 3개 중 최소 한 개의 정답을 82%, 상위 1개 정답으로 68%를 제공한다.
  • 700명의 MTurk 참가자를 대상으로 한 소비자 연구에서 PriBot의 상위 3개 답변이 질문의 89%에 대해 관련성이 있는 것으로 나타났다.
  • 아이콘 예측의 표 기반 평가에서 아이콘 유형별로 정확도가 다르게 나타났고(예: Automated Use 92% 정확도, Data Retention 80%, Children Privacy 98%).
  • Polisis는 전통적인 수동 라벨링과 비교하여 정책 주석 관행과 아이콘 할당의 자동화된 감사 기능을 가능하게 함으로써 확장성을 보여준다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.