[논문 리뷰] ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
ProSA는 PromptSensiScore (PSS)를 도입하여 LLM의 인스턴스 수준 프롬프트 민감도 측정, 목표 및 주관적 평가의 강건성 분석, 민감도와 디코딩 신뢰도 연결.
Large language models (LLMs) have demonstrated impressive capabilities across various tasks, but their performance is highly sensitive to the prompts utilized. This variability poses challenges for accurate assessment and user satisfaction. Current research frequently overlooks instance-level prompt variations and their implications on subjective evaluations. To address these shortcomings, we introduce ProSA, a framework designed to evaluate and comprehend prompt sensitivity in LLMs. ProSA incorporates a novel sensitivity metric, PromptSensiScore, and leverages decoding confidence to elucidate underlying mechanisms. Our extensive study, spanning multiple tasks, uncovers that prompt sensitivity fluctuates across datasets and models, with larger models exhibiting enhanced robustness. We observe that few-shot examples can alleviate this sensitivity issue, and subjective evaluations are also susceptible to prompt sensitivities, particularly in complex, reasoning-oriented tasks. Furthermore, our findings indicate that higher model confidence correlates with increased prompt robustness. We believe this work will serve as a helpful tool in studying prompt sensitivity of LLMs. The project is released at: https://github.com/open-compass/ProSA .
연구 동기 및 목표
- 프롬프트가 인스턴스 수준에서 데이터셋 수준이 아닌 LLM 응답에 어떤 영향을 주는지 동기 부여하고 정량화합니다.
- 인스턴스 수준의 민감도 지표(PSS)를 목표 및 주관적 평가를 위해 개발합니다.
- 프롬프트 민감도가 모델, 데이터셋 및 프롬프트 스타일에 따라 어떻게 달라지는지 조사합니다.
- 디코딩 신뢰도와 프롬프트 강건성 간의 관계를 탐구합니다.
제안 방법
- PSS를 동일 인스턴스에 대해 모든 프롬프트 변형 간 응답의 평균 쌍 차이로 정의합니다.
- objective 작업에서 여러 오픈 소스 LLM과 데이터셋을 대상으로 PSS를 평가합니다.
- Few-shot 프롬프트가 프롬프트 강건성에 미치는 영향을 평가합니다.
- LC AlpacaEval 2.0 및 Arena Hard Auto를 사용한 프롬프트 재작성으로 주관적 평가 분석을 수행합니다.
- 디코딩 신뢰도를 사용하여 프롬프트 민감도의 근본 원인을 분석합니다.

실험 결과
연구 질문
- RQ1인스턴스 수준의 프롬프트 민감도가 데이터셋과 모델 규모에 따라 어떻게 다른가요?
- RQ2Few-shot 프롬 prompting이 프롬프트 민감도를 줄이고 강건성을 높이나요?
- RQ3실제로 LLM의 디코딩 신뢰도와 프롬프트 강건성 사이의 관계는 무엇인가요?
- RQ4주관적 평가에서 프롬프트 민감도는 객관적 평가와 비교하여 어떻게 나타나나요?
- RQ5어떤 작업 범주에서 프롬프트 민감도가 더 높거나 낮은가요?
주요 결과
- 프롬프트 민감도는 데이터셋과 모델에 따라 달라지며, 더 큰 모델일수록 종종 더 큰 강건성을 보입니다.
- Few-shot 프롬 prompting은 특히 0-shot에서 1-shot으로 전환할 때 프롬프트 민감도를 줄이고, 더 큰 모델은 더 많은 샷에서 강건성을 얻습니다.
- 주관적 평가에서는 복잡한 작업에서 더 높은 민감도가 나타나지만, 단순한 작업에서는 강건성이 존재합니다.
- 디코딩 신뢰도는 프롬프트 민감도가 낮아지는 경향과 상관관계가 있으며, 이는 강건성이 내재된 디코딩 역학을 반영합니다.
- 프롬프트 범주가 민감도에 영향을 미치며, 지식 중심 작업은 코딩이나 창의적 작업보다 더 강건한 경향이 있습니다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.