Skip to main content
QUICK REVIEW

[논문 리뷰] When Stability Fails: Hidden Failure Modes Of LLMS in Data-Constrained Scientific Decision-Making

Nazia Riasat|arXiv (Cornell University)|2026. 03. 16.
Scientific Computing and Data Management인용 수 0
한 줄 요약

본 논문은 LLM 보조 데이터 제약된 과학적 의사결정 과제에서 안정성, 정확성, 프롬프트 민감도, 출력 타당성을 구분하는 통제된 행동 평가를 도입하고, 높은 안정성이 ground truth와의 합의나 유효한 출력으로 이어지지 않음을 보여준다.

ABSTRACT

Large language models (LLMs) are increasingly used as decision-support tools in data-constrained scientific workflows, where correctness and validity are critical. However, evaluation practices often emphasize stability or reproducibility across repeated runs. While these properties are desirable, stability alone does not guar- antee agreement with statistical ground truth when such references are available. We introduce a controlled behavioral evaluation framework that explicitly sep- arates four dimensions of LLM decision-making: stability, correctness, prompt sensitivity, and output validity under fixed statistical inputs. We evaluate multi- ple LLMs using a statistical gene prioritization task derived from differential ex- pression analysis across prompt regimes involving strict and relaxed significance thresholds, borderline ranking scenarios, and minor wording variations. Our ex- periments show that LLMs can exhibit near-perfect run-to-run stability while sys- tematically diverging from statistical ground truth, over-selecting under relaxed thresholds, responding sharply to minor prompt wording changes, or producing syntactically plausible gene identifiers absent from the input table. Although sta- bility reflects robustness across repeated runs, it does not guarantee agreement with statistical ground truth in structured scientific decision tasks. These findings highlight the importance of explicit ground-truth validation and output validity checks when deploying LLMs in automated or semi-automated scientific work- flows.

연구 동기 및 목표

  • 데이터 제약된 과학적 워크플로에서 안정성 외에 LLM 평가 필요성 동기 부여.
  • 네 가지 의사결정 차원을 분리하는 제어된 행동 프레임워크 도입: 안정성, 정확성, 프롬프트 민감성, 출력 타당성.
  • LLM 출력과 비교하기 위한 ground-truth 참조로 고정된 differential expression(DE) 표 사용.
  • 다 varied thresholding and prompt wording 아래 통계적 유전자 우선순위화의 일반적인 실패 모드를 특성화한다.

제안 방법

  • 고정된 DESeq2 유래 차등발현 표를 입력으로 제공하고 여러 LLM(ChatGPT, Gemini, Claude)을 서로 다른 체제에서 질의한다.
  • 임계값을 변화시킴(엄격한 FDR ≤ 0.05, 느슨한 0.05 < FDR ≤ 0.10), 경계 순위, 및 약간의 프롬프트 단어 변화(P7a 대 P7b).
  • 네 가지 지표로 출력 평가: 실행 간 안정성(Jaccard), ground truth와의 일치(Jaccard vs truth), 프롬프트 민감도(프롬프트 간 차이), 출력 타당성(잘못된 유전자 식별자 존재 여부).
  • 결정적 프롬프트를 사용하고 구성당 10회 반복 실행으로 데이터 변동성으로부터 모델 행태를 분리.
  • 재현성을 위한 보충 저장소에 코드와 결과 제공.

실험 결과

연구 질문

  • RQ1높은 실행 간 안정성이 통계적 ground truth에 대한 정확성을 암시하는가?
  • RQ2고정된 입력에서 미세한 프롬프트 문구 변화가 LLM 의사결정 출력에 어떤 영향을 미치는가?
  • RQ3통계적 임계값을 완화하는 것이 LLM 기반 유전자 우선순위화에 어떤 영향을 미치는가?
  • RQ4안정적인 출력에도 불구하고 LLM이 잘못된 또는 허구의 유전자 식별자를 생성하는가?

주요 결과

  • LLMs는 ground truth와 일치하지 않으면서도 실행 간 안정성이 거의 완벽한 모습을 보일 수 있다.
  • 프롬프트의 작은 문구 차이가 우선순위화 결과를 현저히 바꿀 수 있다.
  • 경계값의 완화가 신뢰할 수 있는 민감도 개선보다 과선택이나 붕괴를 촉진한다.
  • 모델은 입력에 존재하지 않는 구문상 타당하지만 잘못된 유전자 식별자를 생성할 수 있어 출력 타당성 문제가 나타난다.
  • 안정성은 내부적 강건성을 반영하지만 결정론적 통계 참조와의 합의를 반드시 보장하지는 않는다.
  • 데이터 제약된 과학적 워크플로에서 LLM의 동작을 진단하기 위한 4차원 평가 프레임워크가 필요하다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.