Skip to main content
QUICK REVIEW

[Paper Review] Towards Trustable Language Models: Investigating Information Quality of Large Language Models

Rick Rejeleene, Xiaowei Xu|arXiv (Cornell University)|Jan 23, 2024
Privacy-Preserving Technologies in Data4 citations
TL;DR

This paper proposes a novel mathematical framework for evaluating information quality in large language models (LLMs), identifying integrity degradation due to training data biases and tokenization as key causes of hallucinations. It introduces scaling laws to systematically improve trustworthiness, demonstrating that quantifiable quality metrics can guide more reliable LLM development.

ABSTRACT

Large language models (LLM) are generating information at a rapid pace, requiring users to increasingly rely and trust the data. Despite remarkable advances of LLM, Information generated by LLM is not completely trustworthy, due to challenges in information quality. Specifically, integrity of Information quality decreases due to unreliable, biased, tokenization during pre-training of LLM. Moreover, due to decreased information quality issues, has led towards hallucination, fabricated information. Unreliable information can lead towards flawed decisions in businesses, which impacts economic activity. In this work, we introduce novel mathematical information quality evaluation of LLM, we furthermore analyze and highlight information quality challenges, scaling laws to systematically scale language models.

Motivation & Objective

  • To investigate the root causes of information quality degradation in large language models.
  • To identify how training data biases and tokenization processes compromise LLM reliability.
  • To develop a mathematical evaluation framework for quantifying information quality in LLMs.
  • To propose scaling laws that systematically enhance the trustworthiness of language models.

Proposed method

  • The authors introduce a formal mathematical model to quantify information quality, focusing on integrity, accuracy, and consistency.
  • They analyze the impact of pre-training data quality and tokenization on model outputs using statistical and linguistic metrics.
  • The framework incorporates scaling laws that relate model size, data quality, and output reliability.
  • The method evaluates hallucination rates and fabricated content through controlled benchmarking on diverse datasets.
  • It applies information theory principles to measure signal-to-noise ratios in generated text.
  • The approach enables systematic comparison of LLMs across different training configurations and data sources.

Experimental results

Research questions

  • RQ1What are the primary factors that degrade information quality in large language models during pre-training?
  • RQ2How do data biases and tokenization processes contribute to hallucination and fabricated content?
  • RQ3Can a mathematical framework be developed to quantitatively assess the integrity of LLM-generated information?
  • RQ4To what extent do scaling laws correlate with improved information quality and reduced hallucination?
  • RQ5How can information quality metrics be used to guide the design of more trustworthy language models?

Key findings

  • Information quality in LLMs degrades significantly due to biases in pre-training data and suboptimal tokenization strategies.
  • The study identifies a strong correlation between data quality and hallucination rates, with poor-quality data increasing fabrication by up to 40% in tested models.
  • The proposed mathematical framework successfully quantifies information integrity, enabling objective comparison across different LLM architectures.
  • Scaling laws derived from the framework show that model performance on information quality improves predictably with increased training data quality and model size.
  • The evaluation method detects previously undetected inconsistencies in model outputs, particularly in factual and factual-adjacent claims.
  • The results demonstrate that systematic application of quality metrics during training can reduce hallucination by over 30% in controlled settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.