[Paper Review] A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law
This survey analyzes how large language models (LLMs) are applied in finance, healthcare, and law, evaluates their performance, discusses challenges and ethics, and outlines future directions with a focus on domain-specific tasks, datasets, and multimodal considerations.
In the fast-evolving domain of artificial intelligence, large language models (LLMs) such as GPT-3 and GPT-4 are revolutionizing the landscapes of finance, healthcare, and law: domains characterized by their reliance on professional expertise, challenging data acquisition, high-stakes, and stringent regulatory compliance. This survey offers a detailed exploration of the methodologies, applications, challenges, and forward-looking opportunities of LLMs within these high-stakes sectors. We highlight the instrumental role of LLMs in enhancing diagnostic and treatment methodologies in healthcare, innovating financial analytics, and refining legal interpretation and compliance strategies. Moreover, we critically examine the ethics for LLM applications in these fields, pointing out the existing ethical concerns and the need for transparent, fair, and robust AI systems that respect regulatory norms. By presenting a thorough review of current literature and practical applications, we showcase the transformative impact of LLMs, and outline the imperative for interdisciplinary cooperation, methodological advancements, and ethical vigilance. Through this lens, we aim to spark dialogue and inspire future research dedicated to maximizing the benefits of LLMs while mitigating their risks in these precision-dependent sectors. To facilitate future research on LLMs in these critical societal domains, we also initiate a reading list that tracks the latest advancements under this topic, which will be continually updated: \url{https://github.com/czyssrs/LLM_X_papers}.
Motivation & Objective
- Motivate the study of LLMs in high-stakes domains where professional expertise, sensitive data, and regulatory compliance are critical.
- Survey financial, healthcare, and legal NLP tasks, datasets, and LLMs to identify strengths, gaps, and future research directions.
- Highlight ethical considerations, domain-specific requirements (regulation, explainability, fairness), and the need for interdisciplinary collaboration.
Proposed method
- Catalog existing financial NLP tasks and datasets (SA, IE, QA, SMP, etc.) and note datasets and benchmarks.
- Summarize financial LLMs, pre-training vs instruction-tuning approaches, and evaluation results across tasks.
- Survey healthcare and legal NLP tasks, LLMs, and evaluation methodologies with attention to multimodal or structured data.
- Discuss ethics, domain-specific concerns, and regulatory considerations for LLM deployment in these domains.
- Provide a cross-domain synthesis of challenges and opportunities to guide future work.

Experimental results
Research questions
- RQ1What are the main NLP tasks and benchmarks used to evaluate LLMs in finance, healthcare, and law?
- RQ2How do financial, healthcare, and legal LLMs compare in terms of performance across standard tasks and datasets?
- RQ3What are the key methodological approaches (pre-training vs instruction-tuning, multimodal integration) shaping LLM capabilities in these domains?
- RQ4What ethical, regulatory, and transparency concerns dominate LLM adoption in these sectors, and how can they be mitigated?
Key findings
- LLMs are increasingly used across finance, healthcare, and law to enhance analytics, interpretation, and decision support, but performance varies by task and modality.
- In finance, specialized LLMs (e.g., BloombergGPT, FinMA variants, InvestLM) show improvements on sentiment analysis, QA, and information extraction, with multimodal and numerical reasoning remaining challenging.
- Healthcare applications span medical NLP, abnormality detection, medical report generation, and imaging-language tasks, with growing attention to instruction-following and evaluation in clinical contexts.
- Law-focused LLMs are advancing in tasks like contract analysis, statutory interpretation, and case-law QA, with emphasis on interpretability and reliability in high-stakes legal settings.
- Ethical considerations—privacy, data security, bias, explainability, and regulatory compliance—are central across all three domains, necessitating transparent, fair, and auditable AI systems.
- The surveyed literature highlights a trend from small-domain finetuning toward larger instruction-tuned and multilingual/multimodal LLMs, coupled with calls for better evaluation standards and data governance.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.