Skip to main content
QUICK REVIEW

[論文レビュー] A Guide to Large Language Models in Modeling and Simulation: From Core Techniques to Critical Challenges

Philippe J. Giabbanelli|arXiv (Cornell University)|Feb 5, 2026
Artificial Intelligence in Healthcare and Education被引用数 0
ひとこと要約

この論文は、実務家向けにモデリングとシミュレーションにおけるLLMの活用ガイドを提供し、 prompting、ハイパーパラメータ、 augmentation、評価にわたる核心技術、一般的な落とし穴、および実践的考慮事項を詳述します。

ABSTRACT

Large language models (LLMs) have rapidly become familiar tools to researchers and practitioners. Concepts such as prompting, temperature, or few-shot examples are now widely recognized, and LLMs are increasingly used in Modeling & Simulation (M&S) workflows. However, practices that appear straightforward may introduce subtle issues, unnecessary complexity, or may even lead to inferior results. Adding more data can backfire (e.g., deteriorating performance through model collapse or inadvertently wiping out existing guardrails), spending time on fine-tuning a model can be unnecessary without a prior assessment of what it already knows, setting the temperature to 0 is not sufficient to make LLMs deterministic, providing a large volume of M&S data as input can be excessive (LLMs cannot attend to everything) but naive simplifications can lose information. We aim to provide comprehensive and practical guidance on how to use LLMs, with an emphasis on M&S applications. We discuss common sources of confusion, including non-determinism, knowledge augmentation (including RAG and LoRA), decomposition of M&S data, and hyper-parameter settings. We emphasize principled design choices, diagnostic strategies, and empirical evaluation, with the goal of helping modelers make informed decisions about when, how, and whether to rely on LLMs.

研究の動機と目的

  • LLMsがモデル化・シミュレーションのパイプラインにどのように適合するか、効果的に使用するために実務者が知るべき点を明らかにする。
  • prompting、データ処理、評価におけるエンジニアリング上の決定が再現性と透明性にどう影響するかを強調する。
  • M&Sの成果に影響を及ぼす一般的な誤解と決定論性の問題を特定する。
  • 忘却、ロール prompting、マルチモーダル入力といったLLMベースのM&Sシステムで浮上する概念を議論する。
  • 問題を分解し、M&S文脈でのLLMの知識を評価する実践的演習とガイドラインを提供する。

提案手法

  • promptingとハイパーパラメータといった基本的なLLM構成要素をレビューする。
  • Retrieval-Augmented Generation(RAG)や文脈的知識といった拡張手法を説明する。
  • 非決定論性の源と緩和戦略を論じる。
  • LLMsに過度に依存し、解決策を不適切に再発明することへの警告を行う。
  • モデルと出力の表現選択とそれが性能に与える影響を説明する。
  • prompting設計のコミュニケーション、再現性、自動化に関する実践的なガイダンスを提供する。

実験結果

リサーチクエスチョン

  • RQ1LLMsを効果的にモデリングとシミュレーションに活用する際の核となる構成要素とエンジニアリング上の決定は何か。
  • RQ2 prompting戦略、ハイパーパラメータ、拡張技術が再現性、性能、透明性にどのような影響を与えるか。
  • RQ3M&SにおけるLLMsを取り巻く一般的な誤りや誤信は何で、それをどう緩和できるか。
  • RQ4モデル構造と出力の表現がLLMの性能にどのように影響するか。
  • RQ5研究者がLLM対応のM&Sワークフローを批判的に検討するのに役立つ実践的なガイドラインと演習は何か。

主な発見

  • プロンプトは中心的なインターフェースであり、タスクの分解、検証プロンプト、明示的なタスク定義は信頼性を向上させる。
  • 温度やサンプリング戦略といったハイパーパラメータのデコードは出力と再現性に大きく影響し、影響はLLMとタスクによって異なる。
  • 文脈的拡張(RAG)は取得品質と統合の仕方次第で性能を向上させる場合もあれば妨げる場合もあり、慎重な promptingと評価が必要。
  • モデルの表現(例:エッジリスト対隣接行列)はLLMの性能に実質的な影響を及ぼす可能性がある。 表現の実証的評価を推奨。
  • 非決定論の源は温度以外にも存在し、堅牢なM&S成果のためには徹底した評価と緩和が必要。
  • prompting設計の自動化が進んでいるが、タスクの仕様と解釈には人間の入力が依然不可欠である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。