大規模言語モデルの推論信頼性を向上させる新しいフレームワークを開発(Researchers Develop New Frameworks to Improve LLM Reasoning Reliability)

2026-07-30 合肥物質科学研究院(HFIPS)

中国科学院合肥物質科学研究院の丁増輝教授らと米国カリフォルニア大学などの共同研究チームは、大規模言語モデル(LLM)の推論の信頼性を向上させるための新たな強化学習手法と評価フレームワークを開発した。従来の強化学習は最終回答の正誤のみを評価するOutcome Reward Model(ORM)が主流であったが、この方法では表面的なパターン学習に陥りやすいという課題があった。研究では、人間の認知制御機構に着想を得て、推論途中の各ステップを評価するProcess Reward Model(PRM)が、より適切な推論過程を学習させ、近道的な推論を抑制できることを示した。また、未知の課題への汎化性能を評価する新たな指標「Prism Benchmark OPG-D³」を提案し、Oracle Performance Gap(OPG)により通常学習モデルとテストデータ最適化モデルの性能差を測定することで、従来の評価法では見落とされていた汎化リスクを可視化した。これらの成果は、LLMの学習・推論メカニズムの理解を深め、より信頼性の高いAIシステムの開発に貢献すると期待される。

大規模言語モデルの推論信頼性を向上させる新しいフレームワークを開発(Researchers Develop New Frameworks to Improve LLM Reasoning Reliability)
Illustration of outcome reward-induced shortcuts and the corrective mechanism of process supervision (Image by DING Zenghui)

<関連情報>

結果最適化のパラドックス:LLMにおける推論ショートカットに対する因果情報理論的限界 The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs

Zihan Chen, Yiming Zhang, Wenxiang Geng, Zenghui Ding, Yining Sun
Association for Computational Linguistics  Published:July 2026
DOI:https://doi.org/10.18653/v1/2026.acl-long.925

Abstract

Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) frequently exhibit a critical failure mode: they achieve high performance on in-distribution benchmarks while demonstrating brittle reasoning capabilities on out-of-distribution (OOD) tasks. We term this phenomenon Reward-Induced Manifold Collapse. We establish a theoretical framework bridging Structural Causal Models (SCM) and the Information Bottleneck (IB) principle to explain this paradox. We define reasoning as a high-complexity causal process and shortcut learning as the exploitation of low-complexity spurious correlations. Under the implicit inductive bias of Stochastic Gradient Descent (SGD), models optimized for outcome rewards are biased toward shortcut solutions whenever the training distribution allows for a “Markovian Screening” of the true causal mechanism. We derive a new generalization bound based on Semantic Coverage Measure () rather than sample size, showing why data scaling on homogeneous distributions may fail to correct reasoning flaws. We also show that Process Reward Models (PRMs) function as Topological Filters, enforcing step-wise mutual information constraints that render the low-complexity shortcut manifold inadmissible. These results provide a mathematical grounding for the role of process supervision beyond simple credit assignment.

1603情報システム・データ工学
ad
ad
Follow
ad
タイトルとURLをコピーしました