大規模言語モデルの推論信頼性を向上させる新たなフレームワーク (New Frameworks Improve Large Language Model Reasoning Reliability)

2026-07-30 中国科学院(CAS)

中国科学院(CAS)合肥物質科学研究院の丁増輝教授らの研究グループは、大規模言語モデル(LLM)の推論信頼性を向上させる新たな強化学習手法と、その性能を評価するベンチマーク「Prism Benchmark OPG-D3」を開発した。研究はACL 2026に採択された。従来の強化学習では最終回答の正誤のみを評価するOutcome Reward Model(ORM)が主流であったが、この方法では表面的なパターン学習や推論の近道(ショートカット)に依存しやすいことが判明した。一方、人間の認知制御機構に着想を得たProcess Reward Model(PRM)は、推論途中の各ステップに対してフィードバックを与えることで、より堅牢な推論能力を育成できることを示した。また、学習データを増やすだけでは問題は解決せず、推論過程への構造化された監督が重要であることも明らかになった。さらに、新たに提案したPrism Benchmark OPG-D3は、未知の課題への一般化能力を評価する枠組みであり、「Oracle Performance Gap(OPG)」指標により従来の評価法では見落とされる一般化リスクを検出できる。これらの成果は、より信頼性の高いAIシステムの開発に向けた学習手法と評価法の改善に重要な知見を提供している。

大規模言語モデルの推論信頼性を向上させる新たなフレームワーク (New Frameworks Improve Large Language Model Reasoning Reliability)
Illustration of outcome reward-induced shortcuts and the corrective mechanism of process supervision. (Image by DING Zenghui)

<関連情報>

結果最適化のパラドックス:LLMにおける推論ショートカットに対する因果情報理論的限界 The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs

Zihan Chen, Yiming Zhang, Wenxiang Geng, Zenghui Ding, Yining Sun
Association for Computational Linguistics  Published:July 2026
DOI:https://doi.org/10.18653/v1/2026.acl-long.925

Abstract

Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) frequently exhibit a critical failure mode: they achieve high performance on in-distribution benchmarks while demonstrating brittle reasoning capabilities on out-of-distribution (OOD) tasks. We term this phenomenon Reward-Induced Manifold Collapse. We establish a theoretical framework bridging Structural Causal Models (SCM) and the Information Bottleneck (IB) principle to explain this paradox. We define reasoning as a high-complexity causal process and shortcut learning as the exploitation of low-complexity spurious correlations. Under the implicit inductive bias of Stochastic Gradient Descent (SGD), models optimized for outcome rewards are biased toward shortcut solutions whenever the training distribution allows for a “Markovian Screening” of the true causal mechanism. We derive a new generalization bound based on Semantic Coverage Measure () rather than sample size, showing why data scaling on homogeneous distributions may fail to correct reasoning flaws. We also show that Process Reward Models (PRMs) function as Topological Filters, enforcing step-wise mutual information constraints that render the low-complexity shortcut manifold inadmissible. These results provide a mathematical grounding for the role of process supervision beyond simple credit assignment.


強化学習の評価を再考する:ベンチマークは強化学習手法の失敗を本当に明らかにできるのか? Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

Zihan Chen, Yiming Zhang, Hengguang Zhou, Zenghui Ding, Yining Sun, Cho-Jui Hsieh
Association for Computational Linguistics  Published:July 2026
DOI:https://doi.org/10.18653/v1/2026.findings-acl.769

Abstract

Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs). Despite recent benchmark gains reported for RL, we find that training on these benchmarks’ training sets achieves nearly the same performance as training directly on the test sets, suggesting that the benchmarks cannot reliably separate further progress. To study this phenomenon, we introduce a diagnostic suite and the Oracle Performance Gap (OPG) metric that quantifies the performance difference between training on the train split versus the test split of a benchmark. We further analyze this phenomenon with stress tests and find that, despite strong benchmark scores, existing RL methods struggle to generalize across distribution shifts, varying levels of difficulty, and counterfactual scenarios: shortcomings that current benchmarks fail to reveal. We conclude that current benchmarks are insufficient for evaluating generalization and propose three core principles for designing more faithful benchmarks: sufficient difficulty, balanced evaluation, and distributional robustness.

1603情報システム・データ工学
ad
ad
Follow
ad
タイトルとURLをコピーしました