AIをより信頼できるものにする(Making AI more trustworthy)

2026-09-23 カリフォルニア大学リバーサイド校(UCR)

米カリフォルニア大学リバーサイド校(UCR)主導の研究チームは、大規模言語モデル(LLM)が示す「自信」と「正しさ」が、必ずしも同じ内部情報に基づいていないことを明らかにした。Llama-3.1-8BとGemma-2-9Bを対象に、スパース・オートエンコーダーでモデル内部の特徴を解析したところ、不確実性、誤答、両者に共通する「交絡特徴」を識別できた。交絡特徴を抑制すると、精度を最大1.1%向上させながら不確実性を最大75%低減できた。また、特定の内部特徴から誤答を予測し、該当する質問への回答を拒否させることで、回答した問題に対する正答率を62%から81%へ向上させた。再学習を必要とせず、推論時の内部状態を操作できる点も特徴で、信頼性の高いAIシステム構築につながる可能性が示された。

<関連情報>

LLMの不確実性と正確性は同じ特徴量で符号化されているのか?スパースオートエンコーダによる機能的分離 Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders

Het Patel, Tiejin Chen, Hua Wei, Evangelos E. Papalexakis, Jia Chen
arXiv  Submitted on 21 Apr 2026
DOI:https://doi.org/10.48550/arXiv.2604.19974

AIをより信頼できるものにする(Making AI more trustworthy)

Abstract

Large language models can be uncertain yet correct, or confident yet wrong, raising the question of whether their output-level uncertainty and their actual correctness are driven by the same internal mechanisms or by distinct feature populations. We introduce a 2×2 framework that partitions model predictions along correctness and confidence axes, and uses sparse autoencoders to identify features associated with each dimension independently. Applying this to Llama-3.1-8B and Gemma-2-9B, we identify three feature populations that play fundamentally different functional roles. Pure uncertainty features are functionally essential: suppressing them severely degrades accuracy. Pure incorrectness features are functionally inert: despite showing statistically significant activation differences between correct and incorrect predictions, the majority produce near-zero change in accuracy when suppressed. Confounded features that encode both signals are detrimental to output quality, and targeted suppression of them yields a 1.1% accuracy improvement and a 75% entropy reduction, with effects transferring across the ARC-Challenge and RACE benchmarks. The feature categories are also informationally distinct: the activations of just 3 confounded features from a single mid-network layer predict model correctness (AUROC ~0.79), enabling selective abstention that raises accuracy from 62% to 81% at 53% coverage. The results demonstrate that uncertainty and correctness are distinct internal phenomena, with implications for interpretability and targeted inference-time intervention.

1602ソフトウェア工学
ad
ad
Follow
ad
タイトルとURLをコピーしました