AIを評価するテストは間違っているかもしれない(The tests that grade AI may be getting it wrong)

2026-09-25 スタンフォード大学

スタンフォード大学の研究者らは、AIモデルの性能や安全性、偏り、推論能力などを測るベンチマーク自体が、本当に測定したい能力を測れているのかを検証した。測定科学・心理計量学の手法をAI評価に適用し、56種類の広く利用されているベンチマークを調査したところ、同じ能力を測るとされるベンチマーク同士でも結果が一致しないケースが繰り返し確認された。例えば、偏りを測るテストが、実際にはモデルの読解能力や「ひっかけ問題」への対応力を測ってしまう可能性がある。また、英語以外の言語では、翻訳による問題難度の変化とAIの安全性低下を単一スコアから区別できない。研究者らは、収束的妥当性・弁別的妥当性などの測定理論を導入し、AIベンチマークをより科学的かつ再現可能な評価手段にする必要性を指摘している。

<関連情報>

AIベンチマークが実際に測定しているもの:収束妥当性と弁別妥当性を適用して56のAIベンチマークを検証する What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang
COLM 2026 Conference  Published: 09 Jul 2026

Abstract

Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt the lenses of convergent and discriminant validity from the social sciences into an approach for interrogating AI benchmarks, which we use to analyze 56 capability and safety benchmarks using 53 models. We assign benchmarks with substantively similar \textit{purported concepts} to a shared \textit{assigned concept}, and ask whether model rankings on benchmarks with the same assigned concept correlate strongly with one another, and whether model rankings on benchmarks with different assigned concepts correlate less strongly. We ask analogous questions using model scores at the item-level, drawing on item response theory (IRT) models. We find that correlations between model rankings on benchmarks with the same assigned safety concepts are often weak, suggesting that these assigned concepts may be conceptualized inconsistently across benchmarks. For benchmarks with assigned capability concepts (e.g., reasoning, knowledge), model rankings are often as strongly correlated among benchmarks with the same assigned concept as between benchmarks with different assigned concepts, suggesting that different assigned capability concepts may not discriminate that well from one another. In some cases, benchmarks that share benchmark design elements (e.g., task structure, score format) correlate more strongly with one another than benchmarks with the same assigned concept. Finally, model rankings on some individual benchmarks correlate more strongly with model rankings on benchmarks with a different assigned concept than with model rankings on benchmarks sharing their own assigned concept, suggesting that these benchmarks may measure a different concept than they purport to. For example, model rankings on BBQ-accuracy correlate more strongly with model rankings on benchmarks assigned with reasoning than with model rankings on benchmarks that share its assigned concept, bias. To support future empirical work on the validity of benchmarks, we release our extensive dataset of model outputs and scores at the item- and benchmark-level.


なぜ安全対策の基準は言語によって変化するのか? Why Do Safety Guardrails Degrade Across Languages?

Max Zhang, Ameen Patel, Sang T. Truong, Sanmi Koyejo
arXiv  last revised 11 Aug 2026 (this version, v2)
DOI:https://doi.org/10.48550/arXiv.2605.17173

Abstract

Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several safety-driving factors into one, obscuring the specific cause(s) of safety failure. We introduce a latent variable model, a Multi-Group Item Response Theory (IRT) framework, that decouples language-agnostic safety robustness (θ), intrinsic prompt hardness (β), global language processing difficulty (γ), and a prompt-specific cross-lingual safety gap (τ). Using the MultiJail dataset, we evaluate the safety robustness of 61 model configurations across 5 closed-model families and 10 languages of varying resource, aggregating a dataset of 1.9 million responses. Exploratory Factor Analysis shows safety is primarily unidimensional: models refuse different harm types mainly through a shared mechanism. Contrary to the expected trend that safety degrades largely in low-resource languages, 22 model configurations are more vulnerable in English than in low-resource languages. Low-resource languages produce more uncertain responses (high entropy) than high-resource languages. Also, high-τ prompts cluster in physical harm categories like Theft and Weapons and lower-resource languages, trends validated through cross-dataset generalization. While global translation quality shows low correlation with τ, severe mistranslations drive high-bias outliers, as validated by native speakers. Cultural and conceptual grounding mismatches may also contribute to τ. In predictive validation, the IRT framework achieves AUC=0.940, and unlike rate baselines stays predictive when a whole language is held out (0.875). Our framework reveals concept-language vulnerabilities that aggregate metrics obscure, enabling fairer cross-lingual safety evaluation and targeted improvements in dataset construction.

1603情報システム・データ工学
ad
ad
Follow
ad
タイトルとURLをコピーしました