2026-09-25 スタンフォード大学
<関連情報>
- https://news.stanford.edu/stories/2026/09/ai-benchmarking-measurement-research
- https://hai.stanford.edu/news/the-tests-that-grade-ai-may-be-getting-it-wrong
- https://openreview.net/forum?id=889XnQKyhM
- https://arxiv.org/abs/2605.17173
AIベンチマークが実際に測定しているもの:収束妥当性と弁別妥当性を適用して56のAIベンチマークを検証する What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang
COLM 2026 Conference Published: 09 Jul 2026
Abstract
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt the lenses of convergent and discriminant validity from the social sciences into an approach for interrogating AI benchmarks, which we use to analyze 56 capability and safety benchmarks using 53 models. We assign benchmarks with substantively similar \textit{purported concepts} to a shared \textit{assigned concept}, and ask whether model rankings on benchmarks with the same assigned concept correlate strongly with one another, and whether model rankings on benchmarks with different assigned concepts correlate less strongly. We ask analogous questions using model scores at the item-level, drawing on item response theory (IRT) models. We find that correlations between model rankings on benchmarks with the same assigned safety concepts are often weak, suggesting that these assigned concepts may be conceptualized inconsistently across benchmarks. For benchmarks with assigned capability concepts (e.g., reasoning, knowledge), model rankings are often as strongly correlated among benchmarks with the same assigned concept as between benchmarks with different assigned concepts, suggesting that different assigned capability concepts may not discriminate that well from one another. In some cases, benchmarks that share benchmark design elements (e.g., task structure, score format) correlate more strongly with one another than benchmarks with the same assigned concept. Finally, model rankings on some individual benchmarks correlate more strongly with model rankings on benchmarks with a different assigned concept than with model rankings on benchmarks sharing their own assigned concept, suggesting that these benchmarks may measure a different concept than they purport to. For example, model rankings on BBQ-accuracy correlate more strongly with model rankings on benchmarks assigned with reasoning than with model rankings on benchmarks that share its assigned concept, bias. To support future empirical work on the validity of benchmarks, we release our extensive dataset of model outputs and scores at the item- and benchmark-level.
なぜ安全対策の基準は言語によって変化するのか? Why Do Safety Guardrails Degrade Across Languages?
Max Zhang, Ameen Patel, Sang T. Truong, Sanmi Koyejo
arXiv last revised 11 Aug 2026 (this version, v2)
DOI:https://doi.org/10.48550/arXiv.2605.17173
Abstract
Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several safety-driving factors into one, obscuring the specific cause(s) of safety failure. We introduce a latent variable model, a Multi-Group Item Response Theory (IRT) framework, that decouples language-agnostic safety robustness (θ), intrinsic prompt hardness (β), global language processing difficulty (γ), and a prompt-specific cross-lingual safety gap (τ). Using the MultiJail dataset, we evaluate the safety robustness of 61 model configurations across 5 closed-model families and 10 languages of varying resource, aggregating a dataset of 1.9 million responses. Exploratory Factor Analysis shows safety is primarily unidimensional: models refuse different harm types mainly through a shared mechanism. Contrary to the expected trend that safety degrades largely in low-resource languages, 22 model configurations are more vulnerable in English than in low-resource languages. Low-resource languages produce more uncertain responses (high entropy) than high-resource languages. Also, high-τ prompts cluster in physical harm categories like Theft and Weapons and lower-resource languages, trends validated through cross-dataset generalization. While global translation quality shows low correlation with τ, severe mistranslations drive high-bias outliers, as validated by native speakers. Cultural and conceptual grounding mismatches may also contribute to τ. In predictive validation, the IRT framework achieves AUC=0.940, and unlike rate baselines stays predictive when a whole language is held out (0.875). Our framework reveals concept-language vulnerabilities that aggregate metrics obscure, enabling fairer cross-lingual safety evaluation and targeted improvements in dataset construction.


