AIメンタルヘルス安全性評価の重大な欠陥を明らかに(Study exposes major flaw in AI mental health safety testing)

2026-07-13 スタンフォード大学

スタンフォード大学の研究チームは、AIチャットボットのメンタルヘルス安全性評価に重大な欠陥があることを明らかにした。現在、多くのAI開発企業は精神科医や心理学者などの専門家による評価を基に、AIの応答が「安全」かどうかを判定している。しかし研究では、同じAIの応答を評価しても専門家間の一致率が低く、安全性の判断基準が大きく異なることが判明した。この結果は、現在広く用いられている人手による安全評価だけでは、AIのメンタルヘルス支援能力を信頼性高く検証できないことを示している。研究チームは、安全性評価の標準化や客観的な評価基準の整備、複数の評価者を組み合わせた手法の導入が必要と指摘した。メンタルヘルス用途でAIチャットボットの利用が急速に拡大する中、本研究は利用者保護とAIの安全な社会実装に向けた評価手法の再構築が不可欠であることを示す重要な成果である。

<関連情報>

メンタルヘルスAIの安全性テストにおける専門家評価と人間によるフィードバックの限界 Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

Kiana Jafari, Paul Ulrich Nikolaus Rust, Duncan Eddy, Robbie Fraser, Nina Vasan, Darja Djordjevic, Akanksha Dadlani, Max Lamparth, Eugenia Kim, Mykel Kochenderfer
arXiv  last revised 8 May 2026 (this version, v3)
DOI:https://doi.org/10.48550/arXiv.2601.18061

Abstract

Learning from human feedback~(LHF) assumes that expert judgments, appropriately aggregated, yield valid ground truth for training and evaluating AI systems. We tested this assumption in mental health, where high safety stakes make expert consensus essential. Three certified psychiatrists independently evaluated LLM-generated responses using a calibrated rubric. Despite similar training and shared instructions, inter-rater reliability was consistently poor (ICC 0.087–0.295), falling below thresholds considered acceptable for consequential assessment. Disagreement was highest on the most safety-critical items. Suicide and self-harm responses produced greater divergence than any other category, and was systematic rather than random. One factor yielded negative reliability (Krippendorff’s α=−0.203), indicating structured disagreement worse than chance. Qualitative interviews revealed that disagreement reflects coherent but incompatible individual clinical frameworks, safety-first, engagement-centered, and culturally-informed orientations, rather than measurement error. By demonstrating that experts rely on holistic risk heuristics rather than granular factor discrimination, these findings suggest that aggregated labels function as arithmetic compromises that effectively erase grounded professional philosophies. Our results characterize expert disagreement in safety-critical AI as a sociotechnical phenomenon where professional experience introduces sophisticated layers of principled divergence. We discuss implications for reward modeling, safety classification, and evaluation benchmarks, recommending that practitioners shift from consensus-based aggregation to alignment methods that preserve and learn from expert disagreement.

1603情報システム・データ工学
ad
ad
Follow
ad
タイトルとURLをコピーしました