日本語の表題:研究:生成AIは会話による誤情報の圧力や議論に屈する(Study: Generative AI succumbs to conversational misinformed pressure and argument)

2026-09-04 アリゾナ大学

アリゾナ大学の研究チームは、生成AIの大規模言語モデル(LLM)が、長時間の多ターン対話において誤情報や利用者からの論争的な圧力に影響される特性を調査した。7種類のLLMを比較した結果、モデルによって誤情報への追従性や説得されやすさに大きな違いがあり、特にChatGPT 3.5は、会話中に虚偽の主張を繰り返されると、それを再肯定する傾向が最も強かった。一方、Claude 3.5 Sonnetは最も影響を受けにくかった。また、全モデルで、専門性が低い・あまり知られていない話題ほど誤情報への抵抗力が低下した。さらに、一部のモデルは同じ誤情報を会話の途中で肯定・否定し直す「reverberation(反響)」と呼ばれる不安定な挙動を示した。研究者は、こうした再現性の低い特性が医療などの高リスク領域で安全上の問題になると指摘し、AIを盲目的に信頼せず、人間による検証とAIの診断・評価手法の開発が必要だとしている。

<関連情報>

持続的な会話的誤情報圧力下における大規模言語モデルの誤謬可能性、説得可能性、および修正可能性 Fallibility, persuadability, and correctability of large language models under sustained conversational misinformation pressure

Jordan Rodriguez,Zachary Hansen,Luis De Anda,Katelyn Rohrer,Camila Grubb,Enrique Noriega-Atala,Mihai Surdeanu & Marvin J. Slepian
Scientific Reports  Published:01 September 2026
DOI:https://doi.org/10.1038/s41598-026-68231-0  Early provide 

Abstract

Generative artificial intelligence has emerged as a transformative global force, yet the reliability of large language models (LLMs) under sustained conversational misinformation pressure remains poorly characterized. While recent work has begun to evaluate multi-turn LLM behavior, the full spectrum of conversational susceptibility—encompassing fallibility, persuadability, and correctability—has not been systematically assessed. Here we systematically evaluated seven widely used LLMs—ChatGPT (GPT-3.5, GPT-4o, GPT-4o-mini), Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek—across three dimensions of conversational susceptibility. We assessed fallibility (susceptibility to accepting misinformation under repetitive exposure), persuadability (susceptibility to acceptance under progressively argumentative pressure), and correctability (capacity to recognize and correct self-generated misinformation). One hundred purposefully false statements spanning a range of informational obscurity were presented across 50-repetition conversational sequences. Under repetitive engagement, misinformation affirmation rates ranged from 0.08% to 12.3%, a greater than 150-fold difference across architectures, with ChatGPT 3.5 demonstrating the greatest vulnerability and Claude 3.5 Sonnet the greatest resistance. A novel phenomenon of conversational reverberation was identified, in which models oscillated unpredictably between accepting and rejecting the same false statement across successive turns. Misinformation susceptibility was significantly modulated by informational obscurity under repetitive but not argumentative conditions (χ² = 11.13, p = 0.0038), implicating training data frequency as a determinant of factual resistance. Correctability was heterogeneous: four models achieved 100% self-correction, while the model with the lowest error rate failed to correct any of its rare errors, a dissociation with important implications for deployment. These findings demonstrate that conversational dynamics expose LLM failure modes invisible to standard evaluation, with direct consequences for model selection and use in truth-critical domains.

1602ソフトウェア工学
ad
ad
Follow
ad
タイトルとURLをコピーしました