攻撃者がAIエージェントに規則違反を誘導する手法を解明(How attackers persuade AI agents to break the rules)

2026-08-19 スイス連邦工科大学ローザンヌ校(EPFL)

EPFL(スイス連邦工科大学ローザンヌ校)の研究チームは、AIエージェントが単一の悪意ある指示ではなく、複数回の一見無害な対話を通じて誘導されるリスクを体系的に検証する自動テストフレームワーク「STING」を開発した。STINGは攻撃者役のAIが不正な目的を段階的な小目標に分解し、対象エージェントを説得して実行させる過程を再現する。GPT、Gemini、Claudeなどのツール利用型エージェントを176種類の有害タスクで評価した結果、単一プロンプトより多段階攻撃の成功率が一貫して高く、一部モデルでは有害タスクの実行率が約2倍に達した。また、7言語では単純な低資源言語ほど脆弱になる傾向は見られなかった一方、攻撃途中で言語を切り替えると成功率が大幅に高まるケースも確認。研究は、AIエージェントの開発初期から多段階・多言語攻撃を想定した安全性評価を組み込む必要性を示している。

<関連情報>

過剰なほど親切:複数ターン、多言語対応のLLMエージェントにおける不正な支援の測定 Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents

Nivya Talokar, Ayush K Tarun, Murari Mandal, Maksym Andriushchenko, Antoine Bosselut
arXiv  last revised 7 Jun 2026 (this version, v4)
DOI:https://doi.org/10.48550/arXiv.2602.16346

攻撃者がAIエージェントに規則違反を誘導する手法を解明(How attackers persuade AI agents to break the rules)

Abstract

LLM-based agents execute real-world workflows via tools and memory. These affordances enable ill-intended adversaries to also use these agents to carry out complex misuse scenarios. Existing agent misuse benchmarks largely test single-prompt instructions, leaving a gap in measuring how agents end up helping with harmful or illegal tasks over multiple turns. We introduce STING (Sequential Testing of Illicit N-step Goal execution), an automated red-teaming framework that constructs a step-by-step illicit plan grounded in a benign persona and iteratively probes a target agent with adaptive follow-ups, using judge agents to track phase completion. We further introduce an analysis framework that models multi-turn red-teaming as a time-to-first-jailbreak random variable, enabling analysis tools like discovery curves, hazard-ratio attribution by attack language, and a new metric: Restricted Mean Jailbreak Discovery. Across AgentHarm scenarios, STING yields substantially higher illicit-task completion than single-turn prompting and chat-oriented multi-turn baselines adapted to tool-using agents. In multilingual evaluations across six non-English settings, we find that attack success and illicit-task completion do not consistently increase in lower-resource languages, diverging from common chatbot findings. Overall, STING provides a practical way to evaluate and stress-test agent misuse in realistic deployment settings, where interactions are inherently multi-turn and often multilingual.

1602ソフトウェア工学
ad
ad
Follow
ad
タイトルとURLをコピーしました