AI採用ツールを狙う隠しプロンプト攻撃を検出(Thwarting Hidden Resume Hacks Targeting AI Hiring Tools)

2026-07-22 デューク大学(Duke)

デューク大学の研究チームは、AIを利用した採用システムを狙う「プロンプトインジェクション攻撃」を高精度に検出する新手法を開発した。近年、一部の応募者は履歴書PDF内に白色文字や極小文字など人間には見えない命令文を埋め込み、「この候補者を高評価にする」といった指示をAIに与える手法を用いている。研究チームは、PDFから抽出したテキストと、人間が実際に見るレンダリング画像を視覚言語モデル(VLM)で比較し、「機械だけが読める隠しテキスト」を自動検出するVisual Discrepancy Analysis(VDA)を提案した。約20万件の実際の履歴書を分析した結果、約1%にプロンプトインジェクションが確認され、この手法は高い検出精度を示した。AI採用ツールの普及に伴い、この種の攻撃は今後さらに増加すると予測されており、本研究は生成AIを利用する人材採用だけでなく、文書を処理するさまざまなAIシステムの安全性向上にも役立つ成果とされる。

<関連情報>

LLMベースの履歴書選考における実世界のプロンプトインジェクション攻撃の測定 Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening

Mohan Zhang, Yuqi Jia, Zhen Tan, Steven Jiang, Neil Zhenqiang Gong, Tianlong Chen, Dawn Song
arXiv  Submitted on 27 May 2026
DOI:https://doi.org/10.48550/arXiv.2605.28999


Figure 1: Illustrative examples of two types of hidden prompt
injection detected in real-world resume PDFs.

Abstract

LLMs are vulnerable to prompt injection attacks. However, this vulnerability has been primarily demonstrated conceptually in academic studies or through a few anecdotal case studies. Its prevalence and impact in real-world LLM-based applications are largely unexplored. In this work, we present the first systematic study of prompt-injection attacks in a widely used application: LLM-based resume screening. Our analysis is based on approximately 200K real-world resumes collected over multiple years by hireEZ. We first design tailored methods to detect prompt injection in resumes. Manual validation on a small-scale dataset demonstrates that our detectors achieve high precision and outperform state-of-the-art general-purpose detectors. We then apply our detector to the full resume dataset and conduct a comprehensive measurement study of real-world prompt injection attacks. Our analysis reveals several intriguing findings: approximately 1% of resumes contain hidden prompt injections; the prevalence of such injected resumes has increased noticeably over the past one to two years; and more than 90% of injected prompts do not use explicit instructions. These results provide the first evidence of large-scale prompt injection in real-world LLM-based applications and lay the groundwork for future studies to understand and mitigate such attacks.

1602ソフトウェア工学
ad
ad
Follow
ad
タイトルとURLをコピーしました