職場向けAIに人間らしい文書理解を学習させる新モデルを開発(New Model Teaches Workplace AI to Read More Like Humans)

2026-08-12 ジョージア工科大学

米ジョージア工科大学(Georgia Tech)の研究チームは、職場で利用されるAIが人間の意図や状況をより適切に理解できる新たなモデルを開発した。従来の業務支援AIは、明示的な指示には対応できても、利用者の行動背景や組織内の文脈、暗黙の意図を十分に把握できないという課題があった。新モデルは、人間同士が職場で相手の意図や状況を推測する仕組みに着想を得て設計されており、利用者の行動履歴や作業状況、対話内容などを統合的に解釈して支援を行う。これにより、AIは単なる応答システムではなく、利用者のニーズを先回りして支援する協働パートナーとして機能できる可能性が示された。研究成果は、人間とAIの協調作業の効率化や意思決定支援の高度化につながると期待され、今後の職場向けAIシステム設計やヒューマン・AIインタラクション研究に重要な知見を提供する。

<関連情報>

SlideAgent:複数ページにわたる視覚的文書の理解のための階層型エージェントフレームワーク SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

Yiqiao Jin, Rachneet Kaur, Zhen Zeng, Sumitra Ganesh, Srijan Kumar
arXiv  last revised 5 Jun 2026 (this version, v4)
DOI:https://doi.org/10.48550/arXiv.2510.26615

職場向けAIに人間らしい文書理解を学習させる新モデルを開発(New Model Teaches Workplace AI to Read More Like Humans)

Abstract

Multi-page visual documents such as manuals, brochures, presentations, and posters convey key information through layout, colors, icons, and cross-slide references. While multimodal large language models (MLLMs) offer opportunities in document understanding, current systems struggle with complex, multi-page visual documents, particularly in fine-grained reasoning over elements and pages. We introduce SlideAgent, a versatile agentic framework for understanding multi-modal, multi-page, and multi-layout documents, especially slide decks. SlideAgent employs specialized agents and decomposes reasoning into three specialized levels–global, page, and element–to construct a structured, query-agnostic representation that captures both overarching themes and detailed visual or textual cues. During inference, SlideAgent selectively activates specialized agents for multi-level reasoning and integrates their outputs into coherent, context-aware answers. Extensive experiments show that SlideAgent significantly improves accuracy over both proprietary (+7.9%) and open-source models (+9.8%).

1603情報システム・データ工学
ad
ad
Follow
ad
タイトルとURLをコピーしました