脳の「予測する仕組み」でロボットが複数の介護動作を学習 ―脳科学理論から次世代Physical AIへの道を拓く―

2026-08-17 国立精神・神経医療研究センター,早稲田大学

国立精神・神経医療研究センター(NCNP)と早稲田大学の研究チームは、脳科学の有力理論である「予測処理(Predictive Processing)」および自由エネルギー原理に基づく新しいAIモデル「Scalable PV-RNN」を開発し、多自由度ヒューマノイドロボットに複数の介護動作を学習させることに成功した。研究では、介護ロボットAIRECに約3万次元の視覚情報と身体感覚情報を統合するAIを実装し、「体位変換」と「清拭」という性質の異なる介護作業を単一モデルで学習させた。その結果、AIは予測誤差を最小化するという単一原理のみで複雑な多感覚情報を処理し、複数タスクを柔軟に遂行できた。さらに、視覚情報が不十分な場合に身体感覚で補完する機能や、作業切り替えの自律的学習、遮蔽された対象の推定、不確実性の認識といった、人間の脳に類似した情報処理特性が自然に形成されることも確認された。本成果は、予測処理がロボット知能を実現する汎用的な計算原理となり得ることを示すものであり、介護、物流、製造、災害対応など幅広い分野で活用される次世代Physical AIや脳型AIの実現に向けた重要な前進と評価される。

【図1】AIモデルの概要とロボットが学習した2種類の介護動作
【図1】AIモデルの概要とロボットが学習した2種類の介護動作
A  Scalable PV-RNNのアーキテクチャ:視覚と固有受容感覚(関節角度とトルク)を階層的に統合する。損失関数である変分自由エネルギーは、予測誤差(PE)及び、事前信念と事後信念のカルバックライブラー情報量(KLD)の総和で計算される
B  使用したヒューマノイドロボット(AIREC)
C  学習した2種類の介護動作タスクの動作遷移(体位変換、清拭)

<関連情報>

予測処理は、身体化されたマルチタスク知能のための拡張可能な計算原理である Predictive processing as a scalable computational principle for embodied multitask intelligence

Hayato Idei, Tamon Miyake, Tetsuya Ogata, and Yuichi Yamashita
Science Advances  Published:14 Aug 2026
DOI:https://doi.org/10.1126/sciadv.aed7511

Abstract

Humans exhibit remarkable flexibility in adapting to diverse and uncertain environments—a hallmark arising from the brain’s ability to integrate multimodal sensory streams into coherent predictive models. Drawing on this principle, we introduce a scalable hierarchical multimodal recurrent neural network grounded in predictive processing under the free-energy principle, capable of directly integrating more than 30,000-dimensional visuo-proprioceptive inputs without dimensionality reduction or handcrafted preprocessing. Using sensory data from teleoperation of a full-scale physical humanoid robot performing two caregiving-related tasks—rigid-body repositioning and flexible-towel wiping—the model learns to predict high-dimensional visuo-proprioceptive streams end to end. In open-loop adaptive inference experiments, the framework exhibits three emergent properties: (i) self-organized hierarchical latent dynamics governing task transitions, uncertainty, and occlusion inference; (ii) robustness to degraded vision via multimodal integration; and (iii) asymmetric interference in multitask learning. Although evaluated in simulations, the framework is extensible to closed-loop robot control, with proprioceptive predictions driving action, thereby establishing a generalizable computational foundation bridging brain theory, artificial intelligence, and embodied robotics.

0109ロボット
ad
ad
Follow
ad
タイトルとURLをコピーしました