2026-08-17 国立精神・神経医療研究センター,早稲田大学

【図1】AIモデルの概要とロボットが学習した2種類の介護動作
A Scalable PV-RNNのアーキテクチャ:視覚と固有受容感覚(関節角度とトルク)を階層的に統合する。損失関数である変分自由エネルギーは、予測誤差(PE)及び、事前信念と事後信念のカルバックライブラー情報量(KLD)の総和で計算される
B 使用したヒューマノイドロボット(AIREC)
C 学習した2種類の介護動作タスクの動作遷移(体位変換、清拭)
<関連情報>
- https://www.ncnp.go.jp/topics/detail.php?@uid=Q4rFAKS1F9MpbLhs
- https://www.science.org/doi/10.1126/sciadv.aed7511
予測処理は、身体化されたマルチタスク知能のための拡張可能な計算原理である Predictive processing as a scalable computational principle for embodied multitask intelligence
Hayato Idei, Tamon Miyake, Tetsuya Ogata, and Yuichi Yamashita
Science Advances Published:14 Aug 2026
DOI:https://doi.org/10.1126/sciadv.aed7511
Abstract
Humans exhibit remarkable flexibility in adapting to diverse and uncertain environments—a hallmark arising from the brain’s ability to integrate multimodal sensory streams into coherent predictive models. Drawing on this principle, we introduce a scalable hierarchical multimodal recurrent neural network grounded in predictive processing under the free-energy principle, capable of directly integrating more than 30,000-dimensional visuo-proprioceptive inputs without dimensionality reduction or handcrafted preprocessing. Using sensory data from teleoperation of a full-scale physical humanoid robot performing two caregiving-related tasks—rigid-body repositioning and flexible-towel wiping—the model learns to predict high-dimensional visuo-proprioceptive streams end to end. In open-loop adaptive inference experiments, the framework exhibits three emergent properties: (i) self-organized hierarchical latent dynamics governing task transitions, uncertainty, and occlusion inference; (ii) robustness to degraded vision via multimodal integration; and (iii) asymmetric interference in multitask learning. Although evaluated in simulations, the framework is extensible to closed-loop robot control, with proprioceptive predictions driving action, thereby establishing a generalizable computational foundation bridging brain theory, artificial intelligence, and embodied robotics.
