昆虫の「歩き方」をAIが解読しロボットへ ―数歩分のデータで六脚ロボットが適応歩行―

2026-08-19 東北大学

東北大学とタイ王国ウィタヤシリメティー科学技術大学院大学(VISTEC)の研究グループは、ナナフシのわずか数歩分の歩行データから、歩行を生み出す「報酬構造」と脚の協調ルールをAIが自動抽出する手法を開発した。従来の多脚ロボットでは歩行ルールを人が設計する必要があったが、本研究では敵対的逆強化学習を用いて昆虫の行動原理を推定した。その結果、平地で取得したデータのみを学習したにもかかわらず、凹凸地形や脚を1本失った状況でも、昆虫に類似した柔軟な歩行パターンが自律的に生成された。さらに、抽出した報酬構造は体格や構造の異なる六脚ロボットにも適用可能で、歩行学習速度を約3倍向上させることを確認した。実機ロボットによる検証にも成功しており、生物の適応的な歩行原理をロボットへ移植できることを示した。本成果は、未知環境や災害現場で活動する多脚ロボットの高性能化に加え、生物の運動制御原理の解明にも貢献すると期待される。

昆虫の「歩き方」をAIが解読しロボットへ ―数歩分のデータで六脚ロボットが適応歩行―
図1. ナナフシの歩行データから移植可能な歩行制御則を学習する枠組み

<関連情報>

昆虫の行動から応用可能なロボットの移動まで:敵対的逆強化学習による限られたデータからの身体化された移動原理の推論 From insect behavior to transferable robot locomotion: inferring embodied locomotor principles from limited data via adversarial inverse reinforcement learning

Yuchen Wang, Thirawat Chuthong, Mitsuhiro Hayashibe, Poramate Manoonpong and Dai Owaki
Bioinspiration & Biomimetics  Published: 11 August 2026
DOI:10.1088/1748-3190/ae901f

Abstract

Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. However, the underlying principles that govern leg coordination remain difficult to extract and model computationally. Understanding how insects achieve stable and adaptive locomotion has long provided important inspiration for the development of control strategies in bio-inspired robotics. Nevertheless, many existing approaches rely on predefined coordination rules, manually tuned parameters, or hand-crafted reward functions, which limit the flexibility and transferability of the resulting control strategies. To address this limitation, this study proposes a data-driven framework based on adversarial inverse reinforcement learning , which directly learns continuous locomotion control policies from stick insect walking data and infers latent reward structures and control strategies from biological behavioral demonstrations. Experimental results show that even when trained using only a short segment of flat-terrain demonstration data, the learned policy can still be extracted to learn adaptive and flexible leg coordination patterns under different environmental conditions. Furthermore, the learned reward network can be transferred across different dynamic systems to guide policy learning for robot models with different morphologies. Compared with methods relying solely on reward shaping, the proposed approach achieves faster convergence and produces more biologically consistent gait coordination. A preliminary deployment on a physical bio-inspired robot further demonstrates the potential of the learned policy for sim-to-real application. The proposed method provides a transferable data-driven framework for extracting and learning locomotion control strategies from biological behavior, with potential applications to bio-inspired robotic systems.

0109ロボット
ad
ad
Follow
ad
タイトルとURLをコピーしました