強化学習を用いてドローンに戦略的追跡行動を学習させることに成功 (Teaching Drones to Play Tag)

2026-08-18 サンディア国立研究所(SNL)

米国サンディア国立研究所(Sandia National Laboratories)の研究チームは、ドローン同士が「鬼ごっこ(tag)」を行うことで、自律飛行や追跡・回避行動を学習する技術を開発した。研究では、複数のドローンが動的に変化する相手の行動をリアルタイムで予測しながら、自律的に追跡や回避を行うアルゴリズムを検証した。鬼ごっこという単純なゲーム環境を利用することで、複雑な現実環境に必要な経路計画、意思決定、協調行動、障害物回避などの能力を効率的に訓練できる。さらに、この手法はGPSが利用できない環境や通信制約下でも適応可能な自律システム開発につながる可能性がある。研究成果は、災害対応、捜索救助、インフラ点検、防衛用途などで利用される無人航空機システムの高度化に貢献すると期待される。

<関連情報>

マルチエージェント強化学習を用いたドローンスウォームの協調戦略の学習 Learning Cooperative Strategies for Drone Swarms using Multi-Agent Reinforcement Learning

Christian Llanes, Kyle A. Williams, Spencer W. Jensen, and Samuel Coogan

強化学習を用いてドローンに戦略的追跡行動を学習させることに成功 (Teaching Drones to Play Tag)

Abstract

In this work, we investigate cooperative strategies for an evader drone team of various sizes using multi-agent reinforcement learning in a multi-agent pursuit-evasion scenario. The objective of the evader team is to reach a goal with minimal velocity while not colliding with the pursuer team. The objective of the pursuer team is to defend the goal by catching evaders before they reach it. In this environment, we allow the pursuer to have superior control authority compared to the evader such that reaching the goal is challenging for the evader in a one-on-one scenario. The proposed strategy for an evader is to team up with an ally to lead pursuers into a collision with each other instead of intercepting the evader. We design policies using multi-agent proximal policy optimization, an actor-critic reinforcement learning method, and investigate how the learned strategy changes when we vary the size of the pursuer and evader teams. Finally, we demonstrate the learned policy’s sim-to-real capabilities through a hardware demonstration.

0301機体システム
ad
ad
Follow
ad
タイトルとURLをコピーしました