人と犬の自然なふれあいを捉える大規模4Dデータセット「InterPet4D」を構築 ーインタラクティブペットの動作生成研究を加速ー

2026-10-02 東京科学大学

東京科学大学とカーネギーメロン大学などの研究チームは、人と犬が自然にふれあう様子を映像・音声・3D動作として記録した大規模4Dデータセット「InterPet4D」を構築した。23名と13頭の犬を対象に161回の撮影を行い、約683万フレームを収集。12台のカメラに加えてカメラ付きメガネを用い、第三者視点と人の一人称視点、音声、身体・手の動き、犬の姿勢などを同期して記録した。さらに、このデータから、人の手の動きや音声を入力として犬の自然な3D動作を生成する「InterPetMoGen」を開発し、実際の犬の動きに近い生成性能を確認した。研究成果は、犬の行動理解だけでなく、バーチャルペット、ロボット犬、ペットの見守り、ゲーム・アニメーションなどへの応用が期待される。

人と犬の自然なふれあいを捉える大規模4Dデータセット「InterPet4D」を構築 ーインタラクティブペットの動作生成研究を加速ー

図1. 12台の同期カメラ、一人称映像、音声、3D姿勢推定を組み合わせたInerPet4Dの概要(下段)。

<関連情報>

InterPet4D:ペットの動きを生成するためのマルチモーダル4D人間・ペット相互作用データセット InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation

Yichen Peng, Jyun-Ting Song, Chen-Chieh Liao, Kris Kitani, Hideki Koike, Erwin Wu

The 19th European Conference on Computer Vision

arXiv  Submitted on 11 Jul 2026

DOI:https://doi.org/10.48550/arXiv.2607.10287

Abstract

Human-pet interaction estimation and generation remain underexplored due to the absence of a high-quality large-scale dataset. We present InterPet4D, the first multimodal dataset capturing natural interactions between humans and dogs. Using a synchronized multi-view capture system, we record human-dog obedience tasks and provide annotations for both humans and dogs, including multi-view and egocentric videos, segmentations, 2D and 3D keypoints, meshes, and audio tracks. InterPet4D consists of 6.8 million frames collected from 13 dogs of 11 breeds interacting with 23 human participants. We further introduce the InterPetMoGen framework for human-pet interaction motion generation. Our proposed model achieves an FID score of 11.21 and substantially outperforms the Seq2Seq and DiT baselines, demonstrating the effectiveness of InterPet4D for modeling realistic human-pet interactions.

1603情報システム・データ工学
ad
ad
Follow
ad
タイトルとURLをコピーしました