Publications

OTHERS (INTERNATIONAL) ARMS: Anchor–Relational Motion Streaming for Seamless Solo-Social Motion Transitions

Huakun Liu (Nara Institute of Science and Technology), Qing Yu, Kent Fujiwara, Hideaki Uchiyama (Nara Institute of Science and Technology), Kiyoshi Kiyokawa (Nara Institute of Science and Technology)

The 19th European Conference on Computer Vision (ECCV 2026)

September 09, 2026

Generating temporally continuous and socially coherent human motion from text remains a fundamental challenge, particularly in realistic streams where people act alone, enter interactions, and later disengage. Most existing methods generate fixed-length motion clips under static agent configurations, which makes them brittle to solo–social transitions and unsuitable for incremental generation over long horizons. We propose ARMS, an Anchor–Relational Motion Streaming framework that unifies solo motion and human–human interaction within a single causal generative process. ARMS introduces a dynamics-asymmetric representation that decouples per-person temporal evolution from inter-person alignment via a partner-referenced relative-translation term, enabling seamless switching of social coupling without sacrificing long-horizon stability or spatial consistency between agents. On top of a causal latent space, a causal relational diffusion model progressively refines motion segment by segment using only past context, capturing both intra-person temporal dependencies and inter-person relations. Mode-aware relational gating activates or masks cross-agent connections, allowing the same model to support both solo and interaction generation. Experiments show that ARMS improves transition smoothness and social coherence compared to interaction-centric baselines, while also achieving competitive results on human–human interaction benchmarks.

Paper : ARMS: Anchor–Relational Motion Streaming for Seamless Solo-Social Motion Transitionsopen into new tab or window (external link)