Publications
JOURNAL (INTERNATIONAL) Exploring Motion-Text Matching for Motion Generation and Editing
Qing Yu, Kent Fujiwara
IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)
August 17, 2026
Recent advances in motion-language modeling have led to progress in retrieval and recognition tasks, yet these approaches remain fragmented from essential applications such as motion generation and editing. This disconnect limits their applicability in real-world, video-centric cross-modal systems where unified motion understanding is critical for analyzing and utilizing human motion in videos and animations. In this paper, we propose a unified and lightweight framework, Part-Level Motion-Text Matching (PL-MTM), designed to bridge motion retrieval, generation, and editing through fine-grained, part-aware representation learning. Crucially, PL-MTM and its LLM-enhanced variant, PL-MTM+, are designed to integrate part-level motion features with part-level language. Body-part-specific encoders and temporal modeling align part-wise motion segments with body-part descriptions, providing fine-grained semantic and spatiotemporal correspondence between parts. To enhance alignment robustness, we incorporate spatial and temporal masking strategies. Moreover, PL-MTM integrates seamlessly with pre-trained motion diffusion models, enabling training-free applications for both motion generation and editing. Extensive experiments on standard benchmarks demonstrate that PL-MTM improves motion-text alignment accuracy and delivers competitive performance across retrieval, generation, and editing tasks, making it well-suited for practical deployment in animation, robotics, and video content creation systems.
Paper :
Exploring Motion-Text Matching for Motion Generation and Editing
(external link)