Identity-Aware Human-Object Interaction Motion Captioning
Researchers have proposed a new task for human-object interaction (HOI) motion captioning that requires generated captions to specify both the subject identity and the corresponding HOI motion. They introduce the Identity-Aware Human-Object Interaction Motion Captioning task and design an architecture called ID-HOINet, which includes two core components: Multi-View Identity-Motion Learning Module (MVIML) and Two-Stage Caption Rewriting Strategy (TSCR). Experiments show that I
Researchers have proposed a new task for human-object interaction (HOI) motion captioning that requires generated captions to specify both the subject identity and the corresponding HOI motion. They introduce the Identity-Aware Human-Object Interaction Motion Captioning task and design an architecture called ID-HOINet, which includes two core components: Multi-View Identity-Motion Learning Module (MVIML) and Two-Stage Caption Rewriting Strategy (TSCR). Experiments show that ID-HOINet achieves state-of-the-art performance on the BEHAVE and InterCap datasets. The authors plan to release their code upon acceptance.
---
Why it matters: This work matters because it addresses a limitation in existing HOI motion captioning methods, which often use generic terms to refer to subjects rather than specifying their identities. This can improve the accuracy and specificity of generated captions, especially in applications where subject identity is important.
Source: https://arxiv.org/abs/2608.20690
This article was originally published at: https://arxiv.org/abs/2608.20690