Multimodal Rapport Estimation in Real-World HRI
Researchers have developed a method to estimate the quality of interactions between humans and robots in real-world environments. They used multimodal recordings from 62 sessions in a Japanese drugstore to test various AI models, including zero-shot language models and audio-visual models. The results show that these models can accurately predict rapport scores, but their performance varies depending on factors such as interaction duration and group size. This research sugges
Researchers have developed a method to estimate the quality of interactions between humans and robots in real-world environments. They used multimodal recordings from 62 sessions in a Japanese drugstore to test various AI models, including zero-shot language models and audio-visual models. The results show that these models can accurately predict rapport scores, but their performance varies depending on factors such as interaction duration and group size. This research suggests that current evaluation methods may not be suitable for real-world settings and need to be adapted to account for contextual variability.
---
Why it matters: This study's findings are important because they highlight the limitations of existing automatic evaluation methods in real-world human-robot interactions. Understanding how to accurately estimate interaction quality is crucial for developing robots that can adapt their behavior autonomously, which could improve user experience and safety.
Source: https://arxiv.org/abs/2608.18401
This article was originally published at: https://arxiv.org/abs/2608.18401