ASDnB: Merging Face with Body Cues For Robust Active Speaker Detection
State-of-the-art Active Speaker Detection (ASD) approaches mainly use audio and facial features as input. However, the main hypothesis in this paper is that body dynamics is also highly correlated to "speaking" (and "listening") actions and should be particularly useful in wild conditions (e.

Evidence notes
- Experiments use TIAGo as the physical evaluation platform.
- The reported results provide an independent evaluation on TIAGo hardware.
Company context
Develops service and humanoid robots for research, industrial, and public-facing environments, with a focus on modular platforms and long-term deployment in real-world settings.