Exploring Multimodal Turn-Taking Cues in Face-to-Face Conversation using Voice Activity Projection
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2606. 16731v1 Announce Type: cross Abstract: Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios.
Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses. This work presents a Multimodal Voice Activity Projection (MM-VAP) framework that extends the original audio-only VAP formulation to synchronized audio-visual inputs while preserving its self-supervised future-projection objective.
arXiv:2607. 07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.
arXiv:2609.10394v1 Announce Type: cross Abstract: Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the...
The paper evaluates audio‑visual predictive turn‑taking models trained on clean data when applied to a noisy cocktail‑party scenario derived from the AVCocktail dataset. Results show a consistent performance drop—up to 38% relative in weighted F1—across both audio and visual modalities, with fine‑tuning improving robustness but varying by modality and pre‑training data size. The study highlights differing generalisation and adaptation abilities of audio versus visual inputs and underscores the need for robust modelling strategies in noisy human interactions.
The study investigates how gaze and speech cues, together with perceived interpersonal closeness, predict turn‑taking outcomes in free four‑person conversations. Using the GaMMA corpus, logistic regression models were trained on interpretable features such as gaze transition motifs, entropy, addressee identity, mutual gaze, and speaker loudness to classify floor‑transfer events as gaps or overlaps. Results show that gaze features alone capture predictive structure, and combining them with loudness yields a robust classifier (ROC AUC = 0.76 ± 0.04) that remains effective even under noisy conditions.