Papers by Naomi Harte
Neural Generation of Dialogue Response Timings (2020.acl-main)
Copied to clipboard
| Challenge: | Using neural models, the timings of spoken response offsets in human dialogue can vary based on contextual elements of the dialogue. |
| Approach: | They propose neural models that simulate the distributions of response offsets taking into account the response turn as well as the preceding turn. |
| Outcome: | The proposed models can generate distributions of response offsets based on the response turn and preceding turn based upon human listening tests and offline experiments. |
Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction (2025.findings-acl)
Copied to clipboard
| Challenge: | Predictive turn-taking models are based on speech, yet most rely on audio-only cues. |
| Approach: | They propose a multimodal PTTM which combines speech with visual cues including facial expression, head pose and gaze. |
| Outcome: | The proposed model outperforms the state-of-the-art audio-only turn-taking model (84% vs. 79% hold/shift prediction accuracy) it also outperformed the previous models which aggregated all holds and shifts, but grouped by duration of silence between turns. |
RoomReader: A Multimodal Corpus of Online Multiparty Conversational Interactions (2022.lrec-1)
Copied to clipboard
Justine Reverdy, Sam O’Connor Russell, Louise Duquenne, Diego Garaialde, Benjamin R. Cowan, Naomi Harte
| Challenge: | The corpus of multimodal, multiparty conversational interactions explored in RoomReader can be used to study a wide range of phenomena in online multimodal interaction. |
| Approach: | They propose to use RoomReader to explore multimodal cues of conversational engagement and behavioural aspects of collaborative interaction in online environments. |
| Outcome: | The corpus was developed within the wider RoomReader Project to explore multimodal cues of conversational engagement and behavioural aspects of collaborative interaction in online environments. |