Papers by Antoine Laurent

6 papers
A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case Study (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of spontaneous speech is being developed for educational use . the dataset will be freely available to the research community .
Approach: They propose to use a French speech educational corpus to explore synchronous speech transcription and application in teaching situations.
Outcome: The proposed corpus includes 10 hours of lectures, manually transcribed and segmented . the dataset will be freely available to the research community .
Automatic Speech Interruption Detection: Analysis, Corpus, and System (2024.lrec-main)

Copied to clipboard

Challenge: Interruption detection is a new but challenging task in the field of speech processing.
Approach: They propose to define automatic speech interruption detection and build a specialized corpus to analyze interrupted conversations.
Outcome: The proposed system can detect interruptions in speech with promising results . it can be used to ensure speaking turns are respected during official political debates .
Where are we in Named Entity Recognition from Speech? (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition is usually made through a pipeline process that consists of processing audio and applying a NER to the audio outputs.
Approach: They propose an original 3-pass approach and explore the capability of an E2E system to do structured NER.
Outcome: The proposed system performs better than the current pipeline approach.
Overlaps and Gender Analysis in the Context of Broadcast Media (2022.lrec-1)

Copied to clipboard

Challenge: Using gender and overlap annotations, we characterise interactions between speakers according to their gender and role in broadcast media.
Approach: They propose to characterise interactions between speakers according to their gender and role in broadcast media by using a small dataset of 93 recordings from LCP French channel.
Outcome: The proposed method could improve the efficiency of qualitative studies conducted in human sciences.
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification. (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods for creating diachronic corpus of voices are based on speaker characteristics and require human intervention.
Approach: They propose to use a semi-automatic pipeline to create a diachronic corpus of voices balanced for speaker’s age, gender and recording period, according to 32 categories.
Outcome: The proposed method cut down on manual annotations by ten and provides high quality speech for most of the selected excerpts.
Annotation of Transition-Relevance Places and Interruptions for the Description of Turn-Taking in Conversations in French Media Content (2024.lrec-main)

Copied to clipboard

Challenge: Few speech resources describe interruption phenomena, especially for TV and media content.
Approach: They propose to annotation Transition-Relevance Places (TRPs) and Floor-Taking event types on an existing French TV and Radio broadcast corpus to facilitate studies of interruptions and turn-taking.
Outcome: The proposed annotations on an existing French TV and Radio broadcast corpus show they are reliable and reliable .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations