Papers with diarization
Role-specific Language Models for Processing Recorded Neuropsychological Exams (N18-2)
Copied to clipboard
| Challenge: | Neuropsychological examinations are an important screening tool for the presence of cognitive conditions such as Alzheimer's, Parkinson's and spinal-cord injuries. |
| Approach: | They propose to use audio recordings to determine the cognitive health of 92 subjects from audio that was diarized using an automatic speech recognition system trained on TED talks and on structured language used by testers and subjects. |
| Outcome: | The proposed method can determine the cognitive health of 92 subjects from audio that was diarized using an automatic speech recognition system trained on TED talks and on the structured language used by testers and subjects. |
Computer-assisted Speaker Diarization: How to Evaluate Human Corrections (L18-1)
Copied to clipboard
| Challenge: | a framework to evaluate the human corrections of a speaker diarization is presented for the French National Audiovisual Institute (INA) the speaker diaarization task is a necessary pre-processing step for speaker identification and speech transcription. |
| Approach: | They propose a framework to evaluate the human corrections of a speaker diarization . they propose four elementary actions to correct the diarized speaker and an automaton to simulate the correction sequence. |
| Outcome: | The proposed framework copes with the needs of the French National Audiovisual Institute (INA) due to the increasing number of documents and the limited number of annotators, many documents remain undocumented or only partly documented. |
VAST: A Corpus of Video Annotation for Speech Technologies (L18-1)
Copied to clipboard
| Challenge: | The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development . |
| Approach: | The Video Annotation for Speech Technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development . |
| Outcome: | The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech detection, language identification, speaker identification, and speech recognition . |
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR (2024.emnlp-main)
Copied to clipboard
Shashi Kumar, Srikanth Madikeri, Juan Pablo Zuluaga Gomez, Iuliia Thorbecke, Esaú Villatoro-tello, Sergio Burdisso, Petr Motlicek, Karthik S, Aravind Ganapathiraju
| Challenge: | Existing approaches to automatic speech recognition use cascaded pipelines for tasks like voice activity detection, diarization, transcription and subsequent processing. |
| Approach: | They propose a single Transducer-based model that integrates task-specific tokens into the reference text during ASR model training, streamlining inference and eliminating the need for separate NLP models. |
| Outcome: | The proposed model outperforms the existing pipeline on speaker change detection, endpointing, and NER tasks while outperforming the existing model in individual task performance. |