Computer-assisted Speaker Diarization: How to Evaluate Human Corrections (L18-1)
Copied to clipboard
| Challenge: | a framework to evaluate the human corrections of a speaker diarization is presented for the French National Audiovisual Institute (INA) the speaker diaarization task is a necessary pre-processing step for speaker identification and speech transcription. |
| Approach: | They propose a framework to evaluate the human corrections of a speaker diarization . they propose four elementary actions to correct the diarized speaker and an automaton to simulate the correction sequence. |
| Outcome: | The proposed framework copes with the needs of the French National Audiovisual Institute (INA) due to the increasing number of documents and the limited number of annotators, many documents remain undocumented or only partly documented. |
Similar Papers
Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Tutorials (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | NAACL 2021 tutorials session is a conference for researchers to present on a topic of importance . a total of 35 tutorial submissions were received, of which 6 were selected for presentation . |
| Approach: | NAACL 2021 tutorials session is organized to give conference attendees a comprehensive introduction from expert researchers to a topic of importance drawn from our research field. |
| Outcome: | the tutorials committee selected 6 tutorials for presentation at NAACL 2021 . the topics chosen this year range from transformers to crowdsourcing . |
ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | a meta corpus of audio files is used to gather, annotate and transcribe speech . a large number of speech databases are needed to perform multi-speaker tasks such as speaker diarization and speaker change detection. |
| Approach: | They propose to use human feedback to homogenize and correct speaker labels among the audio files by integrating human feedback within a speaker verification system. |
| Outcome: | The proposed protocol evaluates speech segmentation, speaker diarization, speech transcription and speaker change detection using human feedback. |
Automated Evaluation of Out-of-Context Errors (L18-1)
Copied to clipboard
| Challenge: | Existing methods to modify text understanding systems use only one sentence at a time . however, considering a larger context can improve performance for text understanding tasks. |
| Approach: | They propose to modify existing text data to insert out-of-context errors . they use a 2016 TEDTalk corpus to evaluate computational models for text understanding . |
| Outcome: | The proposed method targets real-world problems of transcription and translation systems by inserting authentic out-of-context errors. |
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation (2024.lrec-main)
Copied to clipboard
Matthias Sperber, Ondřej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polák, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi
| Challenge: | a meta-analysis of human evaluation for speech translation has not been conducted . noisy data and segmentation mismatches are challenges for automatic metrics . |
| Approach: | They propose an evaluation strategy based on automatic resegmentation and direct assessment with segment context. |
| Outcome: | The proposed evaluation strategy is robust and scores well-correlated with other types of human judgements. |
Toward Machine Interpreting: Lessons from Human Interpreting Studies (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current speech translation systems are static and do not adapt to real-world situations in ways human interpreters do. |
| Approach: | They propose to model human interpreting using a new language model to improve usability . they argue that there is great potential to adopt many human interpreted principles . |
| Outcome: | The proposed models can be used to improve human interpreting and improve translation performance. |
Automatic Pronunciation Assessment - A Review (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Pronunciation assessment and its application in computer-aided pronunciation training (CAPT) have seen impressive progress in recent years. |
| Approach: | They review methods employed in computer-aided pronunciation training for both phonemic and prosodic pronunciations. |
| Outcome: | The proposed system should be able to automatically score non-native speech segments and give meaningful feedback. |
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: System Demonstrations) (2024.naacl-demo)
Copied to clipboard
| Challenge: | The System Demonstration Track at NAACL 2024 is a platform for presenting papers that describe system demonstrations. |
| Approach: | the system demonstration track at NAACL 2024 is taking place from June 16 to June 21. |
| Outcome: | the system demonstration track at NAACL 2024 received 48 submissions this year . the acceptance rate was 43 . |
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 5: Tutorial Abstracts) (2024.naacl-tutorials)
Copied to clipboard
| Challenge: | NAACL 2024 tutorial sessions are a cornerstone event of the conference . a total of 27 tutorial submissions were received, and 6 were selected for presentation . |
| Approach: | NAACL 2024 will host a tutorial session featuring top-notch researchers . the tutorials aim to equip attendees with the latest tools and methodologies . a total of 27 tutorial submissions were received, and 6 were selected for presentation . |
| Outcome: | the tutorial sessions are a cornerstone event of the conference . the call, submission, reviewing, and selection of tutorials were coordinated . a total of 27 tutorial submissions were received, and 6 were selected for presentation . |
Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 5: Tutorial Abstracts) (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | NAACL 2025 tutorial sessions are a cornerstone event of the conference . tutorials are designed to equip you with the latest insights, tools, and methodologies . |
| Approach: | NAACL 2025 will host a tutorial session on computational linguistics and natural language processing . the tutorials are a cornerstone event of the conference . |
| Outcome: | the tutorial sessions at NAACL 2025 are a cornerstone event of the conference . each submission received a thorough evaluation by a panel of two to three reviewers . |
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification. (2022.lrec-1)
Copied to clipboard
Rémi Uro, David Doukhan, Albert Rilliard, Laetitia Larcher, Anissa-Claire Adgharouamane, Marie Tahon, Antoine Laurent
| Challenge: | Existing methods for creating diachronic corpus of voices are based on speaker characteristics and require human intervention. |
| Approach: | They propose to use a semi-automatic pipeline to create a diachronic corpus of voices balanced for speaker’s age, gender and recording period, according to 32 categories. |
| Outcome: | The proposed method cut down on manual annotations by ten and provides high quality speech for most of the selected excerpts. |