Papers with speech
AB/BA analysis: A framework for estimating keyword spotting recall improvement while maintaining audio privacy (2022.naacl-industry)
Copied to clipboard
| Challenge: | Keyword spotting systems that detect keywords in speech are difficult to evaluate under privacy constraints. |
| Approach: | They propose to use offline decoding to evaluate a candidate KWS model against a baseline model without requiring negative examples. |
| Outcome: | The proposed method improves the time, privacy, and cost of the evaluation and compares with real data. |
Growing Trees on Sounds: Assessing Strategies for End-to-End Dependency Parsing of Speech (2024.acl-short)
Copied to clipboard
| Challenge: | Direct dependency parsing of the speech signal is proposed as a way of incorporating prosodic information into the parser and bypassing the limitations of a pipeline approach. |
| Approach: | They propose to use graph-based parsing and sequence labeling based parses to integrate prosodic information into the parser and bypass limitations of pipeline approaches. |
| Outcome: | The proposed graph based approach outperforms a pipeline approach on a large treebank of spoken french, despite having 30% fewer parameters. |
Learning When to Translate for Streaming Speech (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods waiting-and-translating for a fixed duration break speech acoustic units . Existing models waiting-for a set duration and generating partial sentences are not effective . |
| Approach: | They propose a monotonic segmentation module inside an encoder-decoder model to detect proper speech unit boundaries for a streaming speech input. |
| Outcome: | The proposed method outperforms existing methods on a speech translation dataset and achieves the best trade-off between translation quality and latency. |
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have demonstrated that Large Language Models possess a form of emotional intelligence, capable of interpreting emotional stimuli in text. |
| Approach: | They propose a method that translates speech characteristics into natural language descriptions and integrates them into LLMs to perform multimodal emotion analysis via text prompts. |
| Outcome: | The proposed method outperforms baseline models that require structural modifications on two datasets showing significant improvements in emotion recognition accuracy. |
When to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants (2023.acl-short)
Copied to clipboard
| Challenge: | Existing models focus on improving the efficiency of self-attention, but in practice they may be slower, especially given modest input lengths that are typical of many tasks. |
| Approach: | They propose a novel local-attention variant of a self-supervised speech model that uses input length thresholds to identify bottlenecks. |
| Outcome: | The proposed model is based on a self-attention-based model with a high input length threshold. |
Grammatical error detection in transcriptions of spoken English (2020.coling-main)
Copied to clipboard
| Challenge: | CrowdED corpus of spoken English monologues on business topics was crowdsourced from native speakers of English and learners of English with German as their first language. |
| Approach: | They propose to use the corpus recordings to correct existing speech transcriptions and edit them to make them more fluent. |
| Outcome: | The proposed transcription corrections and annotations can be used for automatic transcription post-editing and grammatical error correction for spoken English. |
Fluent Translations from Disfluent Speech in End-to-End Speech Translation (N19-1)
Copied to clipboard
| Challenge: | Disfluency removal is an intermediate step between speech recognition and machine translation (MT) with the rise of end-to-end speech translation systems, disfluency recognition and removal needs to be incorporated into the model architectures or handled as a post-processing step. |
| Approach: | They propose to use a sequence-to-sequence model to translate from noisy, disfluent speech to fluent text with disfluencies removed using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset. |
| Outcome: | The proposed model generates fluent translations from disfluent speech using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset. |
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs (2024.naacl-long)
Copied to clipboard
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Ke Li, Junteng Jia, Yuan Shangguan, Jay Mahadeokar, Ozlem Kalinli, Christian Fuegen, Mike Seltzer
| Challenge: | a new model for speech processing and reasoning uses curated data instead of text. |
| Approach: | They extend the instruction-tuned Llama-2 model with end-to-end speech processing and reasoning abilities without using any carefully curated paired data. |
| Outcome: | The proposed model outperforms or outperfects existing models on synthesized and recorded speech QA tests. |
Continual Reinforcement Learning for Controlled Text Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Controlled Text Generation (CTG) aims to steer text generation towards texts possessing a desired attribute. |
| Approach: | They propose an algorithm that steers the generation of continuations of a given context . they propose a Continual Learning problem to learn at every step to steer next-word generation . |
| Outcome: | The proposed algorithm is based on a plug-and-play language model and exhibits promising results. |
Evaluating the Expressive Appropriateness of Speech in Rich Contexts (2026.acl-long)
Copied to clipboard
Tianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang
| Challenge: | Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level. |
| Approach: | They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context. |
| Outcome: | The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set. |
Grounding Task Assistance with Multimodal Cues from a Single Demonstration (2025.findings-acl)
Copied to clipboard
Gabriel Herbert Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi, Vibhav Vineet, Andrew D Wilson
| Challenge: | RGB video often fails to capture fine-grained contextual cues such as intent, safety-critical environmental factors, and subtle preferences embedded in human behavior. |
| Approach: | They propose a framework that integrates eye gaze and speech cues to improve conversational agents for task assistance by integrating eye gaze with speech cuests. |
| Outcome: | The proposed framework captures fine-grained intent and user-specific cues, enabling richer contextual grounding for visual question answering. |
GIL-GALaD: Gender Inclusive Language - German Auto-Assembled Large Database (2024.lrec-main)
Copied to clipboard
| Challenge: | grammatically gendered languages such as German pose unique challenges in generating gender-inclusive language for corrective model training or fine-tuning. |
| Approach: | a corpus of German gender-inclusive language is assembled to help improve model training . grammatically gendered languages such as german pose unique challenges . authors describe most common strategies for gender- inclusive language in german . |
| Outcome: | a corpus of German gender-inclusive language is assembled and will be included in the release. |
Bridging the Granularity Gap for Acoustic Modeling (2023.findings-acl)
Copied to clipboard
Chen Xu, Yuhao Zhang, Chengbo Jiao, Xiaoqian Liu, Chi Hu, Xin Zeng, Tong Xiao, Anxiang Ma, Huizhen Wang, Jingbo Zhu
| Challenge: | Despite the success of speech recognition, how to encode the speech features effectively remains an open problem. |
| Approach: | They propose a Progressive Down-Sampling technique which compresses acoustic features into coarser-grained units containing more complete semantic information, like text-level representation. |
| Outcome: | The proposed method yields comparable or better results on the speech recognition task and inference speedups ranging from 1.20x to 1.47x. |