Papers by Suyoun Kim

5 papers
Acoustic-to-Word Models with Conversational Context Information (N19-1)

Copied to clipboard

Challenge: Existing speech recognition models are built at a sentence level, and therefore it may not capture conversational context information.
Approach: They propose a direct acoustic-to-word, end-to end speech recognition model that integrates a conversational context with other available information and directly recognizes words from speech.
Outcome: The proposed model outperforms a standard end-to-end speech recognition system on the Switchboard conversational speech corpus and shows that it is more accurate than existing models.
Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion (P19-1)

Copied to clipboard

Challenge: Existing speech recognition systems are built at individual, isolated utterance level to make building systems computationally feasible.
Approach: They propose to use text-based external word and/or sentence embeddings to integrate conversational context information into a single neural network model.
Outcome: The proposed model outperforms standard end-to-end speech recognition models on the Switchboard conversational speech corpus and improves word error rate with better conversational-context representation.
Introducing Semantics into Speech Encoders (2023.acl-long)

Copied to clipboard

Challenge: Existing self-supervised speech encoders contain primarily acoustic rather than semantic information.
Approach: They propose a task-agnostic unsupervised way to incorporate semantic information from large language model (LLM) systems into self-supervised speech encoders without labeled audio transcriptions.
Outcome: The proposed approach improves spoken language understanding (SLU) performance by over 5% on intent classification (IC), with modest gains in named entity resolution (NER) and slot filling (SF), and spoken question answering (SQA) score by over 22%.
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding (2024.findings-emnlp)

Copied to clipboard

Challenge: End-to-end models for Spoken Language Understanding have been autoregressive, resulting in higher latencies.
Approach: They propose a method that uses Connectionist Temporal Classification to train robust non-autoregressive deliberation models.
Outcome: The proposed method achieves 10x latency reduction over autoregressive models while preserving ability to correct ASR mistranscriptions.
Joint Audio/Text Training for Transformer Rescorer of Streaming Speech Recognition (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that streaming end-to-end speech recognition models suffer from higher word error rates (WER) compared to non-streaming models, streaming endto-ended ASR models are limited to short audio context or not use future context to satisfy low latency constraints.
Approach: They propose a 2nd-pass rescoring model on top of the 1st-pass streaming model to improve recognition accuracy while keeping latency low.
Outcome: The proposed method improves word error rate significantly compared to the existing model without adding any additional parameters or latency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations