Papers by Srikanth Ronanki

6 papers
In Other News: a Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data (N19-2)

Copied to clipboard

Challenge: Recent advances in text-to-speech synthesis have enabled researchers to generate high-quality speech with a wide range of prosodic variations.
Approach: They propose a model that can synthesise newscaster-style speech with a few hours of data . they propose to factor in contextual word embeddings and evaluate it against neutral synthesis .
Outcome: The proposed model can synthesise newscaster-style speech with just a few hours of data.
AdaBERT-CTC: Leveraging BERT-CTC for Text-Only Domain Adaptation in ASR (2023.emnlp-industry)

Copied to clipboard

Challenge: End-to-end (E2E) automatic speech recognition models struggle to recognize out-of-domain words such as proper nouns and domain-specific terms.
Approach: They propose a domain adaptation technique that relies solely on textual data to adapt to out-of-domain words.
Outcome: The proposed method outperforms the base model by up to 14% relative word error rate improvement on several out-of-domain, publicly available datasets.
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) often fail to scale their performance on long-context tasks performance in line with the context lengths they support.
Approach: They propose a model-agnostic mitigation strategy that transforms a long-context task into a short-concept one by prompting the model to recite the retrieved evidence before attempting to solve the problem.
Outcome: The proposed model improves on a long-context task up to 4% on RULER.
LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: a novel linearization framework is proposed to reduce the cost of training transformers from scratch.
Approach: They propose a linear attention framework that integrates pre-trained transformers into a performant linear attention architecture.
Outcome: The proposed framework improves performance on mistral-7B with 1K-length sequences and BABILong benchmarks.
Retrieve and Copy: Scaling ASR Personalization to Large Catalogs (2023.emnlp-industry)

Copied to clipboard

Challenge: End-to-end ASR models struggle to recognize uncommon domain-specific words due to limited audio context.
Approach: They propose a "Retrieve and Copy" mechanism to improve latency while retaining the accuracy even when scaled to a large catalog.
Outcome: The proposed method achieves 6% more word error rate reduction and 3.6% improvement in F1 when scaled to a large catalog size while retaining the accuracy.
SpeechGuard: Exploring the Adversarial Robustness of Multi-modal Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Integrated Speech and Large Language Models (SLMs) that follow speech instructions and generate relevant text responses have gained popularity lately.
Approach: They propose algorithms that can generate adversarial examples to jailbreak SLMs without human involvement.
Outcome: The proposed algorithms achieve state-of-the-art on spoken question-answering task scoring over 80% on both safety and helpfulness metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations