Challenge: In spoken dialogue, even if two current turns are the same sentence, their responses might differ when they are spoken in different styles.
Approach: They propose a language-to-speech dataset that can model linguistic content and speaking styles.
Outcome: The proposed framework outperforms text-only baselines and prior speech LLMs methods.

Similar Papers

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Unlike textonly large language models (LLMs), SLMs integrate audio encoders and vocoders to support end-to-end speech understanding and generation.
Approach: They evaluate three proprietary and two open-source SLMs and show that none of them can maintain a consistent speaking style when instructed to do so.
Outcome: The proposed models cannot maintain a consistent speaking style after several turns of interaction, but can recall the style instruction when prompted in later turns, but fail to express it.
LLMs for Low Resource Languages in Multilingual, Multimodal and Dialectal Settings (2024.eacl-tutorials)

Copied to clipboard

Challenge: Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting .
Approach: They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models .
Outcome: The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects.
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding (2024.acl-long)

Copied to clipboard

Challenge: Existing large language models struggle to capture some language styles without fine-tuning.
Approach: They propose to meta-trained LLMs based on representative lexicons to recognize new styles they have not been fine-tuned on.
Outcome: The proposed method improves zero-shot transfer across styles on 13 established and 63 novel tasks generated with LLMs.
C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations (2025.emnlp-main)

Copied to clipboard

Challenge: Recent developments in spoken dialogue models have created a gap in understanding their effectiveness in comprehending and emulating human conversations.
Approach: They present a benchmark dataset which comprises 1,079 instances in English and Chinese to examine their effectiveness in emulating human conversations.
Outcome: The proposed model outperforms existing models in English and Chinese by using an LLM-based evaluation method that closely aligns with human judgment.
When Large Language Models Meet Speech: A Survey on Integration Approaches (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have spurred interest in expanding their application beyond text-based tasks.
Approach: They propose to categorize the integration of speech with LLMs into three main approaches . they demonstrate how these methods are applied across various speech-related applications .
Outcome: The proposed methods are applied across speech-related applications and highlight the challenges in this field to offer inspiration for future research.
Comparing Styles across Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Communication practices vary across cultures. Inherent differences in how people think and behave influence cultural norms.
Approach: They propose a framework to extract stylistic differences from multilingual language models (LMs) they use a multilingual lexica to consolidate feature importances into comparable lexical categories .
Outcome: The proposed framework generates comprehensive style lexica in any language and consolidates feature importances from LMs into comparable lexical categories.
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has sought to use large language models to simulate human-human and human-LLM interactions.
Approach: They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts.
Outcome: The proposed models perform similarly in simulating English, Chinese, and Russian dialogues.
Learning LLM Preference over Intra-Dialogue Pairs: A Framework for Utterance-level Understandings (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities in handling complex dialogue tasks without requiring use case-specific fine-tuning.
Approach: They propose a framework that combines the scalability of LLM-generated labels with the precision of human annotations to achieve higher speed and accuracy comparable to larger models.
Outcome: The proposed framework significantly improves accuracy across utterance-level dialogue tasks, including sentiment detection (over 2%), dialogue act classification (over 1.5%), etc.
Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) can simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (N1) knowledge.
Approach: They use large language models to simulate non-native-like English use observed in human second language (L2) learners, and then compare their outputs to real L2 learner data.
Outcome: The proposed models replicate L1-dependent patterns observed in human second language (L2) learners, with distinct influences from various languages.
I Learn Better If You Speak My Language: Understanding the Superior Performance of Fine-Tuning Large Language Models with LLM-Generated Responses (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has demonstrated that a large language model (LLM) can generate training data for another LLM, or for creating supplementary training materials, such as rationales.
Approach: They conduct an in-depth investigation to understand why fine-tuning an LLM with responses generated by a LLM often yields better results than using responses generated from humans.
Outcome: The proposed approach can be used to transfer knowledge from a larger model to a smaller one, or for creating supplementary training materials, such as rationales.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations