Challenge: Existing speech tokenization models lack contextual representations for speech synthesis . absence of contextual representation results in elevated WER and WIL scores .
Approach: They propose a language model-guided distillation method that incorporates contextual information into a comprehensive speech tokenizer.
Outcome: The proposed method outperforms state-of-the-art tokenization models in reducing WER and WIL scores.

Similar Papers

LLM-Codec: Neural Audio Codec Meets Language Model Objectives (2026.findings-acl)

Copied to clipboard

Challenge: Neural audio codecs are optimized for waveform reconstruction rather than autoregressive prediction.
Approach: They propose to augment codec training with language-model-facing objectives while keeping both codec and LLM architectures unchanged.
Outcome: The proposed model improves speech coherence and predictability by preserving the semantic alignment between audio and text representations.
Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated significant strides in generating high-quality speech . discretizing speech by neural audio codecs often results in sequences that differ from text sequences .
Approach: They quantitatively analyze the Discrete Representation Inconsistency phenomenon within popular audio tokenizers such as EnCodec.
Outcome: The proposed method mitigates the DRI phenomenon within popular audio tokenizers such as EnCodec.
RepCodec: A Speech Representation Codec for Speech Tokenization (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have led to discrete speech tokenization, but this discretization can be costly and impedes performance.
Approach: They propose a new speech representation codec for semantic speech tokenization that reconstructs speech representations from speech encoders like HuBERT or data2vec.
Outcome: The proposed method outperforms the widely used k-means clustering approach in speech understanding and generation.
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization (2026.acl-long)

Copied to clipboard

Challenge: Existing speech codecs struggle to balance high-quality reconstruction with semantically rich representations, limiting their effectiveness in both generative and understanding tasks.
Approach: They propose a neural speech codec with semantic-acoustic dual-stream quantization that disentangles semantic and acousian modeling into two dedicated streams.
Outcome: The proposed codec outperforms state-of-the-art speech tokenizers in auto-propagating text-to-speech models.
Enhancing Cross-Tokenizer Knowledge Distillation with Contextual Dynamical Mapping (2025.findings-acl)

Copied to clipboard

Challenge: Knowledge distillation (KD) approaches focus on homogeneous architectures with identical tokenizers, constraining their applicability in cross-architecture scenarios.
Approach: They propose a framework that uses contextual information to enhance sequence alignment precision and dynamically improves vocabulary mapping.
Outcome: The proposed framework shows significant advantages over existing methods for model compression . it can be used across multiple model families and across multiple benchmarks .
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs (2026.acl-long)

Copied to clipboard

Challenge: Existing speech codecs struggle to balance these objectives at low bitrates . XY-Tokenizer achieves stronger semantic alignment than representative semantic-distillation codec .
Approach: They propose a low-bitrate speech codec that aligns discrete speech representations with text while preserving fine-grained acoustic details for reconstruction.
Outcome: The proposed codec outperforms existing low-bitrate speech codecs in speech understanding and generation tasks.
Towards Codec-LM Co-design for Neural Codec Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Neural codec language models (or codec LMs) are emerging as a powerful framework for text-to-speech (TTS) despite the close interdependence of codecs and LM, research on codec and lms has largely remained siloed.
Approach: They propose a frame-wise codec encoder that improves both LM log-likelihood and TTS metrics . they also propose LM codebook level dropout to efficiently navigate a portion of codec-LM design space .
Outcome: The proposed codec-LM co-design improves intelligibility, audio quality and speaker control compared to a siloed baseline.
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing gaps between discrete acoustic codecs and downstream speech language models . initial channel of codebooks contains excessive information, making it difficult to generate tokens from weakly supervised signals such as text.
Approach: They propose a discrete acoustic codec for generating acustic tokens from weakly supervised signals.
Outcome: The proposed language-codec outperforms competing audio compression algorithms and validates on downstream speech language models.
Contextualization Distillation from Large Language Model for Knowledge Graph Completion (2024.findings-eacl)

Copied to clipboard

Challenge: Existing knowledge graph completion models lack textual information, which limits their performance . a plug-in-and-play approach is needed to train small models in descriptive context .
Approach: They propose a plug-in-and-play approach to knowledge graph completion that prompts LLMs to generate descriptive context.
Outcome: The proposed method improves performance on Wikipedia articles and synset definitions.
SelFusion: Self-distillation for Diffusion Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing knowledge distillation methods for autoregressive large language models (LLMs) are not effective for reducing generation quality, but they can be useful for real-time applications.
Approach: They propose a self-distillation framework that allows for effective KD without external teacher . they propose to use two modes of knowledge distillation to determine distillation direction .
Outcome: The proposed framework outperforms existing methods with external teachers on instruction-following tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations