Challenge: Existing methods focus on disentangling speakers and content, while others focus on preserving the source's prosody.
Approach: They propose a rhythm-controllable and efficient zero-shot voice conversion model that transforms the source speaker’s timbre into an unseen one while retaining speech content.
Outcome: The proposed model adapts the linguistic content duration to the desired speaking style, facilitating the transfer of the target speaker’s rhythm.

Similar Papers

Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling (2025.acl-long)

Copied to clipboard

Challenge: Expressive zero-shot voice conversion (VC) aims to modify source timbre to match unseen speaker . existing zero- shot VC systems struggle to reproduce paralinguistic information in highly expressive speech .
Approach: They propose a framework for expressive zero-shot voice conversion that uses hybrid content encoding and memory-augmented context-aware timbre modeling.
Outcome: The proposed framework surpasses state-of-the-art VC systems in speech naturalness, speaker similarity, and speaker similarness.
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion (2025.findings-emnlp)

Copied to clipboard

Challenge: Current voice conversion methods struggle in zero-shot cross-lingual settings . authors develop a method that can be used in zero shot cross-linguistic settings despite advances in technology .
Approach: They propose a voice-conversion model that combines discrete speech representations with a non-autoregressive speech decoder.
Outcome: The proposed approach excels in zero-shot cross-lingual settings even for unseen languages and accents.
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding (2025.acl-demo)

Copied to clipboard

Challenge: Experimental evaluations show RT-VC delivers a 13.3% reduction in latency . voice conversion modifies speech to match the timbre of a target speaker while preserving content information.
Approach: They propose a zero-shot real-time voice conversion system that leverages an articulatory feature space to naturally disentangle content and speaker characteristics.
Outcome: The proposed system achieves a CPU latency of 61.4 ms, representing a 13.3% reduction in latency.
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control (2025.acl-long)

Copied to clipboard

Challenge: Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS systems cannot perform speaker-specific voice generation.
Approach: They propose a style control module that captures codec representations corresponding to timbre, content, and style in a discrete decoupling codec space.
Outcome: The proposed system can fully clone the speaker's voice and perform speech-specific adjustment and control functions.
DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec (2026.acl-long)

Copied to clipboard

Challenge: DisCo-Speech is a zero-shot controllable text-to-speech framework . standard codecs entangle timbre and prosody, which hinders independent control in continuation-based LMs.
Approach: They propose a disentangled speech codec and an LM-based generator to solve this problem . they propose fusion and reconstruction that merges content and prosody into unified tokens .
Outcome: DisCo-Speech achieves competitive voice cloning and superior zero-shot prosody control.
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion (2024.acl-long)

Copied to clipboard

Challenge: Existing LM-based VC models require offline conversion from source semantics to acoustic features, limiting their deployment to real-time applications.
Approach: They propose a streaming LM-based model for zero-shot voice conversion that uses a fully causal context-aware LM with a temporal-independent acoustic predictor to facilitate real-time conversion given arbitrary speaker prompts and source speech.
Outcome: The proposed model achieves comparable performance to non-streaming VC systems while maintaining a fully causal context-aware LM with a temporal-independent acoustic predictor.
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in text-to-speech (TTS) models have led to improvements in speaker prosody and voices modeling.
Approach: They propose an efficient zero-shot TTS model that leverages distilled time-varying style diffusion to capture diverse speaker identities and prosodies.
Outcome: The proposed model surpasses state-of-the-art models in both naturalness and similarity while reducing inference speed by 90%.
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech (2024.acl-long)

Copied to clipboard

Challenge: Existing zero-shot text-to-speech systems require a few seconds of unseen speaker voice prompts to generate high-quality voices.
Approach: They propose a zero-shot text-to-speech system based on mobile devices . they use a discrete speech codec to integrate hierarchical information from the codec .
Outcome: The proposed system achieves RTF of 0.09 on a single A100 GPU and has been successfully deployed on mobile devices.
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild (2024.acl-long)

Copied to clipboard

Challenge: VoiceCraft is a token-infilling neural codec language model for speech editing and zero-shot text-to-speech evaluation.
Approach: They introduce a token infilling neural codec language model that performs on speech editing and zero-shot text-to-speech tasks.
Outcome: The proposed model outperforms previous models on speech editing and zero-shot text-to-speech tasks.
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion (2025.findings-emnlp)

Copied to clipboard

Challenge: Traditional voice conversion methods attempt to separate speaker identity and linguistic information into distinct representations, but this method often leads to information loss during training.
Approach: They propose a method that leverages synthetic speech data generated by a pretrained model . synthetic data pairs that share the same linguistic content are used as input-output pairs .
Outcome: The proposed method outperforms state-of-the-art methods in speaker-to-voice conversions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations