Papers by Chaoren Wang
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment (2025.acl-long)
Copied to clipboard
| Challenge: | Existing zero-shot text-to-speech systems struggle in challenging scenarios such as tongue twisters, repeated words, code-switching, and cross-lingual synthesis. |
| Approach: | They propose a dataset that leverages preference alignment techniques to improve performance . they also extend the Direct Preference Optimization framework to accommodate diverse TTS architectures . |
| Outcome: | The proposed dataset improves intelligibility, similarity, and audio quality for multiple models across domains. |
MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora (2026.findings-acl)
Copied to clipboard
Tao Feng, Yuxiang Wang, Yuancheng Wang, Xueyao Zhang, Dekun Chen, Chaoren Wang, Xun Guan, Zhizheng Wu
| Challenge: | Existing approaches to voice imitation use complex model design and a quality ceiling when synthetic speech is used as training *sources*. |
| Approach: | They propose a model that uses synthetic speech as training *sources* while retaining real recordings as *targets*. |
| Outcome: | The proposed model outperforms existing methods in naturalness while maintaining competitive similarity scores across speaker identity, accent, and emotion dimensions. |
Closing the Modality Reasoning Gap for Speech Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in Speech Large Language Models have a modality reasoning gap that is not addressed by prior work. |
| Approach: | They propose a reinforcement-learning framework that aligns text-conditioned and speech-conditioned trajectories through an asymmetric reward design. |
| Outcome: | Experiments on MMSU and OBQA show that the proposed framework narrows the modality reasoning gap and achieves state-of-the-art performance among 7B-scale Speech LLMs. |