Papers with DiVA
Distilling an End-to-End Voice Assistant Without Instruction Training Data (2025.acl-long)
Copied to clipboard
| Challenge: | Recent efforts to train speech-only LLMs have led to models “forging” speech information from text-only models. |
| Approach: | They propose a paradigm for training Speech Large Language Models without instruction data by using the response of a text-only LLM to transcripts as self-supervision. |
| Outcome: | The proposed model generalizes to Spoken Question Answering, Classification, and Translation and achieves a 72% win rate compared with state-of-the-art models like Qwen 2 Audio . |