Papers with TTM
Kahaani: A Multimodal Co-Creative Storytelling System (2026.eacl-srw)
Copied to clipboard
| Challenge: | Kahaani is a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to address the challenge of sustaining engagement to foster educational narrative experiences. |
| Approach: | They propose a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to help children develop their storytelling skills. |
| Outcome: | The proposed system combines large language models, text-to-speech, and music generation to produce a rich, immersive, and accessible storytelling experience. |
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions (2026.acl-long)
Copied to clipboard
Chunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang, Yuxin Guo, Teng Ma, Ziyu Zhang, Tianrui Wang, Cheng Gong, Yushen Chen, Ruibo Fu, Longbiao Wang, Jianwu Dang
| Challenge: | Generative audio modeling has been fragmented into specialized tasks such as text-to-speech (TTS), text- to-music (TTM), and text-ta (TTA) specialized models require reference audio for timbre cloning and strict phoneme alignment, whereas TTA models generate unstructured textures from open-ended captions. |
| Approach: | They propose a unified flow-matching framework capable of synthesizing speech, music, sound effects . they propose 'token injection mechanism' that projects unstructured environmental sounds into structured temporal latent space . |
| Outcome: | The proposed framework achieves state-of-the-art performance in instruction-based TTS and TTM while maintaining competitive fidelity in TTA. |