Papers by Jiachen Lian
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | AVLM integrates full-face visual cues into a pre-trained expressive speech model. |
| Approach: | They propose an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. |
| Outcome: | The proposed model incorporates full-face visual cues into a pre-trained expressive speech model. |
Towards Hierarchical Spoken Language Disfluency Modeling (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing solutions to speech dysfluency modeling are limited and expensive for low-income families. |
| Approach: | They propose a hierarchical unconstrained dysfluency modeling approach that addresses both dysfluencies transcription and detection to eliminate the need for extensive manual annotation. |
| Outcome: | The proposed approach eliminates the need for extensive manual annotation and improves the accuracy of the proposed model in phonetic transcription. |