Papers by Jiachen Lian

2 papers
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: AVLM integrates full-face visual cues into a pre-trained expressive speech model.
Approach: They propose an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model.
Outcome: The proposed model incorporates full-face visual cues into a pre-trained expressive speech model.
Towards Hierarchical Spoken Language Disfluency Modeling (2024.eacl-long)

Copied to clipboard

Challenge: Existing solutions to speech dysfluency modeling are limited and expensive for low-income families.
Approach: They propose a hierarchical unconstrained dysfluency modeling approach that addresses both dysfluencies transcription and detection to eliminate the need for extensive manual annotation.
Outcome: The proposed approach eliminates the need for extensive manual annotation and improves the accuracy of the proposed model in phonetic transcription.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations