Papers by Huadai Liu

7 papers
Masked Text-to-Audio Flow-Matching and Reward Feedback Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Experimental results show that Flow-matching generative models can scale training by increasing data, computational resources, and model size.
Approach: They propose a flow-matching transformer with masked generative modeling for scaling text-to-audio inference-time prediction.
Outcome: The proposed model scales inference-time computations by masking generation and re-predicting them through iterative decoding.
RMSSinger: Realistic-Music-Score based Singing Voice Synthesis (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for singing voice synthesis are limited to fine-grained music scores . manual adjustment destroys regularity of note durations, making fine-grain music scores "crushed"
Approach: They propose a method to synthesize singing voices given realistic music scores . they use real-music-score-based Singing Voice Synthesis to generate high-quality voices .
Outcome: The proposed method eliminates manual annotation and simplifies phoneme-level mel-note alignment.
FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment.
Approach: They propose to learn straight flow for fast simulation by using flashAudio with rectified flows and immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment.
Outcome: The proposed method can learn straight flow for fast simulations and reduce noise distribution.
AntCritic: Argument Mining for Free-Form and Visually-Rich Financial Comments (2024.lrec-main)

Copied to clipboard

Challenge: Argument mining is a thriving task in natural language processing, but its generalization is limited by existing datasets.
Approach: They propose to use a dataset to help model argument mining . the dataset AntCritic supports both argument component detection and argument relation prediction tasks.
Outcome: The proposed model can detect arguments and identify their relationships automatically.
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer (2023.emnlp-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) performance has improved with the advent of denoising Diffusion Probabilistic Models . however, perceived quality of audio depends on content, pitch, rhythm, and energy .
Approach: They propose a visual TTS model with scalable diffusion transformers that complement phoneme sequences with visual information to generate high-perceived audio.
Outcome: The proposed model outperforms existing models regardless of visibility of the scene . it can generate high-perceived audio, opening up new avenues for AR and VR applications .
AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation (2023.acl-long)

Copied to clipboard

Challenge: Existing models for speech-to-speech translation suffer from distinct degradation in noisy environments and fail to translate visual speech.
Approach: They propose a text-based audio-visual speech-to-speech translation model that integrates visual information with audio-only data to improve system robustness.
Outcome: The proposed model outperforms models trained on audio-only corpus in two languages . it also improves with low-resource audio-visual data, compared with baselines .
Wav2SQL: Direct Generalizable Speech-To-SQL Parsing (2024.findings-acl)

Copied to clipboard

Challenge: Existing models for speech-driven SQL parsing are based on a cascaded approach, resulting in data scarcity and inconsistent performance.
Approach: They propose a direct generalizable speech-to-SQL parsing model which avoids error compounding across cascaded systems.
Outcome: The proposed model avoids error compounding and achieves state-of-the-art results by 4.7% improvement over baseline.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations