Papers by Ha-Yeong Choi
Hierarchical Representation Alignment Learning of Diffusion Transformers for Neural Audio Codec (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in diffusion and conditional flow matching models for low-resolution domains are underexplored. |
| Approach: | They propose a CFM-based model that iteratively generates raw waveform in low-bitrate conditions . they propose DVQ, a factorized quantization method that uses a single quantizer . |
| Outcome: | The proposed model outperforms state-of-the-art neural audio codecs in audio quality and semantic intelligibility under low-bitrate conditions. |