Papers with Regularization
AlignCap: Aligning Speech Emotion Captioning to Human Preferences (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for speech emotion capture often produce hallucinations and lose generalization on unseen speech. |
| Approach: | They propose to align speech emotion captioning to human preference based on large language model (LLM) and human preference regularization to eliminate factuality and faithfulness hallucinations. |
| Outcome: | Experiments show that AlignCap performs better than existing methods on Zero-shot SEC task. |