Papers by HangChen HangChen
MISP-Meeting: A Real-World Dataset with Multimodal Cues for Long-form Meeting Transcription and Summarization (2025.acl-long)
Copied to clipboard
| Challenge: | Existing systems that can recognize spoken content, extract key information, and produce concise summaries are lacking in meeting transcription and summarization. |
| Approach: | They propose a multimodal dataset that integrates information from speech, vision, and text modalities to facilitate automatic meeting transcription and summarization (AMTS). |
| Outcome: | The proposed dataset reduces the character error rate (CER) by 36.60% to 20.27% and improves speech recognition and large language models. |