Papers by Shuichiro Shimizu
When Large Language Models Meet Speech: A Survey on Integration Approaches (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models have spurred interest in expanding their application beyond text-based tasks. |
| Approach: | They propose to categorize the integration of speech with LLMs into three main approaches . they demonstrate how these methods are applied across various speech-related applications . |
| Outcome: | The proposed methods are applied across speech-related applications and highlight the challenges in this field to offer inspiration for future research. |
VISA: An Ambiguous Subtitles Dataset for Visual Scene-aware Machine Translation (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing multimodal machine translation datasets contain images and video captions or general subtitles which rarely contain linguistic ambiguity. |
| Approach: | They propose a dataset that consists of Japanese-English parallel sentence pairs and corresponding video clips. |
| Outcome: | The proposed dataset is challenging for the latest MMT system and can facilitate MMT research. |
MELD-ST: An Emotion-aware Speech Translation Dataset (2024.findings-acl)
Copied to clipboard
Sirou Chen, Sakiko Yahata, Shuichiro Shimizu, Zhengdong Yang, Yihang Li, Chenhui Chu, Sadao Kurohashi
| Challenge: | Emotion plays a crucial role in human conversation. |
| Approach: | They present a MELD-ST dataset for the emotion-aware speech translation task . they show that fine-tuning with emotion labels can enhance translation performance . |
| Outcome: | The proposed dataset shows that fine tuning with emotion labels can improve translation performance in some settings. |
Video-Helpful Multimodal Machine Translation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing multimodal machine translation datasets contain images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity. |
| Approach: | They propose an MMT dataset that contains ambiguous subtitles and a video-helpful evaluation set. |
| Outcome: | The proposed model performs significantly better than existing models on ambiguous subtitles dataset . it is based on a training set and video-helpful evaluation set . |
Towards Speech Dialogue Translation Mediating Speakers of Different Languages (2023.findings-acl)
Copied to clipboard
| Challenge: | a new task is proposed to mediate speakers of different languages using speech dialogue translation . we consider context as an important aspect that needs to be addressed in this task . speech translation (ST) has also recently shown success in monologue translation - but no study has focused on ST of dialogues . |
| Approach: | They propose a task to mediate speakers of different languages using speech dialogue translation . they construct a speechBSD dataset and conduct baseline experiments . |
| Outcome: | The proposed task mediates speakers of different languages using speech dialogue translation dataset . it shows that bilingual context performs better in our settings . |