BigVideo: A Large-scale Video Subtitle Translation Dataset for Multimodal Machine Translation (2023.findings-acl)
Copied to clipboard
Liyan Kang, Luyang Huang, Ningxin Peng, Peihao Zhu, Zewei Sun, Shanbo Cheng, Mingxuan Wang, Degen Huang, Jinsong Su
| Challenge: | Existing datasets focus on captions describing images or videos, which are not large and diverse enough. |
| Approach: | They propose a large-scale video subtitle translation dataset to facilitate multi-modality machine translation. |
| Outcome: | The proposed dataset is 10 times larger than the widely used *How2* and *VaTeX* datasets. |
Similar Papers
Video-Helpful Multimodal Machine Translation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing multimodal machine translation datasets contain images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity. |
| Approach: | They propose an MMT dataset that contains ambiguous subtitles and a video-helpful evaluation set. |
| Outcome: | The proposed model performs significantly better than existing models on ambiguous subtitles dataset . it is based on a training set and video-helpful evaluation set . |
Large Scale Multi-Lingual Multi-Modal Summarization Dataset (2023.eacl-main)
Copied to clipboard
| Challenge: | a large dataset of document-image pairs and annotated multi-modal summarization data is needed for multi-lingual modeling . encoder-decoder models represent information comprising multiple modalities. |
| Approach: | They propose to use a multi-lingual summarization dataset to analyze multi-modal summarizing using multi-linguistic annotated data. |
| Outcome: | The proposed dataset is the largest multi-lingual multi-modal summarization dataset for 13 languages and consists of cross-lingual summarizing data for 2 languages. |
VISA: An Ambiguous Subtitles Dataset for Visual Scene-aware Machine Translation (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing multimodal machine translation datasets contain images and video captions or general subtitles which rarely contain linguistic ambiguity. |
| Approach: | They propose a dataset that consists of Japanese-English parallel sentence pairs and corresponding video clips. |
| Outcome: | The proposed dataset is challenging for the latest MMT system and can facilitate MMT research. |
Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation (2023.findings-acl)
Copied to clipboard
| Challenge: | Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision. |
| Approach: | They propose a framework for multimodal machine translation that utilizes large-scale non-triple data and a multimodal translation dataset. |
| Outcome: | The proposed method can significantly improve translation performance with more non-triple data. |
3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset (2024.lrec-main)
Copied to clipboard
Xinyu Ma, Xuebo Liu, Derek F. Wong, Jun Rao, Bei Li, Liang Ding, Lidia S. Chao, Dacheng Tao, Min Zhang
| Challenge: | Existing studies have shown that visual information in existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. |
| Approach: | They propose to use 3AM to create an ambiguity-aware multimodal machine translation dataset. |
| Outcome: | The proposed dataset includes more ambiguity and a greater variety of captions and images than other MMT datasets. |
MultiSubs: A Large-scale Multimodal and Multilingual Dataset (2022.lrec-1)
Copied to clipboard
| Challenge: | a large-scale multimodal and multilingual dataset is used to facilitate research on visual grounding of words to images in their contextual usage in language. |
| Approach: | They propose a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. |
| Outcome: | The proposed dataset will facilitate research on visual grounding of words in their contextual usage in language. |
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)
Copied to clipboard
Pei Fu, Tongkun Guan, Zining Wang, Zhentao Guo, Chen Duan, Hao Sun, Boming Chen, Qianyi Jiang, Jiayao Ma, Kai Zhou, Junfeng Luo
| Challenge: | Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms. |
| Approach: | They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks. |
| Outcome: | The proposed models perform well on mainstream benchmarks and are compared with other models. |
Learning Translations via Images with a Massively Multilingual Image Dataset (P18-1)
Copied to clipboard
| Challenge: | Existing datasets for learning translations of words are limited to a few high-resource languages and unrealistically easy settings. |
| Approach: | They propose a large-scale multilingual corpus of images labeled with the word they represent to facilitate translation research. |
| Outcome: | The proposed method improves on an unsupervised technique that has been limited to a few languages and unrealistic settings. |
MM-AVS: A Full-Scale Dataset for Multi-modal Summarization (2021.naacl-main)
Copied to clipboard
| Challenge: | Multimodal summarization materials lacking a holistic organization by integrating resources from various modalities. |
| Approach: | They propose a multimodal article and video summarization dataset that integrates resources from different modalities. |
| Outcome: | The proposed dataset validates the important assistance role of external information for multimodal summarization. |
Video-guided Machine Translation with Spatial Hierarchical Attention Network (2021.acl-srw)
Copied to clipboard
| Challenge: | Existing studies use pretrained motion detection models as verb sense ambiguity representations to solve the verb sense problem. |
| Approach: | They propose to use video contents as auxiliary information to address the word sense ambiguity problem in machine translation. |
| Outcome: | Experiments on the VATEX dataset show that the proposed system achieves 35.86 BLEU-4 score, which is 0.51 score higher than the single model of the SOTA method. |