| Challenge: | mTVR is a multilingual video moment retrieval dataset with 218K queries in English and Chinese . Various datasets have been proposed or adapted for the task, but they are all created for a single language (English). |
| Approach: | They propose a multilingual video moment retrieval dataset with 218K queries from 21.8K TV show video clips. |
| Outcome: | The proposed model outperforms strong monolingual baselines while using fewer parameters. |
Similar Papers
Cross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing multilingual video corpus moment retrieval methods are based on a two-stream structure. |
| Approach: | They propose a multilingual video corpus moment retrieval task that uses a two-stream structure to generate a query-visual similarity and a subtitle stream exploits the query-subtitle similarity. |
| Outcome: | The proposed method improves accuracy on a large-scale video corpus moment retrieval dataset. |
Modal-specific Pseudo Query Generation for Video Corpus Moment Retrieval (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown promising results in video corpus moment retrieval . however, they relied on the expensive query annotations for the VCMR . |
| Approach: | They propose a self-supervised learning framework to localize video corpus moment without annotations. |
| Outcome: | The proposed framework can localize the video corpus moment without any explicit annotation on TVR dataset. |
MMAR: Multilingual and Multimodal Anaphora Resolution in Instructional Videos (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to multilingual anaphora resolution include images and video inputs. |
| Approach: | They propose to include multimodal information in the form of images in anaphora resolution tasks. |
| Outcome: | The proposed approach improves resolution by 10% for unseen languages. |
MM-AVS: A Full-Scale Dataset for Multi-modal Summarization (2021.naacl-main)
Copied to clipboard
| Challenge: | Multimodal summarization materials lacking a holistic organization by integrating resources from various modalities. |
| Approach: | They propose a multimodal article and video summarization dataset that integrates resources from different modalities. |
| Outcome: | The proposed dataset validates the important assistance role of external information for multimodal summarization. |
Large Scale Multi-Lingual Multi-Modal Summarization Dataset (2023.eacl-main)
Copied to clipboard
| Challenge: | a large dataset of document-image pairs and annotated multi-modal summarization data is needed for multi-lingual modeling . encoder-decoder models represent information comprising multiple modalities. |
| Approach: | They propose to use a multi-lingual summarization dataset to analyze multi-modal summarizing using multi-linguistic annotated data. |
| Outcome: | The proposed dataset is the largest multi-lingual multi-modal summarization dataset for 13 languages and consists of cross-lingual summarizing data for 2 languages. |
MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction (2023.acl-long)
Copied to clipboard
| Challenge: | Natural language video localization (NLVL) aims to localize a temporal moment from an untrimmed video that semantically corresponds to a given text query. |
| Approach: | They propose a proposal-based solution that generates proposals and selects the best matching proposal. |
| Outcome: | The proposed solution is faster than existing approaches on three public datasets. |
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in embedding resources have led to a lack of representation of the Dutch language in multilingual resources. |
| Approach: | They introduce Massive Text Embedding Benchmark for Dutch (MTEB-NL) which includes existing Dutch datasets and newly created ones, covering a wide range of tasks. |
| Outcome: | The proposed models demonstrate strong performance across multiple tasks. |
Cross-Lingual Phrase Retrieval (2022.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to cross-lingual phrase retrieval learn word or sentence representations in word or sentences. |
| Approach: | They propose a cross-lingual phrase retrieval model that extracts phrase representations from unlabeled example sentences. |
| Outcome: | The proposed model outperforms state-of-the-art methods on a large-scale cross-lingual phrase retrieval dataset, showing it can perform in an unseen language pair during training. |
ViLL-E: Video LLM Embeddings for Retrieval (2026.acl-long)
Copied to clipboard
| Challenge: | Video Large Language Models excel at video understanding tasks where outputs are textual . however, they underperform specialized embedding-based models in Retrieval tasks . |
| Approach: | They propose a video-LLM-based model with an embedding generation mechanism that allows the model to "think longer" for complex videos and stop early for easy ones. |
| Outcome: | The proposed model outperforms specialized embedding-based models in video understanding tasks while remaining competitive on VideoQA tasks. |
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)
Copied to clipboard
| Challenge: | BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese . |
| Approach: | They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages . |
| Outcome: | The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese . |