Challenge: mTVR is a multilingual video moment retrieval dataset with 218K queries in English and Chinese . Various datasets have been proposed or adapted for the task, but they are all created for a single language (English).
Approach: They propose a multilingual video moment retrieval dataset with 218K queries from 21.8K TV show video clips.
Outcome: The proposed model outperforms strong monolingual baselines while using fewer parameters.

Similar Papers

Cross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval (2022.findings-naacl)

Copied to clipboard

Challenge: Existing multilingual video corpus moment retrieval methods are based on a two-stream structure.
Approach: They propose a multilingual video corpus moment retrieval task that uses a two-stream structure to generate a query-visual similarity and a subtitle stream exploits the query-subtitle similarity.
Outcome: The proposed method improves accuracy on a large-scale video corpus moment retrieval dataset.
Modal-specific Pseudo Query Generation for Video Corpus Moment Retrieval (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown promising results in video corpus moment retrieval . however, they relied on the expensive query annotations for the VCMR .
Approach: They propose a self-supervised learning framework to localize video corpus moment without annotations.
Outcome: The proposed framework can localize the video corpus moment without any explicit annotation on TVR dataset.
MMAR: Multilingual and Multimodal Anaphora Resolution in Instructional Videos (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multilingual anaphora resolution include images and video inputs.
Approach: They propose to include multimodal information in the form of images in anaphora resolution tasks.
Outcome: The proposed approach improves resolution by 10% for unseen languages.
MM-AVS: A Full-Scale Dataset for Multi-modal Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Multimodal summarization materials lacking a holistic organization by integrating resources from various modalities.
Approach: They propose a multimodal article and video summarization dataset that integrates resources from different modalities.
Outcome: The proposed dataset validates the important assistance role of external information for multimodal summarization.
Large Scale Multi-Lingual Multi-Modal Summarization Dataset (2023.eacl-main)

Copied to clipboard

Challenge: a large dataset of document-image pairs and annotated multi-modal summarization data is needed for multi-lingual modeling . encoder-decoder models represent information comprising multiple modalities.
Approach: They propose to use a multi-lingual summarization dataset to analyze multi-modal summarizing using multi-linguistic annotated data.
Outcome: The proposed dataset is the largest multi-lingual multi-modal summarization dataset for 13 languages and consists of cross-lingual summarizing data for 2 languages.
MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction (2023.acl-long)

Copied to clipboard

Challenge: Natural language video localization (NLVL) aims to localize a temporal moment from an untrimmed video that semantically corresponds to a given text query.
Approach: They propose a proposal-based solution that generates proposals and selects the best matching proposal.
Outcome: The proposed solution is faster than existing approaches on three public datasets.
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in embedding resources have led to a lack of representation of the Dutch language in multilingual resources.
Approach: They introduce Massive Text Embedding Benchmark for Dutch (MTEB-NL) which includes existing Dutch datasets and newly created ones, covering a wide range of tasks.
Outcome: The proposed models demonstrate strong performance across multiple tasks.
Cross-Lingual Phrase Retrieval (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to cross-lingual phrase retrieval learn word or sentence representations in word or sentences.
Approach: They propose a cross-lingual phrase retrieval model that extracts phrase representations from unlabeled example sentences.
Outcome: The proposed model outperforms state-of-the-art methods on a large-scale cross-lingual phrase retrieval dataset, showing it can perform in an unseen language pair during training.
ViLL-E: Video LLM Embeddings for Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Video Large Language Models excel at video understanding tasks where outputs are textual . however, they underperform specialized embedding-based models in Retrieval tasks .
Approach: They propose a video-LLM-based model with an embedding generation mechanism that allows the model to "think longer" for complex videos and stop early for easy ones.
Outcome: The proposed model outperforms specialized embedding-based models in video understanding tasks while remaining competitive on VideoQA tasks.
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations