Challenge: a new cross-lingual information retrieval system for low-resource languages is available in less-frequently-taught languages . a multilingual system can search for relevant information in a haystack of documents in swahili or Somali . human-driven approaches to this problem are complicated in 'low-resourced' languages aaron sagar: "the key role played by humans in triaging results is complicated"
Approach: They propose an end-to-end cross-lingual information retrieval system for low-resource languages . the system enables English speakers to search foreign language repositories using English queries . it summarizes the retrieved documents in English with respect to a particular information need .
Outcome: The proposed system achieves top performance in the most recent IARPA MATERIAL CLIR+summarization evaluations.

Similar Papers

Weakly Supervised Attentional Model for Low Resource Ad-hoc Cross-lingual Information Retrieval (D19-61)

Copied to clipboard

Challenge: Low resource languages often lack relevance annotations for cross-lingual information retrieval . when available, the training data has limited coverage for possible queries .
Approach: They propose a weakly supervised neural model for Cross-lingual information retrieval from low-resource languages using weak supervision instead of relevance annotations.
Outcome: The proposed model achieves 19 MAP points improvement compared to CNNs and 12 points improvement from machine translation-based CLIR models.
Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages (2025.coling-main)

Copied to clipboard

Challenge: lexical gaps between dialects in cross-lingual information retrieval (CLIR) are caused by orthographic variations and different regional expressions.
Approach: They propose a dataset that consists of seven German dialects extracted from Wikipedia.
Outcome: The proposed dataset consists of seven German dialects extracted from Wikipedia.
ACROSS: An Alignment-based Framework for Low-Resource Many-to-One Cross-Lingual Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies ignore data imbalance in multilingual settings and do not utilize monolingual data.
Approach: They propose a cross-lingual summarization model that aligns cross-linguistic data with high-resource monolingual data via contrastive and consistency loss.
Outcome: The proposed model outperforms baseline models and consistently dominates on 45 language pairs.
Cross-Lingual Learning-to-Rank with Shared Representations (N18-2)

Copied to clipboard

Challenge: Cross-lingual information retrieval (CLIR) is a document retrieval task where the documents are written in a language different from that of the user's query.
Approach: They propose a large-scale dataset derived from Wikipedia to support CLIR research in 25 languages.
Outcome: The proposed model can improve the results of Swahili-English CLIR in Japanese and Japanese.
Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages (2026.findings-eacl)

Copied to clipboard

Challenge: Recent studies suggest that summarization in English may be solved, or even "dead" However, there are no accessible, high-quality summarizing datasets in under-represented languages.
Approach: They propose a method for collecting naturally occurring summaries via front-page teasers, where editors summarize full length articles.
Outcome: The proposed method is suited to varying linguistic resources and is available in seven languages.
The Challenges of Optimizing Machine Translation for Low Resource Cross-Language Information Retrieval (D19-1)

Copied to clipboard

Challenge: Existing studies do not investigate the effectiveness of MT metrics in predicting performance of downstream IR models.
Approach: They examine the relationship between MT performance and IR quality in a CLIR-based system . they find that the choice of IR collection can significantly affect MT tuning decisions .
Outcome: The proposed model can predict CLIR performance better from MT quality, the authors show . the proposed model is based on a BLEU-based model with a bag of words constraint .
Design Challenges in Low-resource Cross-lingual Entity Linking (2020.emnlp-main)

Copied to clipboard

Challenge: Existing techniques for grounding mentions of entities in a foreign language do not rise to the challenges introduced by text in low-resource languages (LRL) and fail to generalize to text not taken from Wikipedia, on which they are usually trained.
Approach: They propose a cross-lingual XEL technique that uses search engines to locate and search for foreign language entries in Wikipedia.
Outcome: The proposed system shows an increase of 25% in gold candidate recall and 13% in end-to-end linking accuracy over state-of-the-art baselines.
Learning Cross-Lingual IR from an English Retriever (2022.naacl-main)

Copied to clipboard

Challenge: DR.DECR is a cross-lingual information retrieval system trained using multi-stage knowledge distillation (KD) DRDECR demonstrates superior accuracy over direct fine-tuning with labeled CLIR data.
Approach: They propose a cross-lingual information retrieval system with multi-stage knowledge distillation . they teach powerful multilingual representations and CLIR by optimizing two corresponding KD objectives .
Outcome: The proposed system is the best single-model retriever on the XOR-TyDi benchmark . it is based on a multi-stage knowledge distillation process that can be expensive .
Cross-Lingual Link Discovery for Under-Resourced Languages (2022.lrec-1)

Copied to clipboard

Challenge: Linked data paradigms can be used to solve under-resourced languages' problem of under-utilization of resources.
Approach: They propose a paradigm for cross-lingual link discovery that can be applied to under-resourced languages . they argue that techniques for cross language linking can be readily applied .
Outcome: The proposed technologies can be applied to under-resourced languages, the authors argue . the authors show that the Linked Data paradigm can be used to solve the problem .
Multilingual Dependency Parsing for Low-Resource Languages: Case Studies on North Saami and Komi-Zyrian (L18-1)

Copied to clipboard

Challenge: Developing systems for low-resource languages is a crucial issue for Natural Language Processing (NLP).
Approach: They propose a method for parsing low-resource languages with very small training corpora using multilingual word embeddings and annotated corporata of larger languages.
Outcome: The proposed method improves dependency parsing for low-resource languages with very small training corpora compared to previous work . it also explores whether contemporary contact languages or genetically related languages would be the most fruitful starting point for multilingual parsers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations