Challenge: Low resource languages often lack relevance annotations for cross-lingual information retrieval . when available, the training data has limited coverage for possible queries .
Approach: They propose a weakly supervised neural model for Cross-lingual information retrieval from low-resource languages using weak supervision instead of relevance annotations.
Outcome: The proposed model achieves 19 MAP points improvement compared to CNNs and 12 points improvement from machine translation-based CLIR models.

Similar Papers

Cross-Lingual Learning-to-Rank with Shared Representations (N18-2)

Copied to clipboard

Challenge: Cross-lingual information retrieval (CLIR) is a document retrieval task where the documents are written in a language different from that of the user's query.
Approach: They propose a large-scale dataset derived from Wikipedia to support CLIR research in 25 languages.
Outcome: The proposed model can improve the results of Swahili-English CLIR in Japanese and Japanese.
The Challenges of Optimizing Machine Translation for Low Resource Cross-Language Information Retrieval (D19-1)

Copied to clipboard

Challenge: Existing studies do not investigate the effectiveness of MT metrics in predicting performance of downstream IR models.
Approach: They examine the relationship between MT performance and IR quality in a CLIR-based system . they find that the choice of IR collection can significantly affect MT tuning decisions .
Outcome: The proposed model can predict CLIR performance better from MT quality, the authors show . the proposed model is based on a BLEU-based model with a bag of words constraint .
SARAL: A Low-Resource Cross-Lingual Domain-Focused Information Retrieval System for Effective Rapid Document Triage (P19-3)

Copied to clipboard

Challenge: a new cross-lingual information retrieval system for low-resource languages is available in less-frequently-taught languages . a multilingual system can search for relevant information in a haystack of documents in swahili or Somali . human-driven approaches to this problem are complicated in 'low-resourced' languages aaron sagar: "the key role played by humans in triaging results is complicated"
Approach: They propose an end-to-end cross-lingual information retrieval system for low-resource languages . the system enables English speakers to search foreign language repositories using English queries . it summarizes the retrieved documents in English with respect to a particular information need .
Outcome: The proposed system achieves top performance in the most recent IARPA MATERIAL CLIR+summarization evaluations.
Cross-Lingual Text Classification with Minimal Resources by Transferring a Sparse Teacher (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for transferring supervision across languages require expensive cross-lingual resources.
Approach: They propose a cross-lingual teacher-student method that generates "weak" supervision in a target language using minimal cross-linguistic resources.
Outcome: The proposed method outperforms state-of-the-art methods with a student classifier in 18 languages . it extracts and transfers only the most important task-specific seed words across languages based on translated seed words .
Multi-Source Cross-Lingual Model Transfer: Learning What to Share (P19-1)

Copied to clipboard

Challenge: Cross-lingual transfer learning (CLTL) is a viable method for building NLP models for a low-resource target language . however, many languages lack the labeled training data necessary for training deep neural nets for varying NLP tasks.
Approach: They propose a cross-lingual transfer learning method that leverages annotated data from other languages to build NLP models for a target language.
Outcome: The proposed model achieves significant performance gains over prior art over multiple text classification and sequence tagging tasks including a large-scale industry dataset.
Distant Supervision from Disparate Sources for Low-Resource Part-of-Speech Tagging (D18-1)

Copied to clipboard

Challenge: Low-resource languages lack manual annotated data to learn basic models such as part-of-speech (POS) taggers.
Approach: They propose a cross-lingual neural part-of-speech tagger that learns from disparate sources of distant supervision in a uniform framework.
Outcome: The proposed model scales to hundreds of low-resource languages without access to gold annotated data.
Improving Low-Resource Languages in Pre-Trained Multilingual Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Pre-trained multilingual language models are the foundation of many NLP approaches, but are often not well-supported by these models due to small available monolingual corpora.
Approach: They propose an unsupervised approach to improve cross-lingual representations of low-resource languages by bootstrapping word translation pairs from monolingual corpora and using them to improve language alignment.
Outcome: The proposed approach improves cross-lingual representations on low-resource languages using word retrieval and zero-shot named entity recognition.
Learning with Limited Data for Multilingual Reading Comprehension (D19-1)

Copied to clipboard

Challenge: Existing approaches to support question answering in a new language with limited training resources introduce noises to the training data due to translation or generation errors.
Approach: They propose a weakly-supervised framework that quantifies noises from automatically generated labels to deemphasize or fix noisy data in training.
Outcome: The proposed framework can deemphasize or fix noisy data in training on low-resource languages with varying similarity to English.
A Little Annotation does a Lot of Good: A Study in Bootstrapping Low-resource Named Entity Recognizers (D19-1)

Copied to clipboard

Challenge: Named entity recognition models rely on large amounts of labeled data, making them challenging to extend to new, lower-resource languages.
Approach: They propose a method for bootstrapping named entity recognition models in under-resourced languages . they use cross-lingual transfer learning and targeted annotation of only uncertain entities .
Outcome: The proposed method achieves competitive accuracy with just one-tenth of training data.
Low-resource Cross-lingual Event Type Detection via Distant Supervision with Minimal Effort (C18-1)

Copied to clipboard

Challenge: Currently, few or no language processing tools or resources exist for most languages . a problem is that there is not enough available training data even in resource-rich languages if the task is complex.
Approach: They propose to use a bilingual dictionary to train machine learning in a resource-poor language . they also explore adversarial training of bilingual word representations .
Outcome: The proposed approach gives similar performance in event-type detection tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations