Somali Information Retrieval Corpus: Bridging the Gap between Query Translation and Dedicated Language Resources (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing research on the Somali language information retrieval relies on query translation . lack of digital resources is key obstacle to advancing language technologies . |
| Approach: | They develop an annotated corpus for Somali information retrieval using query expansion technique. |
| Outcome: | The proposed corpus comprises 2335 documents collected from well-known online sites . it can be used for text classification-related tasks and question-answering research purposes. |
Similar Papers
Corpus-Steered Query Expansion with Large Language Models (2024.eacl-short)
Copied to clipboard
| Challenge: | Recent studies show query expansions generate hypothetical documents that answer queries as expansions. |
| Approach: | They propose a corpus-steered query expansion to promote incorporation of knowledge embedded within the corpus. |
| Outcome: | et al. analyzed corpus-based Query Expansion (CSQE) using LLMs to generate hypothetical documents that answer the query. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets for cross-lingual information retrieval are limited in many languages, especially those spoken in Africa. |
| Approach: | They propose to build a test collection for cross-lingual information retrieval in 15 diverse African languages. |
| Outcome: | AfriCLIRMatrix contains 6 million queries in English and 23 million relevance judgments automatically mined from Wikipedia inter-language links, covering many more African languages than any existing information retrieval test collection. |
The Hebrew Essay Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotated corpus of argumentative essays authored by prospective higher-education students . corpus includes essays by native speakers and essays by non-native speakers . |
| Approach: | They propose to use an annotated corpus of Hebrew argumentative essays to analyze non-native language use. |
| Outcome: | The proposed corpus includes essays by native speakers and essays authored by non-native speakers with three different native languages. |
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)
Copied to clipboard
| Challenge: | a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics . |
| Approach: | They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development . |
| Outcome: | This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development . |
ELQA: A Corpus of Metalinguistic Questions and Answers about English (2023.acl-long)
Copied to clipboard
| Challenge: | ELQA corpus is metalinguistic—it consists of language about language. |
| Approach: | They present a corpus of questions and answers in and about the English language . they use a free-form question answering task and multiple LLMs to analyze their capacity . |
| Outcome: | The ELQA corpus covers grammar, meaning, fluency, and etymology . the results can be used to investigate metalinguistic capabilities of NLU models . |
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead (2025.emnlp-main)
Copied to clipboard
| Challenge: | African languages are often left behind in state-of-the-art natural language processing systems and large language models. |
| Approach: | They analyze 884 research papers on NLP for African languages published over past five years . they identify key trends shaping the field and outline promising directions . |
| Outcome: | The findings identify key trends shaping the field and outline promising directions . the authors analyze 884 research papers on NLP for African languages published over the past five years . |
A description and demonstration of SAFAR framework (2021.eacl-demos)
Copied to clipboard
Karim Bouzoubaa, Younes Jaafar, Driss Namly, Ridouane Tachicart, Rachida Tajmout, Hakima Khamar, Hamid Jaafar, Lhoussain Aouragh, Abdellah Yousfi
| Challenge: | Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language . |
| Approach: | They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework" |
| Outcome: | The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect. |
Masader: Metadata Sourcing for Arabic Text and Speech Data Resources (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, there is no online catalogue for Arabic datasets with annotated attributes . this paper aims to identify the publicly available Arabic dataset and provide a catalogue of them to researchers. |
| Approach: | They propose to create the largest public catalogue for Arabic NLP datasets with 25 attributes and a metadata annotation strategy that could be extended to other languages. |
| Outcome: | The proposed approach could be extended to other languages and regions. |
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)
Copied to clipboard
| Challenge: | Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate. |
| Approach: | They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs. |
| Outcome: | The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis. |