Challenge: Existing research on the Somali language information retrieval relies on query translation . lack of digital resources is key obstacle to advancing language technologies .
Approach: They develop an annotated corpus for Somali information retrieval using query expansion technique.
Outcome: The proposed corpus comprises 2335 documents collected from well-known online sites . it can be used for text classification-related tasks and question-answering research purposes.

Similar Papers

Corpus-Steered Query Expansion with Large Language Models (2024.eacl-short)

Copied to clipboard

Challenge: Recent studies show query expansions generate hypothetical documents that answer queries as expansions.
Approach: They propose a corpus-steered query expansion to promote incorporation of knowledge embedded within the corpus.
Outcome: et al. analyzed corpus-based Query Expansion (CSQE) using LLMs to generate hypothetical documents that answer the query.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for cross-lingual information retrieval are limited in many languages, especially those spoken in Africa.
Approach: They propose to build a test collection for cross-lingual information retrieval in 15 diverse African languages.
Outcome: AfriCLIRMatrix contains 6 million queries in English and 23 million relevance judgments automatically mined from Wikipedia inter-language links, covering many more African languages than any existing information retrieval test collection.
The Hebrew Essay Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Annotated corpus of argumentative essays authored by prospective higher-education students . corpus includes essays by native speakers and essays by non-native speakers .
Approach: They propose to use an annotated corpus of Hebrew argumentative essays to analyze non-native language use.
Outcome: The proposed corpus includes essays by native speakers and essays authored by non-native speakers with three different native languages.
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)

Copied to clipboard

Challenge: a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics .
Approach: They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development .
Outcome: This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development .
ELQA: A Corpus of Metalinguistic Questions and Answers about English (2023.acl-long)

Copied to clipboard

Challenge: ELQA corpus is metalinguistic—it consists of language about language.
Approach: They present a corpus of questions and answers in and about the English language . they use a free-form question answering task and multiple LLMs to analyze their capacity .
Outcome: The ELQA corpus covers grammar, meaning, fluency, and etymology . the results can be used to investigate metalinguistic capabilities of NLU models .
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead (2025.emnlp-main)

Copied to clipboard

Challenge: African languages are often left behind in state-of-the-art natural language processing systems and large language models.
Approach: They analyze 884 research papers on NLP for African languages published over past five years . they identify key trends shaping the field and outline promising directions .
Outcome: The findings identify key trends shaping the field and outline promising directions . the authors analyze 884 research papers on NLP for African languages published over the past five years .
A description and demonstration of SAFAR framework (2021.eacl-demos)

Copied to clipboard

Challenge: Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language .
Approach: They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework"
Outcome: The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect.
Masader: Metadata Sourcing for Arabic Text and Speech Data Resources (2022.lrec-1)

Copied to clipboard

Challenge: Currently, there is no online catalogue for Arabic datasets with annotated attributes . this paper aims to identify the publicly available Arabic dataset and provide a catalogue of them to researchers.
Approach: They propose to create the largest public catalogue for Arabic NLP datasets with 25 attributes and a metadata annotation strategy that could be extended to other languages.
Outcome: The proposed approach could be extended to other languages and regions.
Scoping natural language processing in Indonesian and Malay for education applications (2022.acl-srw)

Copied to clipboard

Challenge: Limited natural language processing resources are available for Indonesian and Malay varieties and are difficult to locate.
Approach: They propose to encourage collaboration and efficiency within NLP in Indonesian and Malay by identifying most published authors and research hubs.
Outcome: The findings suggest that the field is dominated by exploratory corpus work, machine reading of text gathered from the Internet, and sentiment analysis.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations