Challenge: ad-hoc information retrieval methods usually require large amounts of annotated data to be effective.
Approach: They propose an open-source toolkit to automatically build large-scale English information retrieval datasets based on Wikipedia.
Outcome: The proposed toolkit builds large-scale English information retrieval datasets based on Wikipedia with 59,252 queries and 2,617,003 pairs.

Similar Papers

WikiTableT: A Large-Scale Data-to-Text Dataset for Generating Wikipedia Article Sections (2021.findings-acl)

Copied to clipboard

Challenge: Existing datasets for data-to-text generation focus on single-sentence generation or long-form generation.
Approach: They create a dataset that pairs Wikipedia sections with tabular data and various metadata.
Outcome: The proposed dataset can generate fluent and high quality texts but struggle with coherence and factuality.
X-WikiRE: A Large, Multilingual Resource for Relation Extraction as Machine Comprehension (D19-61)

Copied to clipboard

Challenge: Existing knowledge bases are heavily biased towards English, but Wikipedias cover very different topics in different languages.
Approach: They propose a multilingual dataset that frams relation extraction as a machine reading problem.
Outcome: The proposed model can be used to transfer models cross-lingually and improves knowledge base completion across languages.
MAKED: Multi-lingual Automatic Keyword Extraction Dataset (2022.lrec-1)

Copied to clipboard

Challenge: a large dataset of news articles spanning 20 languages is lacking for keyword extraction.
Approach: They propose a large-scale multi-lingual keyword extraction dataset for 11 of 20 languages . authors believe it will help advance the field of automatic keyword extraction .
Outcome: The proposed dataset is the first for 11 of 20 languages and is based on 540K+ news articles from the BBC News network.
Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus (L18-1)

Copied to clipboard

Challenge: et al. (2017): WiFiNE annotated with fine-grained entity types . lack of a well-established training corpus makes it difficult to manually annotate the amount of data needed for training.
Approach: They propose an English corpus annotated with fine-grained entity types based on Wikipedia . they use heuristics to build a large, high quality, annotating corpus using 2 manually annotized benchmarks .
Outcome: The proposed system outperforms the existing systems with two datasets and gains a 2.8 macro F1 score.
Deep Neural Networks at the Service of Multilingual Parallel Sentence Extraction (C18-1)

Copied to clipboard

Challenge: Existing models for parallel data harvesting from Wikipedia are language-independent, robust and highly scalable.
Approach: They propose an end-to-end neural model for large-scale parallel data harvesting from Wikipedia . their model is language-independent, robust, and highly scalable .
Outcome: The proposed model is language-independent, robust, and highly scalable.
LSOIE: A Large-Scale Dataset for Supervised Open Information Extraction (2021.eacl-main)

Copied to clipboard

Challenge: Open Information Extraction (OIE) systems extract factual propositions into n-ary tuples . current datasets are limited in size and diversity .
Approach: They propose to convert QA-SRL 2.0 dataset to large-scale OIE dataset LSOIE.
Outcome: The proposed dataset is 20 times larger than the next largest human-annotated OIE dataset.
Massively Multilingual Pronunciation Modeling with WikiPron (2020.lrec-1)

Copied to clipboard

Challenge: WikiPron is an open-source command-line tool for extracting pronunciation data from Wiktionary . the tool generates a database of 1.7 million pronunciations from 165 languages .
Approach: They propose a command-line tool for extracting pronunciation data from Wiktionary . they use it to generate a database of 1.7 million pronunciations from 165 languages .
Outcome: The proposed software generates a database of pronunciations for 165 languages . the proposed model is then validated by a grapheme-to-phoneme model .
A Multilingual Wikified Data Set of Educational Material (L18-1)

Copied to clipboard

Challenge: a crowdsourcing effort to annotate and link parallel texts has been unsuccessful . a data set of parallel texts in eleven languages is presented .
Approach: They present a wikified data set of English sentences linked to Wikipedia pages . they use crowdsourcing to annotate the texts and perform crowdsourcing for complex annotations .
Outcome: The proposed data set is valuable as it constitutes a rich resource . it includes annotated data of English sentences linked to translations in eleven languages .
CLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval (2020.emnlp-main)

Copied to clipboard

Challenge: Cross-Lingual Information Retrieval (CLIR) is a retrieval task in which search queries and candidate documents are written in different languages.
Approach: They present a massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval extracted automatically from Wikipedia.
Outcome: The proposed datasets are the largest and most comprehensive CLIR dataset to date.
WiCE: Real-World Entailment for Claims in Wikipedia (2023.emnlp-main)

Copied to clipboard

Challenge: Textual entailment models are increasingly used in fact-checking, presupposition verification in question answering, or summary evaluation.
Approach: They propose a new fine-grained textual entailment dataset built on natural claim and evidence pairs extracted from Wikipedia that provides en-tailment judgments over sub-sentence units of the claim and a minimal subset of evidence sentences that support each subclaim.
Outcome: The proposed dataset improves on multiple datasets at test time and shows that real claims involve verification and retrieval problems that existing models fail to address.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations