Challenge: a corpus of 43 million atomic edits is available for Wikipedia edit history . edits are instances in which a human editor has inserted a single contiguous phrase into, or deleted a contigous phrase from, an existing sentence.
Approach: They use Wikipedia edit history to mine atomic edits across 8 languages . they find edits contain instances in which a human editor has inserted a single phrase into, or deleted a contiguous phrase from, an existing sentence.
Outcome: The data show that edits differ from the language observed in standard corpora and that models trained on edits encode different aspects of semantics and discourse than models trained in raw text.

Similar Papers

The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
The Million Authors Corpus: A Cross-Lingual and Cross-Domain Wikipedia Dataset for Authorship Verification (2025.findings-acl)

Copied to clipboard

Challenge: Authorship verification (AV) is a crucial task for identity verification, accountlinking, historical linguistics, and AI-generated text identification.
Approach: They propose to use Wikipedia's Million Authors Corpus to examine authorship verification models on a broad scale.
Outcome: The proposed dataset includes 60.08M textual chunks, contributed by 1.29M Wikipedia authors.
ESCAPE: a Large-scale Synthetic Corpus for Automatic Post-Editing (L18-1)

Copied to clipboard

Challenge: eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far.
Approach: a team of researchers develops a Synthetic Corpus for Automatic Post-Editing . eSCAPE is the largest freely-available Synthetic corpus for automatic post-editing released so far . the results prove that the models always improve MT quality with statistically significant gains .
Outcome: eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far.
ELQA: A Corpus of Metalinguistic Questions and Answers about English (2023.acl-long)

Copied to clipboard

Challenge: ELQA corpus is metalinguistic—it consists of language about language.
Approach: They present a corpus of questions and answers in and about the English language . they use a free-form question answering task and multiple LLMs to analyze their capacity .
Outcome: The ELQA corpus covers grammar, meaning, fluency, and etymology . the results can be used to investigate metalinguistic capabilities of NLU models .
The Multilingual Microblog Translation Corpus: Improving and Evaluating Translation of User-Generated Text (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of over 200,000 microblog translations supports translation of thirteen languages into English . large collections of parallel text, or bitext, are increasingly available in many languages .
Approach: They propose a corpus of over 200,000 microblog posts that supports translation of thirteen languages into English.
Outcome: The proposed corpus contains over 200,000 translations of microblog posts in 13 languages . fine-tuning showed significant improvements in translation quality .
Models and Datasets for Cross-Lingual Summarisation (2021.emnlp-main)

Copied to clipboard

Challenge: Recent years have witnessed increased interest in abstractive summarisation thanks to the popularity of neural network models and the availability of datasets containing hundreds of thousands of document-summary pairs.
Approach: They propose to create a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in . target language.
Outcome: The proposed task can be applied to several other languages and covers twelve languages and directions.
A Corpus of Encyclopedia Articles with Logical Forms (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of annotated typed lambda calculus translations is described in this paper . typed Lambda Calculus expressions are intended to serve as a theory-neutral formal representation .
Approach: They describe an annotated corpus of typed lambda calculus translations for 2,000 sentences in Simple English Wikipedia.
Outcome: The annotated typed lambda calculus translations are used in a corpus of 2,000 sentences in Simple English Wikipedia.
Wiki-40B: Multilingual Language Model Dataset (2020.lrec-1)

Copied to clipboard

Challenge: We propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families.
Approach: They propose a multilingual language model benchmark composed of 40+ languages . they train monolingual causal language models using a state-of-the-art model .
Outcome: The proposed model is composed of 40+ languages spanning several scripts and linguistic families.
Summarization Corpora of Wikipedia Articles (2020.lrec-1)

Copied to clipboard

Challenge: Using Wikipedia articles, we extract summarization data for other languages.
Approach: They propose a process to extract Wikipedia summarization corpora and apply it to the German language.
Outcome: The proposed method can be applied to the German language and compares to baselines.
Transforming Wikipedia into a Large-Scale Fine-Grained Entity Type Corpus (L18-1)

Copied to clipboard

Challenge: et al. (2017): WiFiNE annotated with fine-grained entity types . lack of a well-established training corpus makes it difficult to manually annotate the amount of data needed for training.
Approach: They propose an English corpus annotated with fine-grained entity types based on Wikipedia . they use heuristics to build a large, high quality, annotating corpus using 2 manually annotized benchmarks .
Outcome: The proposed system outperforms the existing systems with two datasets and gains a 2.8 macro F1 score.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations