Generation of a Spanish Artificial Collocation Error Corpus (L18-1)

Copied to clipboard

Challenge: collocations are combinations of two elements where one (the base) is freely chosen, despite the limitations of the other (collocate) current tools for collocation error detection and correction focus on collocation validation and identification of miscollocations .
Approach: They propose an algorithm for automatic generation of an artificial collocation error corpus of american English learners of Spanish that includes 17 different types of collocation errors.
Outcome: The proposed algorithm can detect and classify collocation errors in learners' writings . collocation error detection and correction has not received the attention it deserves .

Similar Papers

Evaluating language models for the retrieval and categorization of lexical collocations (2021.eacl-main)

Copied to clipboard

Challenge: Lexical collocations are idiosyncratic combinations of two syntactically bound lexical items.
Approach: They perform an exhaustive analysis of current language models for collocation understanding . they first construct a dataset of apparitions of lexical collocations in context .
Outcome: The proposed models perform well in distinguishing light verb constructions, especially if the collocation’s first argument acts as subject, but often fail to distinguish, first, different syntactic structures within the same semantic category, and second, fine-grained semantic categories which restrict the use of small sets of valid collocates for a given base.
Pay Attention when you Pay the Bills. A Multilingual Corpus with Dependency-based and Semantic Annotation of Collocations. (P19-1)

Copied to clipboard

Challenge: resulting corpus can be useful for different NLP tasks such as natural language understanding or natural language generation.
Approach: They propose to annotate 155k tokens and 1,526 collocations in context in a multilingual corpus in English, Portuguese, and Spanish.
Outcome: The new corpus can be used to evaluate different approaches for collocation identification, which can be useful for different NLP tasks such as natural language understanding or natural language generation.
Wronging a Right: Generating Better Errors to Improve Grammatical Error Detection (D18-1)

Copied to clipboard

Challenge: grammatical error correction is a labor-intensive task that requires large amounts of training data.
Approach: They propose to use a human-annotated corpus of human-generated grammatical errors to generate a synthetic model.
Outcome: The proposed method outperforms the current state of the art in grammatical error correction . human annotators achieve 39.39 F1 scores, suggesting the model generates mostly human-like instances .
A Lexical Tool for Academic Writing in Spanish based on Expert and Novice Corpora (L18-1)

Copied to clipboard

Challenge: a corpus of academic texts provides lexical combinations for the production of academic text.
Approach: They describe the extraction of data from a corpus of academic texts and the use of those data to develop a lexical tool oriented to the production of academic text.
Outcome: The proposed tool will provide indications as to how to use vocabulary typical of the academic genre in order to build complete texts.
Correcting the Autocorrect: Context-Aware Typographical Error Correction via Training Data Augmentation (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that typographical errors are now ubiquitous . traditional spelling correction software is inadequate to correct typographical mistakes .
Approach: They propose to generate typographical errors based on annotated spelling errors . they then use annotations to introduce errors into substantially larger corpora .
Outcome: The proposed method generates typographical errors that require context-aware error detection . it also shows that machine learning can correct typographical mistakes based on the data .
Automatic Grammatical Error Correction for Sequence-to-sequence Text Generation: An Empirical Study (P19-1)

Copied to clipboard

Challenge: Sequence-to-sequence (seq2sequ) models have a weakness: they cannot always generate sentences without grammatical errors.
Approach: They propose to use automatic grammatical error correction to improve seq2seq models . they conduct experiments on machine translation, formality style transfer, sentence compression and simplification .
Outcome: The proposed system can improve grammaticality of generated text and improve formal style tasks.
A Simple Recipe for Multilingual Grammatical Error Correction (2021.acl-short)

Copied to clipboard

Challenge: Modern approaches view the task of Grammatical Error Correction (GEC) as monolingual text-to-text rewriting and employ encoderdecoder neural architectures.
Approach: They propose a language-agnostic method to generate a large number of synthetic examples and use large-scale multilingual language models to train state-of-the-art GEC models.
Outcome: The proposed method surpasses state-of-the-art results on GEC benchmarks in English, Czech, German and Russian.
Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling (2022.acl-long)

Copied to clipboard

Challenge: a corpus of Spanish newswire rich in unassimilated lexical borrowings is used to identify the language of a word.
Approach: They propose to annotate a corpus of Spanish newswire rich in unassimilated lexical borrowings and evaluate how models perform on this task.
Outcome: The proposed model outperforms models fed with subword embeddings and Transformer-based embeddables on the Spanish newswire corpus.
Minimally-Augmented Grammatical Error Correction (D19-55)

Copied to clipboard

Challenge: Existing approaches to automatic grammatical error correction require error-labelled training data to achieve their best performance.
Approach: They propose an unsupervised method that generates noise from inverted spell-checkers by using a synthetic error generation method.
Outcome: The proposed method outperforms the current state-of-the-art for German and Russian GEC tasks without using real error-labelled training data.
Low-Resource Grammatical Error Correction: Selective Data Augmentation with Round-Trip Machine Translation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for grammatical error correction require large amounts of parallel training data.
Approach: They propose to generate synthetic data through round-trip machine translation by generating a set of character-level errors using a technique known as SeLex-RT.
Outcome: The proposed technique produces errors similar to those observed with language learners, but lacks gold-labeled training data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations