Challenge: a dataset of real errors in premodern Greek is presented to improve error detection methods . scribal errors are more difficult to detect than print or digitization errors.
Approach: They propose to annotate 1,000 words more likely to contain errors and annotated them as errors or not . they propose to evaluate new error detection methods that outperform other methods .
Outcome: The proposed method outperforms all other methods, improving true positive rate by 5%.

Similar Papers

AGILe: The First Lemmatizer for Ancient Greek Inscriptions (2022.lrec-1)

Copied to clipboard

Challenge: Existing models for ancient Greek inscriptions are not performant on epigraphic data due to language differences . a lemmatizer for ancient inscription data can enable meaningful generalizations, we show .
Approach: They propose to train an automatic lemmatizer for ancient Greek inscriptions with 80% accuracy . they also show that existing models are not performant on epigraphic data .
Outcome: The proposed model achieves above 80% accuracy on epigraphic data, and makes it available to the community.
On the Robustness of Language Encoders against Grammatical Errors (2020.acl-main)

Copied to clipboard

Challenge: Pre-trained language encoders are effective in facilitating downstream natural language processing tasks, but they often assume training and test corpora are clean and it is unclear how the models behave when confronted with noisy input.
Approach: They conduct adversarial attacks to simulate grammatical errors on clean text data.
Outcome: The proposed model performs better when confronted with natural grammatical errors than when faced with noisy input.
BERT-Proof Syntactic Structures: Investigating Errors in Discontinuous Constituency Parsing (2021.findings-acl)

Copied to clipboard

Challenge: Recent results show that pretrained language models can be used for many tasks with high accuracy and high performance.
Approach: They propose two methods for automatically analysing discontinuous parsers' errors.
Outcome: The proposed methods characterize errors of a state-of-the-art transition-based discontinuous parser and provide an overview of the contribution of BERT to this task.
Detecting Label Errors by Using Pre-Trained Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for label error detection focus on label errors in training data.
Approach: They propose a method for introducing realistic, human-originated label noise into existing crowdsourced datasets such as SNLI and TweetNLP.
Outcome: The proposed method outperforms existing methods for detecting label errors in natural language datasets.
Grammatical error detection in transcriptions of spoken English (2020.coling-main)

Copied to clipboard

Challenge: CrowdED corpus of spoken English monologues on business topics was crowdsourced from native speakers of English and learners of English with German as their first language.
Approach: They propose to use the corpus recordings to correct existing speech transcriptions and edit them to make them more fluent.
Outcome: The proposed transcription corrections and annotations can be used for automatic transcription post-editing and grammatical error correction for spoken English.
Enriching Grammatical Error Correction Resources for Modern Greek (2022.lrec-1)

Copied to clipboard

Challenge: Davidson and Kilgarriff, 2011) have focused on the English language, but there are limited efforts to expand GEC in other languages.
Approach: They develop and test a multilingual text-to-text transformer for Greek . they provide a model that can be fully-fledged for Greek with annotation corrections .
Outcome: The proposed model achieves 52.63% F0.5 on part of the Greek Native Corpus, 16% below the winning system on English GEC.
Lemmatisation of Medieval Greek: Against the Limits of Transformer’s Capabilities? (2024.lrec-main)

Copied to clipboard

Challenge: Existing lemmatisation algorithms display an accuracy drop of around 30pp when tested on unedited, Byzantine Greek epigrams.
Approach: They propose to use transformer-based embeddings and a dictionary look-up to lemmatise unedited, Byzantine Greek epigrams.
Outcome: The proposed method outperforms existing methods and provides detailed error analysis revealing why unedited, Byzantine Greek is so challenging for lemmatisation.
Development of Numerical Error Detection Tasks to Analyze the Numerical Capabilities of Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing language models are difficult to detect numerical errors because of their finite set of tokens.
Approach: They use a benchmark dataset to classify numerical errors using automatically generated numerical errors and investigate their ability to detect errors.
Outcome: The proposed model performs well in the numerical error detection task, but not as accurate as humans.
TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language Models (2021.acl-long)

Copied to clipboard

Challenge: Using pretrained language models, we propose an error-annotated dataset for text generation . we use carefully selected prompt words to guide GPT-2 to generate candidate sentences .
Approach: They propose an error-annotated dataset with multiple benchmark tasks for text generation from pretrained language models.
Outcome: The proposed dataset covers 24 types of errors according to common sense and linguistics.
Sprucing up the trees – Error detection in treebanks (C18-1)

Copied to clipboard

Challenge: a method for detecting annotation errors in manually annotated dependency trees is presented . the method is based on ensemble parsing and Bayesian inference guided by active learning .
Approach: They propose a method for detecting annotation errors in manually annotated dependency parse trees . they use ensemble parsing in combination with Bayesian inference guided by active learning .
Outcome: The proposed method detects errors in annotated dependency treebanks and improves parsing accuracy on in- and out-of-domain data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations