| Challenge: | a crowdsourcing-based approach to model sentence complexity is proposed . word-level predictors shown to correlate with greater processing difficulties are e.g. word frequency, age of acquisition, root frequency effect, orthographic neighbourhood frequency . |
| Approach: | They propose a crowdsourcing-based approach to model human perception of sentence complexity using a corpus of sentences rated with judgments of complexity for two typologically-different languages. |
| Outcome: | The proposed model predicts agreement among annotators independently from the assigned judgment and the perception of sentence complexity in Italian and English. |
Similar Papers
Crowdsourced Corpus of Sentence Simplification with Core Vocabulary (L18-1)
Copied to clipboard
| Challenge: | a crowdsourced corpus of simplified sentences is used to generate complex sentences from more complex ones. |
| Approach: | They propose to use crowdsourced data set of simplified sentences from Japanese textbooks and reference books to generate simplified sentences. |
| Outcome: | The proposed set of simplified sentences is a good quality data set for machine learning. |
Word Complexity is in the Eye of the Beholder (2021.naacl-main)
Copied to clipboard
| Challenge: | Lexical complexity is a subjective notion, yet it is often neglected in lexical simplification and readability systems which use a ”one-size-fits-all” approach. |
| Approach: | They propose to use a dataset of complex words annotated by readers with different backgrounds to investigate which aspects contribute to the notion of lexical complexity. |
| Outcome: | The proposed approach can be replicated in a dataset of complex words annotated by readers with different backgrounds. |
Subjective Text Complexity Assessment for German (2022.lrec-1)
Copied to clipboard
| Challenge: | Often, readability is defined as how easily a written text is to read. |
| Approach: | They propose to use a corpus of sentences provided by a German IT service provider to assess the readability of German text. |
| Outcome: | The proposed model can predict complexity of German text by using linguistically motivated features. |
Crowd-sourcing annotation of complex NLU tasks: A case study of argumentative content annotation (D19-59)
Copied to clipboard
| Challenge: | Recent advances in machine reading and listening comprehension involve the annotation of long texts. |
| Approach: | They propose a way to perform a sentence-by-sentence annotation task with crowd annotators. |
| Outcome: | The proposed approach can be used to identify claims in a debate speech. |
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English (2024.lrec-main)
Copied to clipboard
| Challenge: | Syntactic acceptance dataset is a resource being designed for syntax and computational linguistics research. |
| Approach: | They propose to use the Syntactic Acceptability Dataset to examine the syntactical discourse. |
| Outcome: | The proposed dataset is the largest of its kind that is publicly accessible. |
A Word-Complexity Lexicon and A Neural Readability Ranking Model for Lexical Simplification (D18-1)
Copied to clipboard
| Challenge: | Current lexical simplification approaches rely on heuristics and corpus level features that do not align with human judgment. |
| Approach: | They propose a human-rated word-complexity lexicon and a neural readability ranking model that uses human ratings to measure the complexity of any given word or phrase. |
| Outcome: | The proposed model performs better than state-of-the-art models for lexical simplification tasks and evaluation datasets. |
Estimating Lexical Complexity from Document-Level Distributions (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for complexity estimation are limited to entire documents . health assessment tools are too short for existing methods to apply . |
| Approach: | They propose a two-step approach for estimating lexical complexity that does not rely on pre-annotated data. |
| Outcome: | The proposed method is tested on the Norwegian language and compares with other assessment tools. |
Complex Word Identification: A Comparative Study between ChatGPT and a Dedicated Model for This Task (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to assess lexical complexity are used to evaluate the difficulty of vocabulary for language learners. |
| Approach: | They propose to use pre-trained language models to assess the complexity of a word based on its context. |
| Outcome: | The proposed method outperforms the best systems in SemEval-2021. |
Using Eye-tracking Data to Predict the Readability of Brazilian Portuguese Sentences in Single-task, Multi-task and Sequential Transfer Learning Approaches (2020.coling-main)
Copied to clipboard
Sidney Evaldo Leal, João Marcos Munguba Vieira, Erica dos Santos Rodrigues, Elisângela Nogueira Teixeira, Sandra Aluísio
| Challenge: | Sentence complexity assessment is a relatively new task in Natural Language Processing. |
| Approach: | They propose to use Brazilian Portuguese to evaluate sentences with linguistic features to improve readability. |
| Outcome: | The proposed model reaches the state-of-the-art for Brazilian Portuguese with 97.8% accuracy with linguistic features. |
A linguistically-motivated evaluation methodology for unraveling model’s abilities in reading comprehension tasks (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing models fail for linguistic characteristics of input examples, despite the impressive quantity of scientific studies dedicated to them, the capabilities, limitations, and risks of these models remain largely unknown. |
| Approach: | They propose to use semantic frame annotation to characterize examples by a small number of complexity factors to account for model’s difficulty. |
| Outcome: | The proposed evaluation methodology is based on the intuition that certain examples consistently yield lower scores regardless of model size or architecture. |