Measure Children’s Mindreading Ability with Machine Reading (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing scoring models do not take the features of the stories and video clips into account when scoring, which will reduce the accuracy of the models. |
| Approach: | They propose to leverage the features extracted from stories and videos related to the questions being asked during the children’s mindreading evaluation. |
| Outcome: | The proposed framework agrees well with human experts on scores produced by the models. |
Similar Papers
“What is on your mind?” Automated Scoring of Mindreading in Childhood and Early Adolescence (2020.coling-main)
Copied to clipboard
Venelin Kovatchev, Phillip Smith, Mark Lee, Imogen Grumley Traynor, Irene Luque Aguilera, Rory Devine
| Challenge: | Existing studies show that children who excel at mindreading are more likely to be identified as popular by classmates and have reciprocated friendships. |
| Approach: | They propose to automate the scoring of mindreading ability in middle childhood and early adolescence using a new corpus of 11,311 question-answer pairs in English from 1,066 children aged from 7 to 14 . |
| Outcome: | The proposed scoring system is based on 11,311 question-answer pairs in English from 1,066 children aged from 7 to 14 . the results demonstrate the applicability of state-of-the-art NLP solutions to a new domain and task. |
Can vectors read minds better than experts? Comparing data augmentation strategies for the automated scoring of children’s mindreading ability (2021.acl-long)
Copied to clipboard
| Challenge: | In-domain experts are recruited to reannotate augmented samples and determine to what extent each strategy preserves the original rating. |
| Approach: | They implement 7 different data augmentation strategies for the task of automatic scoring of children’s ability to understand others’ thoughts, feelings, and desires. |
| Outcome: | The data augmentation strategies outperform task-agnostic augmentations and automatic augmentation systems perform worst on the MIND-CA corpus. |
Neural Automated Essay Scoring and Coherence Modeling for Adversarially Crafted Input (N18-1)
Copied to clipboard
| Challenge: | Existing approaches to Automated Essay Scoring (AES) are not well-suited to capture adversarially crafted input of grammatical but incoherent sequences of sentences. |
| Approach: | They propose a neural model of local coherence that can effectively learn connectedness features between sentences. |
| Outcome: | The proposed approach strengthens the validity of neural essay scoring models. |
Evaluating Theory of Mind in Question Answering (D18-1)
Copied to clipboard
| Challenge: | a dataset is proposed for question answering models with respect to their capacity to reason about beliefs. |
| Approach: | They propose a dataset for evaluating question answering models with respect to their capacity to reason about beliefs. |
| Outcome: | The proposed dataset is inspired by theory-of-mind experiments that examine whether children are able to reason about beliefs of others. |
StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference. |
| Approach: | They propose a novel Story Evaluation method that mimics human preference when judging a story . the model is based on a well-annotated dataset and a longformer-encoder-decoder . |
| Outcome: | The proposed method is applicable to machine-generated and human-written stories. |
MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing reading comprehension metrics rely on token overlap and are agnostic to the nuances of reading comprehension. |
| Approach: | They propose a benchmark for training and evaluating generative reading comprehension metrics: MOdeling Correctness with Human Annotations. |
| Outcome: | The proposed benchmark outperforms baseline metrics by 10 to 36 absolute Pearson points on held-out annotations. |
SkillQG: Learning to Generate Question for Reading Comprehension Assessment (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing question generation systems focus on the literal nature of questions and rarely consider comprehension types of the generated questions. |
| Approach: | They propose a question generation framework with controllable comprehension types for machine reading comprehension models. |
| Outcome: | Empirical results show that SkillQG outperforms baselines in quality, relevance, and skill-controllability while showing a performance boost in downstream question answering task. |
Predicting Multidimensional Subjective Ratings of Children’ Readings from the Speech Signals for the Automatic Assessment of Fluency (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a novel framework, we estimate the reading performance of young readers using linguistic and phonetic features. |
| Approach: | They propose a framework for performing such an estimation that exploits multiple references performed by adults and demonstrate its efficiency using recordings of 273 pupils. |
| Outcome: | The proposed framework exploits multiple references performed by adults and shows that it is efficient. |
An Automatic Tool For Language Evaluation (2020.lrec-1)
Copied to clipboard
| Challenge: | standardized tests are used to assess and screen developmental language impairments but require manual laborious transcription, annotation and calculation. |
| Approach: | They propose to use the correct sentence and the sentence produced by patients to evaluate the level of verbal production and return a score. |
| Outcome: | The proposed system evaluates the level of the verbal production and returns a score. |
Automated Essay Scoring: A Reflection on the State of the Art (2024.emnlp-main)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a key application of natural language processing . it is based on a holistic score that summarizes the essay's overall quality . |
| Approach: | aaron carroll: automated essay scoring is one of the most important applications in NLP . carroll says the task is still far from being solved, but it's still progressing steadily . he says it'll be interesting to see how researchers can improve performance numbers . |
| Outcome: | a new neural model can beat existing models on a standard evaluation dataset, authors say . authors: the current model is not enough to improve performance numbers . they say it could spark discussion among researchers on how to move forward . |