Challenge: Existing scoring models do not take the features of the stories and video clips into account when scoring, which will reduce the accuracy of the models.
Approach: They propose to leverage the features extracted from stories and videos related to the questions being asked during the children’s mindreading evaluation.
Outcome: The proposed framework agrees well with human experts on scores produced by the models.

Similar Papers

“What is on your mind?” Automated Scoring of Mindreading in Childhood and Early Adolescence (2020.coling-main)

Copied to clipboard

Challenge: Existing studies show that children who excel at mindreading are more likely to be identified as popular by classmates and have reciprocated friendships.
Approach: They propose to automate the scoring of mindreading ability in middle childhood and early adolescence using a new corpus of 11,311 question-answer pairs in English from 1,066 children aged from 7 to 14 .
Outcome: The proposed scoring system is based on 11,311 question-answer pairs in English from 1,066 children aged from 7 to 14 . the results demonstrate the applicability of state-of-the-art NLP solutions to a new domain and task.
Can vectors read minds better than experts? Comparing data augmentation strategies for the automated scoring of children’s mindreading ability (2021.acl-long)

Copied to clipboard

Challenge: In-domain experts are recruited to reannotate augmented samples and determine to what extent each strategy preserves the original rating.
Approach: They implement 7 different data augmentation strategies for the task of automatic scoring of children’s ability to understand others’ thoughts, feelings, and desires.
Outcome: The data augmentation strategies outperform task-agnostic augmentations and automatic augmentation systems perform worst on the MIND-CA corpus.
Neural Automated Essay Scoring and Coherence Modeling for Adversarially Crafted Input (N18-1)

Copied to clipboard

Challenge: Existing approaches to Automated Essay Scoring (AES) are not well-suited to capture adversarially crafted input of grammatical but incoherent sequences of sentences.
Approach: They propose a neural model of local coherence that can effectively learn connectedness features between sentences.
Outcome: The proposed approach strengthens the validity of neural essay scoring models.
Evaluating Theory of Mind in Question Answering (D18-1)

Copied to clipboard

Challenge: a dataset is proposed for question answering models with respect to their capacity to reason about beliefs.
Approach: They propose a dataset for evaluating question answering models with respect to their capacity to reason about beliefs.
Outcome: The proposed dataset is inspired by theory-of-mind experiments that examine whether children are able to reason about beliefs of others.
StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference.
Approach: They propose a novel Story Evaluation method that mimics human preference when judging a story . the model is based on a well-annotated dataset and a longformer-encoder-decoder .
Outcome: The proposed method is applicable to machine-generated and human-written stories.
MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics (2020.emnlp-main)

Copied to clipboard

Challenge: Existing reading comprehension metrics rely on token overlap and are agnostic to the nuances of reading comprehension.
Approach: They propose a benchmark for training and evaluating generative reading comprehension metrics: MOdeling Correctness with Human Annotations.
Outcome: The proposed benchmark outperforms baseline metrics by 10 to 36 absolute Pearson points on held-out annotations.
SkillQG: Learning to Generate Question for Reading Comprehension Assessment (2023.findings-acl)

Copied to clipboard

Challenge: Existing question generation systems focus on the literal nature of questions and rarely consider comprehension types of the generated questions.
Approach: They propose a question generation framework with controllable comprehension types for machine reading comprehension models.
Outcome: Empirical results show that SkillQG outperforms baselines in quality, relevance, and skill-controllability while showing a performance boost in downstream question answering task.
Predicting Multidimensional Subjective Ratings of Children’ Readings from the Speech Signals for the Automatic Assessment of Fluency (2020.lrec-1)

Copied to clipboard

Challenge: Using a novel framework, we estimate the reading performance of young readers using linguistic and phonetic features.
Approach: They propose a framework for performing such an estimation that exploits multiple references performed by adults and demonstrate its efficiency using recordings of 273 pupils.
Outcome: The proposed framework exploits multiple references performed by adults and shows that it is efficient.
An Automatic Tool For Language Evaluation (2020.lrec-1)

Copied to clipboard

Challenge: standardized tests are used to assess and screen developmental language impairments but require manual laborious transcription, annotation and calculation.
Approach: They propose to use the correct sentence and the sentence produced by patients to evaluate the level of verbal production and return a score.
Outcome: The proposed system evaluates the level of the verbal production and returns a score.
Automated Essay Scoring: A Reflection on the State of the Art (2024.emnlp-main)

Copied to clipboard

Challenge: Automated essay scoring (AES) is a key application of natural language processing . it is based on a holistic score that summarizes the essay's overall quality .
Approach: aaron carroll: automated essay scoring is one of the most important applications in NLP . carroll says the task is still far from being solved, but it's still progressing steadily . he says it'll be interesting to see how researchers can improve performance numbers .
Outcome: a new neural model can beat existing models on a standard evaluation dataset, authors say . authors: the current model is not enough to improve performance numbers . they say it could spark discussion among researchers on how to move forward .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations