Papers with validity
Grading Massive Open Online Courses Using Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Massive open online courses (MOOCs) offer free education globally, but the massive enrollment in these courses makes it impractical for an instructor to assess every student’s writing assignment. |
| Approach: | They propose to use large language models to replace peer grading in MOOCs by using zero-shot chain-of-thought prompts to automate feedback process. |
| Outcome: | The proposed method automates the feedback process once the LLM assigns a score to an assignment. |
M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database (2022.acl-long)
Copied to clipboard
| Challenge: | Existing data resources to support multimodal affective analysis in dialogues are limited in scale and diversity. |
| Approach: | They propose a multimodal multi-scene multi-label Emotional Dialogue dataset, M3ED, which contains 990 dyadic emotional dialogues from 56 different TV series. |
| Outcome: | The proposed dataset contains 990 dyadic emotional dialogues from 56 different TV series, a total of 9,082 turns and 24,449 utterances. |
Are the Reasoning Models Good at Automated Essay Scoring? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | o3-mini and o4-mini reasoning models perform poorly on automated essay scoring tasks, despite excellent performance on many benchmarks. |
| Approach: | They evaluated OpenAI’s o3-mini and o4-mini reasoning models in automated essay scoring tasks by measuring agreement with expert ratings and consistency in repeated evaluations. |
| Outcome: | The models’ performance on the TOEFL11 dataset is evaluated by measuring agreement with expert ratings and consistency in repeated evaluations. |
What Makes Reading Comprehension Questions Easier? (D18-1)
Copied to clipboard
| Challenge: | Recent studies have shown that questions require a deeper understanding of language to answer beyond using superficial cues. |
| Approach: | They propose to use simple heuristics to split MRC datasets into easy and hard subsets and manually annotate questions from each subset with validity and reasoning skills to investigate which skills explain the difference between easy and harder questions. |
| Outcome: | The proposed model performs better for hard and easy questions than for easy questions. |
FFSTC: Fongbe to French Speech Translation Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Fongbe to French Speech Translation Corpus is a comprehensive dataset compiled through various collection methods and the efforts of dedicated individuals. |
| Approach: | They introduce the Fongbe to French Speech Translation Corpus (FFSTC) which encompasses approximately 31 hours of collected Fongbbe language content. |
| Outcome: | The proposed corpus includes both transcriptions and voice recordings in Fongbe and French. |
Context versus Prior Knowledge in Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies have investigated how often a model will rely on prior knowledge over conflicting contextual information in answering questions. |
| Approach: | They propose two mutual information-based metrics to measure a model’s dependency on a context and on its prior about an entity. |
| Outcome: | The proposed metrics show that language models can integrate prior knowledge and new information in a predictable way across different questions and contexts. |