Papers by Andrea Horbach
EVil-Probe - a Composite Benchmark for Extensive Visio-Linguistic Probing (2024.lrec-main)
Copied to clipboard
| Challenge: | Visual question answering, image-text retrieval and retrieving image patches that match an expression are some of the tasks visio-linguistic models show impressive performance on. |
| Approach: | They propose a composite benchmark that processes existing probing datasets into a unified format and reorganizes them based on the linguistic categories they probe. |
| Outcome: | The proposed benchmark is challenging for all models as they are sensitive to linguistic categories and only handles nouns. |
Rainbow - A Benchmark for Systematic Testing of How Sensitive Visio-Linguistic Models are to Color Naming (2024.eacl-long)
Copied to clipboard
| Challenge: | Visio-linguistic models have been gaining popularity for tasks that require a deeper understanding of multimodalities. |
| Approach: | They compile a probing dataset to test multi-modal alignment around color . they show that models have trouble with prepositions and verbs . |
| Outcome: | The proposed model is superior to models that do not rely on pre-extracted image features and is able to perform well with noisy pre-training data. |
When Argumentation Meets Cohesion: Enhancing Automatic Feedback in Student Writing (2024.lrec-main)
Copied to clipboard
| Challenge: | Argumentative essays require a high degree of cohesion, defined as a network of semantic relationships that link together. |
| Approach: | They investigate the role of arguments in the automatic scoring of cohesion in argumentative essays. |
| Outcome: | The proposed model improves on a multi-task learning process by adding argumentative elements as an auxiliary task. |
Similarity-Based Content Scoring - A more Classroom-Suitable Alternative to Instance-Based Scoring? (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent work suggests that similarity-based content scoring methods can yield comparable results to instance-based supervised learning. |
| Approach: | They propose to use similarity-based scoring to achieve similar results . they compare different instance-based and similarity based methods on multiple data sets . |
| Outcome: | The proposed approach has a lower need for annotated training data and better zero-shot performance, but the results are not consistent with previous studies. |
DARIUS: A Comprehensive Learner Corpus for Argument Mining in German-Language Essays (2024.lrec-main)
Copied to clipboard
Nils-Jonathan Schaller, Andrea Horbach, Lars Ingver Höft, Yuning Ding, Jan Luca Bahr, Jennifer Meyer, Thorben Jansen
| Challenge: | Existing corpora focus on specific out-of-school domains, such as legal documents. |
| Approach: | They present a digital argumentation instruction for science corpus on 4589 essays written by 1839 german secondary school students. |
| Outcome: | The proposed corpus is annotated according to a fine-grained annotation scheme on 4589 essays written by 1839 german secondary school students. |
Don’t take “nswvtnvakgxpm” for an answer –The surprising vulnerability of automatic content scoring systems to adversarial input (2020.coling-main)
Copied to clipboard
| Challenge: | Automated content scoring systems can be used on short answer tasks to save human effort, but can invite cheating strategies such as writing irrelevant answers. |
| Approach: | They generate adversarial answers for benchmark content scoring datasets based on different methods of increasing sophistication and examine countermeasures such as adversarials. |
| Outcome: | The proposed methods show that even simple methods can reduce content scoring performance but do not solve the problem. |
Linguistic Appropriateness and Pedagogic Usefulness of Reading Comprehension Questions (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing evaluation measures for automatic generation of reading comprehension questions focus on linguistic quality only, ignoring educational value and appropriateness of questions. |
| Approach: | They propose a new evaluation scheme where questions are structured in a hierarchical way . they also create and evaluate two new evaluation data sets for Basque and German . |
| Outcome: | The proposed evaluation scheme can be applied, but expert annotators are needed. |
FEAT-writing: An Interactive Training System for Argumentative Writing (2025.coling-demos)
Copied to clipboard
| Challenge: | Argumentative writing is a critical skill for academic success, but many students struggle to develop these skills. |
| Approach: | They developed an online system that provides students with automated feedback and exercises for argumentative writing. |
| Outcome: | The proposed system improves argumentative writing quality among native English speakers and english-as-a-foreign-language learners. |
ESCRITO - An NLP-Enhanced Educational Scoring Toolkit (L18-1)
Copied to clipboard
| Challenge: | Existing implementations are very specific to specific use cases and datasets. |
| Approach: | ESCRITO is a toolkit for scoring student writings using NLP techniques . authors propose teachers and NLP researchers to use APIs for scoring pipelines . |
| Outcome: | ESCRITO is a toolkit for scoring student writings using NLP techniques . it addresses two main user groups: teachers and NLP researchers . |
Semi-Supervised Clustering for Short Answer Scoring (L18-1)
Copied to clipboard
| Challenge: | Existing approaches to SAS use unsupervised clustering and have teachers label some items after clustering. |
| Approach: | They propose to use semi-supervised clustering to provide structured groups of answers in addition to a score. |
| Outcome: | The proposed method improves clustering performance from 0.504 kappa for unsupervised clustering to 0.566 kppa. |
Score It All Together: A Multi-Task Learning Study on Automatic Scoring of Argumentative Essays (2023.findings-acl)
Copied to clipboard
| Challenge: | a multi-task learning approach outperforms sequential approaches for scoring argumentative essays . segmentation and classification of argumentative elements are important steps towards providing feedback on writing structure, but assessing the quality of arguments is less researched . |
| Approach: | They use a student essay dataset to study how argumentative essays are scored . they use automated span detection, type and quality prediction to combine these tasks . |
| Outcome: | The proposed method outperforms sequential approaches for segmentation and quality prediction. |
Chinese Content Scoring: Open-Access Datasets and Features on Different Segmentation Levels (2020.aacl-main)
Copied to clipboard
| Challenge: | Unlike English, which uses spaces as natural separators between words, segmentation of Chinese texts into tokens is challenging. |
| Approach: | They present two data sets for Chinese content scoring that use Chinese short answer questions and a new scoring system that uses Chinese short-answer questions. |
| Outcome: | The proposed system performs better on lower segmentation levels than on token level. |
LeSpell - A Multi-Lingual Benchmark Corpus of Spelling Errors to Develop Spellchecking Methods for Learner Language (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing spellcheckers do not work well with learner data. |
| Approach: | They propose a multi-lingual evaluation data set of spelling mistakes in context that is highly customizable for the DKPro architecture. |
| Outcome: | The proposed spellchecker improves performance in many settings and can be customized to meet learners' needs. |