Papers by Marie Bexte
EVil-Probe - a Composite Benchmark for Extensive Visio-Linguistic Probing (2024.lrec-main)
Copied to clipboard
| Challenge: | Visual question answering, image-text retrieval and retrieving image patches that match an expression are some of the tasks visio-linguistic models show impressive performance on. |
| Approach: | They propose a composite benchmark that processes existing probing datasets into a unified format and reorganizes them based on the linguistic categories they probe. |
| Outcome: | The proposed benchmark is challenging for all models as they are sensitive to linguistic categories and only handles nouns. |
Rainbow - A Benchmark for Systematic Testing of How Sensitive Visio-Linguistic Models are to Color Naming (2024.eacl-long)
Copied to clipboard
| Challenge: | Visio-linguistic models have been gaining popularity for tasks that require a deeper understanding of multimodalities. |
| Approach: | They compile a probing dataset to test multi-modal alignment around color . they show that models have trouble with prepositions and verbs . |
| Outcome: | The proposed model is superior to models that do not rely on pre-extracted image features and is able to perform well with noisy pre-training data. |
Similarity-Based Content Scoring - A more Classroom-Suitable Alternative to Instance-Based Scoring? (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent work suggests that similarity-based content scoring methods can yield comparable results to instance-based supervised learning. |
| Approach: | They propose to use similarity-based scoring to achieve similar results . they compare different instance-based and similarity based methods on multiple data sets . |
| Outcome: | The proposed approach has a lower need for annotated training data and better zero-shot performance, but the results are not consistent with previous studies. |
Linguistic Appropriateness and Pedagogic Usefulness of Reading Comprehension Questions (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing evaluation measures for automatic generation of reading comprehension questions focus on linguistic quality only, ignoring educational value and appropriateness of questions. |
| Approach: | They propose a new evaluation scheme where questions are structured in a hierarchical way . they also create and evaluate two new evaluation data sets for Basque and German . |
| Outcome: | The proposed evaluation scheme can be applied, but expert annotators are needed. |
Score It All Together: A Multi-Task Learning Study on Automatic Scoring of Argumentative Essays (2023.findings-acl)
Copied to clipboard
| Challenge: | a multi-task learning approach outperforms sequential approaches for scoring argumentative essays . segmentation and classification of argumentative elements are important steps towards providing feedback on writing structure, but assessing the quality of arguments is less researched . |
| Approach: | They use a student essay dataset to study how argumentative essays are scored . they use automated span detection, type and quality prediction to combine these tasks . |
| Outcome: | The proposed method outperforms sequential approaches for segmentation and quality prediction. |
LeSpell - A Multi-Lingual Benchmark Corpus of Spelling Errors to Develop Spellchecking Methods for Learner Language (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing spellcheckers do not work well with learner data. |
| Approach: | They propose a multi-lingual evaluation data set of spelling mistakes in context that is highly customizable for the DKPro architecture. |
| Outcome: | The proposed spellchecker improves performance in many settings and can be customized to meet learners' needs. |
PictureStories: Predicting the Task Adherence of Language Learner Answers to a Picture Story-Based Writing Task (2026.eacl-long)
Copied to clipboard
| Challenge: | a lack of suitable training and evaluation data limits the evaluation of language learning tasks to language proficiency only. |
| Approach: | They develop a marking rubric that covers task adherence with respect to form and content. |
| Outcome: | The proposed model can predict the adherence of learners to written tasks using picture stories. |