Papers by Carina Silberer
Object Naming in Language and Vision: A Survey and a New Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | Object naming has been studied in Psycholinguistics, but has received little attention in Computational Linguistics. |
| Approach: | They propose a dataset that provides 36 name annotations for each of 25K objects in images selected from VisualGenome. |
| Outcome: | The proposed dataset shows that people choose certain names for objects, on average. |
Are Visual-Linguistic Models Commonsense Knowledge Bases? (2022.coling-1)
Copied to clipboard
| Challenge: | PTLMs are used to extract knowledge from text on demand. |
| Approach: | They compare visual-linguistic and language-only visual-language models in a zero-shot commonsense question answering inference task. |
| Outcome: | The proposed models are highly promising on certain types of commonsense knowledge associated with the visual world. |
Combining Tradition with Modernness: Exploring Event Representations in Vision-and-Language Models for Visual Goal-Step Inference (2023.acl-srw)
Copied to clipboard
| Challenge: | Existing methods for representing procedural knowledge are limited to capturing the most crucial information, namely actions and the participants, to learn stereotypical event sequences. |
| Approach: | They propose a task that uses images to identify steps towards achieving a goal in the multimodal domain. |
| Outcome: | The proposed task uses images that represent steps towards achieving a textually expressed goal in the multimodal domain. |
A Couch Potato is not a Potato on a Couch: Prompting Strategies, Image Generation, and Compositionality Prediction for Noun Compounds (2025.findings-acl)
Copied to clipboard
| Challenge: | a new method to predict the compositionality of English noun compounds is proposed . |
| Approach: | They propose a visual modality and vision transformers to predict the compositionality of English noun compounds. |
| Outcome: | The proposed method compared with a state-of-the-art text-based approach reveals complementary contributions regarding features and degrees of abstractness in English noun compounds. |
Humans Meet Models on Object Naming: A New Dataset and Analysis (2020.coling-main)
Copied to clipboard
| Challenge: | Existing object naming datasets that use only images with a bounding box are noisy . a human-like model behavior is not stable across domains, a study finds . |
| Approach: | They use MN v2 to verify object naming datasets with dozens of valid names per object . they find that human-like model behavior is not stable across domains . |
| Outcome: | The proposed model confuses people and clothing objects more frequently than humans do. |
Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts (2025.acl-long)
Copied to clipboard
| Challenge: | Accurate modeling of subjective phenomena requires data annotated with authors’ intentions. |
| Approach: | They collect study-created and genuine social media posts labeled for emotion and compare them on several dimensions, including model performance. |
| Outcome: | The results show that study-created posts are longer, rely more on text and less on images for emotion expression, and focus more on emotion-prototypical events. |
What do Entity-Centric Models Learn? Insights from Entity Linking in Multi-Party Dialogue (N19-1)
Copied to clipboard
| Challenge: | a recent study suggests that models that incorporate a bias towards learning entity representations are not effective at modeling entities. |
| Approach: | They propose to use two entity-centric models for a referential task . they show they outperform the state of the art and do better on lower frequency entities . |
| Outcome: | The proposed models outperform the state of the art on a referential task . they do better on lower frequency entities than a counterpart model not entity-centric . |
Grounding Semantic Roles in Images (D18-1)
Copied to clipboard
| Challenge: | Experimental results show that visual semantic role labeling is useful for text understanding . image-based role annotations are prohibitive, but the model induces frame-semantic visual representations . |
| Approach: | They propose to train a visual semantic role labeling model without prohibitive image annotations . they render candidate participants as image regions of objects and train vSRL model which learns to ground roles in the regions which depict the corresponding participant . |
| Outcome: | The proposed model trains without prohibitive image-based role annotations without prohibiting image-related annotations. |