Papers by Stella Frank

6 papers
Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era (2025.findings-acl)

Copied to clipboard

Challenge: danoneata, et al., 2021): human learning and conceptual representation is grounded in sensorimotor experience.
Approach: They evaluate image encoders and language-only models to learn which attributes are salient to the models.
Outcome: The proposed models outperform language-only models on attributes predicting extended denser McRae norms and newer Binder datasets.
VISaGE: Understanding Visual Generics and Exceptions (2025.emnlp-main)

Copied to clipboard

Challenge: atypical evaluation instances disrupt incontext instance understanding and in-weight conceptual knowledge.
Approach: They propose to use a dataset to analyze atypical visual and textual images to test their models.
Outcome: The proposed model is based on a dataset consisting of typical and exceptional images.
Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers (2021.emnlp-main)

Copied to clipboard

Challenge: Pretrained vision-and-language BERTs aim to learn representations that combine information from both modalities.
Approach: They propose a diagnostic method based on cross-modal input ablation to assess the extent to which pretrained models integrate cross-module information.
Outcome: The proposed method evaluates the model's performance on the other modality based on inputs from one or both modality.
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)

Copied to clipboard

Challenge: Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages.
Approach: They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices .
Outcome: The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices.
CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning (2020.acl-main)

Copied to clipboard

Challenge: Approaches to Grounded Language Learning focus on a single task-based final performance measure which may not depend on desirable properties of the learned hidden representations.
Approach: They propose an evaluation framework for Grounded Language Learning with Attributes based on three sub-tasks: 1) Goal-oriented evaluation; 2) Object attribute prediction evaluation; and 3) Zero-shot evaluation.
Outcome: The proposed framework evaluates the quality of learned representations with respect to attribute grounding.
Multilingual Multimodal Learning with Machine Translated Text (2022.findings-emnlp)

Copied to clipboard

Challenge: Currently, most vision-and-language pretraining research focuses on English tasks due to the availability of datasets.
Approach: They propose a framework for machine translating English multimodal data to improve training data . they propose two metrics to prevent models from learning from low-quality translated text .
Outcome: The proposed framework can be applied to any multimodal dataset and model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations