Papers by Michael Saxon

14 papers
Can Vision Language Models Understand Mimed Actions? (2025.findings-acl)

Copied to clipboard

Challenge: Nonverbal communication (NVC) is an integral part of human language, but it has been overlooked in natural language processing research.
Approach: They propose a multimodal multimodal recognition task that uses a corpus of mimed gestures to evaluate their understanding of NVC.
Outcome: The proposed task is based on 86 unique gestures with perturbations applied to avatar, background, and viewpoint for evaluating recognition robustness.
Not All Errors are Equal: Learning Text Generation Metrics using Stratified Error Synthesis (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing learning metrics are limited to tasks where large human ratings are available.
Approach: They propose a model-based natural language generation (NLG) evaluation metric that is highly correlated with human judgements without requiring human annotation.
Outcome: The proposed metric outperforms all prior unsupervised metrics on multiple NLG tasks including translation, image captioning, and WebNLG text generation.
Let’s Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies show vision-language systems can reason about images using natural language, but their capacity for video reasoning remains underexplored.
Approach: They propose to frame video reasoning as the sequential understanding of a small number of keyframes, thereby leveraging the power and robustness of vision-language systems' capacity to reason about images using natural language.
Outcome: The proposed models can generate multiple intermediate keyframes and predict future keyframe, and they perform poorly on GPT-4, GPT-3, and VICUNA.
Multilingual Conceptual Coverage in Text-to-Image Models (2023.acl-long)

Copied to clipboard

Challenge: Neural text-to-image systems generate coherent, visually-appealing images with novel combinations of objects, scenarios, and styles.
Approach: They propose a technique to benchmark the degree to which a generative text-to-image system provides multilingual parity to its training language in terms of tangible nouns.
Outcome: The proposed technique can be used to benchmark T2I models in terms of multilinguality and identify model-specific weaknesses, spurious correlations, and biases without a-priori assumptions.
CausalDialogue: Modeling Utterance-level Causality in Conversations (2023.findings-acl)

Copied to clipboard

Challenge: Despite widespread adoption, neural conversation models have yet to exhibit natural chat capabilities with humans . despite their widespread adoption in society, chatbots have yet not shown natural chat capability .
Approach: They propose a causality-enhanced method to enhance the impact of causality at the utterance level in training neural conversation models.
Outcome: The proposed method improves diversity and agility of loss functions and still needs improvement . the proposed method is based on a CausalDialogue dataset .
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts (2024.naacl-short)

Copied to clipboard

Challenge: With growth in the popularity of text-to-image models has come interest in assessing their multilingual capabilities, including multilingual accessibility.
Approach: They propose to correct translation errors in a concept list translated to seven languages and compare the outputs of the benchmark to those conditioned on the old.
Outcome: The proposed benchmark contains translation errors in Spanish, Japanese, and Chinese.
Modeling Disclosive Transparency in NLP Application Descriptions (2021.emnlp-main)

Copied to clipboard

Challenge: Broader disclosive transparency is difficult to define and quantify, authors say . previous work has demonstrated trade-offs and negative consequences to disclosing transparency .
Approach: They propose to use neural language model-based probabilistic metrics to model disclosive transparency . they demonstrate that they correlate with user and expert opinions of system transparency a valid objective proxy .
Outcome: The proposed metrics correlate with user and expert opinions of system transparency, making them a valid objective proxy.
Do You Know About My Nation? Investigating Multilingual Language Models’ Cultural Literacy Through Factual Knowledge (2025.emnlp-main)

Copied to clipboard

Challenge: Multilingual question-answering benchmarks do not factor in regional diversity in the information they capture and tend to be Western-centric.
Approach: They propose to benchmark eight standard multilingual LLMs on XNationQA and evaluate them using two novel transference metrics.
Outcome: The proposed model shows greater knowledge of cultural information in English than in the dominant language of the respective culture.
Culture is Everywhere: A Call for Intentionally Cultural Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to evaluate cultural alignment of large language models are too trivial and focus on static facts and values.
Approach: They argue for intentionally cultural evaluation: an approach that examines cultural assumptions . they characterize what, how, and circumstances by which culturally contingent considerations arise in evaluation .
Outcome: The authors argue for intentionally cultural evaluation: an approach that examines cultural assumptions embedded in all aspects of evaluation, not just in explicitly cultural tasks.
CAIRE: Cultural Attribution of Images with Retrieval (2026.eacl-long)

Copied to clipboard

Challenge: Current text-to-image models produce homogeneous outputs given under-specified prompts and their outputs are disproportionately biased toward Western cultures.
Approach: They propose a framework that assesses the degree of cultural relevance of an image, given a user-defined set of labels.
Outcome: The proposed evaluation metric surpasses baselines on a manually curated dataset of culturally salient but rare items built using language models by 22% F1 points.
PECO: Examining Single Sentence Label Leakage in Natural Language Inference Datasets through Progressive Evaluation of Cluster Outliers (2023.eacl-main)

Copied to clipboard

Challenge: Efforts to debias NLI have led to datasets that exhibit different kinds of bias than those shown before.
Approach: They propose a new technique to detect and reduce single sentence label leakage . leakage is a problem with many modern NLI datasets, they argue . future work must prioritize reducing this problem, they write .
Outcome: a new model-driven technique can detect leakage and detect subpopulations in the datasets which exhibit it . the proposed technique is based on the progressive evaluation of cluster outliers (PECO) . it allows objective measurement of leakage, and automatic detection of subpopulations in the data which exhibit leakage.
Investigating Memorization of Conspiracy Theories in Text Generation (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies examine conspiracy theories in social media, but they have not evaluated their presence in generative language models.
Approach: They examine the ability of generative language models to generate conspiracy theory text . they highlight the difficulties of this task and discuss the drawbacks .
Outcome: The proposed model can generate conspiracy theories without access to training data.
TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing video generation models struggle to interpret compositional changes and synthesize components across different time steps.
Approach: They propose a temporal compositionality benchmark that uses text prompts and ground truth videos to evaluate compositional changes in video.
Outcome: The proposed benchmark can be used for text-to-video and image-to video generation.
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts (2024.findings-emnlp)

Copied to clipboard

Challenge: evaluators of long-context vision language models (VLMs) have not kept up with the rapid development of open-weight long-constraint language models.
Approach: They propose a dynamic benchmark generator for evaluating long-context reasoning in vision language models.
Outcome: The proposed model can ignore irrelevant information when answering queries, showing that current models lack this capability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations