Papers by Sua Lee

1 papers
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge (2026.acl-long)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) are increasingly used as automatic judges . however, their reliability and vulnerabilities to biases remain underexplored .
Approach: They propose a benchmark to evaluate MLLMs that fail to integrate visual cues . they also introduce a test to evaluate the reliability of MLMLs based on a set of asymmetric evaluation tendencies.
Outcome: Experiments on 26 state-of-the-art MLLMs reveal modality neglect and asymmetric evaluation tendencies . a standardized model with a benchmark enables a fine-grained diagnosis of nine bias types .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations