Papers by Yixin Wan

20 papers
Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks conflate factual correctness and normative fairness . a model may generate responses that are factually accurate but socially unfair .
Approach: They propose a benchmark to examine the boundary between fact and fair . they draw on representativeness bias, attribution bias and ingroup–outgroup bias to explain why models often misalign fact and faireness.
Outcome: The proposed model is based on ten frontier models and is available on github . it is compared with a standard model that generates people of color in Nazi-era uniforms .
V-ALPHASOCIAL: Benchmark and Self-Reflective Chain-of-Thought Generation for Visual Social Commonsense Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Social commonsense reasoning is a multimodal task that requires both textual and visual cues.
Approach: They propose a method that integrates visual cues into social commonsense reasoning tasks.
Outcome: The proposed method improves social commonsense reasoning on a multimodal foundation model.
LUME: LLM Unlearning with Multitask Evaluations (2025.findings-emnlp)

Copied to clipboard

Challenge: Unlearning aims to remove copyrighted, sensitive, or private content from large language models without a full retraining.
Approach: They propose a multi-task unlearning benchmark LUME that unlearns short novels, biographies and public biographie .
Outcome: The proposed benchmark unlearns short novels, biographies and public biographie . it also releases fine-tuned models with 1B and 7B parameter sizes as targets .
MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues (2026.findings-acl)

Copied to clipboard

Challenge: Existing contextual safety benchmarks are mostly single-turn and miss how malicious intent can emerge gradually or how the same scene can support both benign and exploitative goals.
Approach: They propose a benchmark that evaluates contextual safety in multimodal large language models . they observe persistent trade-offs between contextual safety and utility .
Outcome: The proposed model combines multi-turn and multi-switch scenarios to evaluate safety in multimodal large language models.
TheoremQA: A Theorem-driven Question Answering Dataset (2023.emnlp-main)

Copied to clipboard

Challenge: Recent LLMs like GPT-4 and PaLM-2 have made tremendous progress in solving fundamental math problems like GSM8K by achieving over 90% accuracy.
Approach: They propose to use theorem-driven question-answering dataset to evaluate AI models' ability to apply theoretic concepts to solving challenging science problems.
Outcome: TheoremQA is curated by domain experts and contains 800 high-quality questions covering 350 theoremics from Math, Physics, EE&CS, and Finance.
Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: In the era of evaluating large language models, data contamination is an increasingly prominent concern . static benchmarking has been used for evaluation, but there are limitations of *dynamic* benchmarks .
Approach: They propose a series of optimal design principles for *dynamic* benchmarking and analyze the limitations of existing *static* benchmarks.
Outcome: The proposed benchmarks highlight a critical gap in the evaluation of LLMs.
InsideOut: Measuring and Mitigating Insider–Outsider Bias in Interview Script Generation (2026.acl-long)

Copied to clipboard

Challenge: Recent research has raised concerns about culture-related fairness issues in LLM-generated content.
Approach: They propose to use 4,000 generation prompts and three evaluation metrics to quantify LLMs' **insider-outsider bias** .
Outcome: The proposed method reduces bias in Llama model by 89.70% and mitigates bias on Qwen by 82.54% on cultural alignment gap metric.
PIP: Parse-Instructed Prefix for Syntactically Controlled Paraphrase Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing fine-tuning methods for this task are costly and require updating the parameters of the entire model to adapt to the newly included syntax information.
Approach: They propose a method to instruct model’s encoder prefix to capture syntax-related knowledge by direct initiation and indirect optimization.
Outcome: The proposed methods are 10 times more efficient and learnable than existing methods.
“Kelly is a Warm Person, Joseph is a Role Model”: Gender Biases in LLM-Generated Reference Letters (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are an effective tool to assist individuals in writing documents.
Approach: They examine gender biases in large language models (LLMs)-generated reference letters . they find that models are biased because they are hallucinated .
Outcome: The proposed model-generated reference letters are evaluated on 2 popular LLMs- ChatGPT and Alpaca.
MACAROON: Training Vision-Language Models To Be Your Engaged Partners (2024.findings-emnlp)

Copied to clipboard

Challenge: Large vision-language models (LVLMs) generate detailed responses even when questions are ambiguous or unanswerable, leading to hallucinations and bias issues.
Approach: They propose a three-tiered hierarchy for questions of invalid, ambiguous, and personalizable nature to measure the proactive engagement capabilities of LVLMs.
Outcome: The proposed model generates contrastive response pairs for unlabeled questions, achieving 0.84 AAR, while maintaining comparable performance on general tasks.
Not Every Token Needs Forgetting: Selective Unlearning Balancing Forgetting and Utility in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Conventional unlearning approaches forget all tokens in a target document, including common tokens that carry general knowledge.
Approach: They propose a method that identifies a critical subset of tokens within the forgetting set that is relevant to the unwanted information and unlearns only those tokens.
Outcome: Experiments on two benchmarks and six baseline unlearning algorithms show that selective unlearning achieves effective unlearning on the targeted forget data.
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference (2025.findings-naacl)

Copied to clipboard

Challenge: Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences.
Approach: They propose a system-level evaluation framework that ranks LLMs based on their alignment with human preferences.
Outcome: The proposed framework aims to rank LLMs based on their performance and alignment with human preferences.
The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention (2024.emnlp-main)

Copied to clipboard

Challenge: Prompt-based “diversity interventions” are commonly adopted to improve the diversity of Text-to-Image models depicting individuals with diverse racial or gender traits.
Approach: They propose a benchmark to quantify the trade-off between using diversity interventions and preserving demographic factuality in T2I models.
Outcome: The proposed model significantly improves the demographic factuality under diversity interventions while preserving diversity.
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects (2025.acl-long)

Copied to clipboard

Challenge: Recent large-scale T2I models like DALLE-3 have made progress in reducing gender stereotypes when generating single-person images.
Approach: They propose a framework that queries T2I models to depict two individuals with gender-stereotyped social identities to evaluate gender biases.
Outcome: The proposed framework reduces gender stereotypes when generating images with more than one person.
LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected? (2024.findings-naacl)

Copied to clipboard

Challenge: Current research focuses on purely MGT detection without adequately addressing mixed scenarios including AI-revised Human-Written Text (HWT) and human-revealed MGT.
Approach: They define mixtext, a form of mixed text involving both AI and human-generated content, and then use a MixSet dataset to assess their effectiveness.
Outcome: The proposed detectors struggle to identify mixtext, particularly in dealing with subtle modifications and style adaptability.
Improving the Adversarial Robustness of NLP Models by Information Bottleneck (2022.findings-acl)

Copied to clipboard

Challenge: Existing studies have shown that adversarial examples can be directly attributed to the presence of non-robust features.
Approach: They propose to capture task-specific robust features while eliminating non-robust ones . they show that models can achieve significant improvement in robust accuracy .
Outcome: The proposed method outperforms all defense methods on SST-2, AGNEWS and IMDB datasets while achieving no performance drop.
Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue Systems (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models enable them to follow freeform instructions, including imitating generic or specific demographic personas in conversations.
Approach: They propose to investigate persona biases by experimenting with UNIVERSALPERSONA, a model that incorporates both generic and specific personas.
Outcome: The proposed model systematically measures persona biases in harmful expression and harmful agreement.
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Text-to-image retrieval is challenging because of cross-modal embeddings are bags of concepts, underrepresenting structured visual relationships.
Approach: They propose a retrieval paradigm that embeds textual queries into the image modality via T2I generation and performs retrieval within the image mode to bypass weaknesses of cross-modal retrievers in recognizing subtle visual-spatial features.
Outcome: The proposed retrieval paradigm outperforms previous approaches in visual-spatial retrieval benchmarks.
Knowledge Control for Responsible Generative AI: Bridging Academia, Industry, and Society (2026.acl-tutorials)

Copied to clipboard

Challenge: This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods.
Approach: This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods.
Outcome: This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods . key motivations and failure modes, harmful generation and stereotype reinforcement, are addressed . core methods such as machine unlearning, knowledge editing, and inference-time interventions are also included .
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs (2025.acl-long)

Copied to clipboard

Challenge: Social biases manifest in language agency, but there is no comprehensive benchmark for evaluating such biase in language models.
Approach: They propose a benchmark to evaluate language agency biases in large language models . they propose 'Mitigation via Selective Rewrite' to selectively revise parts of generated texts .
Outcome: The proposed language agency bias evaluation benchmark identifies gender, racial, and intersectional biases in 3 recent LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations