Papers by Donglei Yu

5 papers
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions (2025.acl-long)

Copied to clipboard

Challenge: Existing research addresses ambiguous visual questions by rephrasing questions, but it fails to address the inherently interactive nature of user interactions with visual language models (VLMs). Existing studies focus on re-phrase questions, and lack of a benchmark to assess VLMs’ capacity for resolving ambiguities through interaction.
Approach: They propose a visual question answering task that provides a natural language answer to a question based on a given image and an automated pipeline to generate ambiguity-clarification question pairs.
Outcome: The proposed benchmark targets three common categories of ambiguity in visual question answering (VQA) context and encompasses various VQA scenarios.
Fixing Distribution Shifts of LLM Self-Critique via On-Policy Self-Play Training (2025.acl-long)

Copied to clipboard

Challenge: Large language models show impressive performance in a wide range of linguistic tasks, but their performance on complex reasoning tasks is still signif-icantly lower than the human level.
Approach: They propose a reinforcement learning framework to synchronize the reasoning and critique capabilities of language models by using Monte Carlo sampling to give appropriate rewards to the model's critique content.
Outcome: The proposed framework improves the model's reasoning and critique capabilities by 5.40 and 3.66 points, respectively, compared to the best baseline approach.
Investigating Hallucinations in Simultaneous Machine Translation: Knowledge Distillation Solution and Components Analysis (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods to mitigate hallucinations in siMT generate fluency but unfaithful translation.
Approach: They propose a method that utilizes the OMT model to mitigate hallucinations in SiMT.
Outcome: The proposed method reduces hallucinations and improves the SiMT performance.
Self-Modifying State Modeling for Simultaneous Machine Translation (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for simultaneous machine translation fail to optimize the policy . existing methods require building a decision path to learn the policy, but they cannot explore all potential paths .
Approach: They propose a new training paradigm that uses a read/write policy to optimize the policy . existing methods usually require building a decision path to learn a suitable policy a user makes .
Outcome: The proposed model outperforms strong baselines and allows offline models to acquire SiMT ability with fine-tuning.
Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQA (2024.emnlp-main)

Copied to clipboard

Challenge: Existing visual language models struggle to capture longtail knowledge in the real world due to redundant visual information.
Approach: They propose a method leveraging the reasoning capability of a large language model to identify key visual entities.
Outcome: The proposed method outperforms other strong visual language model-based systems in two knowledge-intensive VQA benchmarks and performs comparably to models with 1-2 orders larger parameters.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations