Papers by Mingyang Zhou

21 papers
Learning to Plan for Retrieval-Augmented Large Language Models from Knowledge Graphs (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have attempted to enhance the performance of large language models (LLMs) in complex question-answering (QA) tasks by combining step-wise planning with external retrieval.
Approach: They propose a framework for enhancing LLMs’ planning capabilities by using planning data derived from knowledge graphs (KGs).
Outcome: The proposed framework improves LLMs’ planning capabilities by using knowledge graphs (KGs) the proposed framework is compared with existing frameworks on multiple datasets and shows that it is effective for large language models.
M2-TabFact: Multi-Document Multi-Modal Fact Verification with Visual and Textual Representations of Tabular Data (2025.findings-acl)

Copied to clipboard

Challenge: Existing fact-checking systems that can reason over structured data are inefficient compared to humans.
Approach: They propose a multi-modal table-based fact verification task that requires reasoning over visual and textual representations of structured data.
Outcome: The proposed model can reason over visual and textual representations of structured data.
A Joint Learning Framework for Restaurant Survival Prediction and Explanation (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances in deep learning have various models that research reviews and interactions for different kinds of tasks, such as predicting restaurant survival.
Approach: They propose a joint learning framework for explainable restaurant survival prediction based on multi-modal data of user-restaurant interactions and users’ textual reviews.
Outcome: The proposed framework improves on two datasets showing that it can model restaurant interactions and users’ textual reviews.
Knowing-but-Doing: Diagnosing and Defending Role-Play-Driven LLMs Jailbreaks via Moral Disengagement (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in role-play scenarios, but their safety implications remain under-characterized.
Approach: They propose a diagnostic benchmark for role-play jailbreaks based on Bandura’s Moral Disengagement theory and propose 'MD-Trace' based defense that reduces attack success while maintaining Role Fidelity.
Outcome: The proposed framework improves safety behavior for benign personas while increasing unsafe compliance for malicious ones.
Eliminating Out-of-Domain Recommendations in LLM-based Recommender Systems: A Unified View (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to reduce OOD recommendations fall into three grounding paradigms: retrieval, constrained generation and discrete item tokenizer generation.
Approach: They propose a framework that instantiates three grounding paradigms under a single architecture . embedding-based retrieval, constrained generation and discrete item-tokenizer methods are implemented .
Outcome: The proposed framework eradicates OOD recommendations across all variants and achieves state-of-the-art accuracy compared to strong ID-based and LLM-based baselines.
Pretraining Context Compressor for Large Language Models with Embedding-Based Memory (2025.acl-long)

Copied to clipboard

Challenge: Efficient processing of long contexts in large language models is essential for real-world applications such as retrieval-augmented generation and in-context learning.
Approach: They propose a decoupled compressor-LLM framework that preserves contextual information within condensed embedding representations.
Outcome: The proposed framework outperforms baseline models in three domains and across eight datasets while adapting to different downstream LLMs.
A Visual Attention Grounding Neural Model for Multimodal Machine Translation (D18-1)

Copied to clipboard

Challenge: Existing approaches to multimodal machine translation do not integrate visual information into the translation process.
Approach: They propose a multimodal machine translation model that utilizes parallel visual and textual information.
Outcome: The proposed model outperforms existing methods on the Multi30K and Ambiguous COCO datasets.
Building Task-Oriented Visual Dialog Systems Through Alternative Optimization Between Dialog Policy and Language Generation (D19-1)

Copied to clipboard

Challenge: Current approaches to visual dialog learning involve an end-to-end framework that maps the multi-modal context to a deep vector and in order to decode a natural dialog response.
Approach: They propose a framework that trains a RL policy for image guessing and a seq2seq model to improve dialog quality.
Outcome: The proposed framework achieves state-of-the-art performance on a guessWhich task . it can be applied to a wide range of tasks including assisting blind people .
PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models (2025.acl-long)

Copied to clipboard

Challenge: Recent large language models (LLMs) have achieved significant performance in complex reasoning tasks such as mathematics and code generation.
Approach: They propose a process-level benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs.
Outcome: The proposed model measures the accuracy, soundness, and sensitivity of 25 models across open-source and closed-source large language models.
Training Verifier to Assessing Complex Real-World Tool-Use Trajectories (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for training effective AI agents often resort to synthetic data generation.
Approach: They propose a plug-and-play framework for data quality control in tool-use scenarios . they construct a tool-verify dataset and release a benchmark to assess its performance .
Outcome: The proposed framework surpasses Qwen2.5-72B-Instruct on Tool-V-Bench and the previous APIGen-MT dataset.
Optimizing Length Compression in Large Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Large Reasoning Models suffer from producing unnecessary and verbose reasoning chains.
Approach: They propose a post-training method that uses a Length Reward and a Compress Reward to remove the invalid portion of the thinking process.
Outcome: The proposed method reduces sequence length by 50% with only a marginal (2%) drop in accuracy.
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to learn internal world models rely on one-step supervision . however, standard MTP suffers from structural hallucinations .
Approach: They propose a method which anchors predictions to ground-truth hidden state trajectories.
Outcome: The proposed method bridges the gap between discrete tokens and continuous state representations, reducing structural hallucinations, and improving robustness to perturbations.
Gunrock: A Social Bot for Complex and Engaging Long Conversations (D19-3)

Copied to clipboard

Challenge: Gunrock is a speech-based social chatbot that can be used to understand complex sentences and have in-depth conversations.
Approach: They propose a system that allows users to understand complex sentences and have in-depth conversations in open domains.
Outcome: The proposed system produces longer sentences, which are directly related to user engagement (e.g., ratings, number of turns).
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing research focuses on character-level settings and static evaluation formats fail to capture the complexity of everyday social interactions.
Approach: They propose a dynamic simulation framework for evaluating and improving persona-level role-playing in large language models (LLMs).
Outcome: The proposed framework leverages user-generated social content to construct a nuanced persona bank and elicits multi-turn, context-rich interactions within simulated social environments.
Hierarchical Reward Modeling for Fault Localization in Large Code Repositories (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have limited fault localization capabilities due to limited context length.
Approach: They propose a hierarchical localization reward model to evaluate and select the most accurate fault localization candidates from the outputs of LLMs.
Outcome: The proposed model improves the final line-level localization recall by 12% on the SWE-Bench-Lite dataset.
Explainable Recommendation with Personalized Review Retrieval and Aspect Learning (2023.acl-long)

Copied to clipboard

Challenge: Recent years have witnessed a growing interest in the development of explainable recommendation models.
Approach: They propose a model that combines prediction and generation tasks to produce more persuasive explanations by obtaining additional information from the training sets.
Outcome: The proposed model outperforms state-of-the-art models on three datasets and shows that it is more persuasive than previous models.
Focus! Relevant and Sufficient Context Selection for News Image Captioning (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent work only coarsely leverages the article to extract the necessary context, which makes it difficult for models to identify relevant events and named entities.
Approach: They propose to use a vision and language retrieval model CLIP to localize the visually grounded entities in the news article and then capture the non-visual entities via an open relation extraction model.
Outcome: The proposed model significantly improves on existing models and achieves state-of-the-art on multiple benchmarks.
Aligning Large Language Models for Controllable Recommendations (2024.acl-long)

Copied to clipboard

Challenge: Existing literature focuses on integrating domain-specific knowledge into LLMs to enhance accuracy using a fixed task template.
Approach: They propose a collection of supervised learning tasks augmented with labels derived from a conventional recommender model to improve LLMs’ proficiency in adhering to recommendation-specific instructions.
Outcome: The proposed approach significantly improves the capability of LLMs to respond to instructions within recommender systems, reducing formatting errors while maintaining a high level of accuracy.
R-CHAR: A Metacognition-Driven Framework for Role-Playing in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing role-playing structures lack cognitive consistency in complex scenarios . Existing models excel in math and coding tasks but lack coherent reasoning .
Approach: They propose a metacognition-driven framework that enhances role-playing performance . experimental results show performance improvements across varying scenario complexities .
Outcome: The proposed framework outperforms existing models in social intelligence tasks and shows strength in long-context comprehension and group-level social interactions.
Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning (2024.findings-acl)

Copied to clipboard

Challenge: LVLMs are known for producing text that is factually inconsistent with visual input . factuality of generated captions for structured visuals has not been studied as much .
Approach: They propose a typology of factual errors in captions generated by large vision-language models . they propose CHOCOLATE, a visual entailment model that outperforms current models based on this analysis .
Outcome: The proposed model outperforms current models in evaluating caption factuality.
Enhanced Chart Understanding via Visual Language Pre-training on Plot Table Pairs (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to understand chart plots are difficult to apply to visual-language tasks.
Approach: They propose a V+L model that learns how to interpret table information from chart images via cross-modal pre-training on plot table pairs.
Outcome: The proposed model outperforms state-of-the-art models on the chartQA benchmark by over 8% performance gains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations