Papers by Yi Zhu

66 papers
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing large language model (LLM) agents are unable to adapt to changing domain knowledge and rules.
Approach: They propose an LLM agent framework that continuously learns updated domain knowledge at test time.
Outcome: The proposed agent improves on a customer due diligence name screening task on . the agent learns updated domain knowledge at test time.
SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks conflate coordination ability with role-based priors.
Approach: They propose a role-free benchmark for evaluating free-form collaboration under information silos.
Outcome: The proposed benchmark systematically probes coordination capabilities under information silos using 54 configurations and 3 frontier LLMs.
Slot Attention with Value Normalization for Multi-Domain Dialogue State Tracking (2020.emnlp-main)

Copied to clipboard

Challenge: Existing dialogue state tracking approaches rely on ontology already defined, where all slots and their possible values are given.
Approach: They propose a new architecture to exploit domain ontology by using Slot Attention and Value Normalization . they supplement the annotation of supporting span for MultiWOZ 2.1, which is the shortest span in utterances to support the labeled value.
Outcome: The proposed architecture exploits ontology and can convert supporting spans to values.
FanLoRA: Fantastic LoRAs and Where to Find Them in Large Language Model Fine-tuning (2024.emnlp-industry)

Copied to clipboard

Challenge: Lowrank adaptation and its variants introduce significant latency in multi-tenant settings, hindering their applications in the industry.
Approach: They propose a framework to fine-tune LoRA modules on a large-scale instruction tuning dataset.
Outcome: The proposed framework outperforms existing PEFT methods and significantly reduces inference latency.
Sequence Structure Aware Retriever for Procedural Document Retrieval: A New Dataset and Baseline (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing retrieval methods neglect the execution sequence structures inherent in procedural documents.
Approach: They propose a retrieval model which integrates procedural graphs with document representations.
Outcome: The proposed model integrates procedural graphs with document representations to improve document retrieval.
WhitenedCSE: Whitening-based Contrastive Learning of Sentence Embeddings (2023.acl-long)

Copied to clipboard

Challenge: Extensive experiments on seven semantic textual similarity tasks show our method achieves consistent improvement over the contrastive learning baseline and sets new states of the art.
Approach: They propose a whitening-based contrastive learning method for sentence embedding learning which combines contrastive and shuffled group whitening.
Outcome: The proposed method achieves better alignment and uniformity on seven semantic textual similarity tasks.
A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots Matters (2021.acl-long)

Copied to clipboard

Challenge: Few-shot crosslingual transfer outperforms zero-shot with pretrained encoders like multilingual BERT.
Approach: They conduct an experimental study on 40 sets of sampled few shots for six diverse NLP tasks across up to 40 languages.
Outcome: The proposed model outperforms state-of-the-art approaches on lexical features and a full model finetuning approach outperformed several state- of-the art approaches.
On the Vulnerability of Safety Alignment in Open-Access LLMs (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are susceptible to malicious exploitation, but are often rejected and limited harmfulness is limited.
Approach: They propose two types of reverse alignment techniques: reverse supervised fine-tuning (RSFT) and reverse preference optimization (RPO).
Outcome: The proposed methods can significantly enhance the success rate and harmfulness of jailbreak attacks, but they face high rejection rates and limited harmfulness.
JECC: Commonsense Reasoning Tasks Derived from Interactive Fictions (2023.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on a single reasoning type and ask human annotators to write candidate statements related to the particular type of commonsense.
Approach: They propose a new commonsense reasoning dataset based on human’s Interactive Fiction (IF) gameplaywalkthroughs.
Outcome: The proposed dataset is challenging to previous machine reading models and large language models with a significant 20%performance gap compared to human experts.
Knowledge Graph-Guided Retrieval Augmented Generation (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on RAG focus on semantic retrieval of isolated relevant chunks, which ignore their intrinsic relationships.
Approach: They propose a framework that utilizes knowledge graphs to provide fact-level relationships between chunks, improving the diversity and coherence of the retrieved results.
Outcome: Extensive experiments on the HotpotQA dataset and its variants demonstrate the advantages of KG2RAG compared to existing RAG-based approaches in terms of response quality and retrieval quality.
Chinese Lexical Substitution: Dataset and Method (2023.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for lexical substitution (LS) are limited and limited in coverage . despite extensive research on Lexical Substitution in various languages, there is limited evidence for LS in Chinese.
Approach: They propose to use human and machine collaboration to construct a Chinese LS dataset . they combine four unsupervised LS methods to generate candidate substitutes .
Outcome: The proposed method outperforms existing benchmarks on the Chinese lexical substitution task.
A Systematic Study of Leveraging Subword Information for Learning Word Representations (N19-1)

Copied to clipboard

Challenge: Existing word representation models for morphologically rich languages use subword-level information, but their systematic comparative analysis across typologically diverse languages and tasks is still missing.
Approach: They propose a framework for learning subword-informed word representations that allows for easy experimentation with different segmentation and composition components.
Outcome: The proposed framework allows for easy experimentation with different segmentation and composition components, as well as advanced techniques based on position embeddings and self-attention.
MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning (2024.findings-emnlp)

Copied to clipboard

Challenge: Low-rank adaptation and its mixture-of-experts (MOE) methods are highly effective but introduce significant latency in multi-tenant settings due to the LoRA modules and MOE routers added to multiple linear modules.
Approach: They propose a low-rank adaptation variant that considers each LoRA module as an expert and employs a prompt-aware routing mechanism.
Outcome: Extensive analysis on commonsense reasoning tasks and math reasoning tasks show that MiLoRA outperforms strong PEFT baselines with comparable tunable parameter budgets.
PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing reward models evaluate empathy from a single perspective, overlooking bidirectional interaction nature of empathy.
Approach: They propose a reward model that evaluates empathy from a single perspective . they propose PERM to integrate a bystander perspective to monitor overall interaction quality .
Outcome: a new reward model outperforms state-of-the-art models on an emotional intelligence benchmark and an industrial daily conversation dataset.
When Efficiency Becomes a Vulnerability: Computational Cost Attacks on WebAgents (2026.acl-long)

Copied to clipboard

Challenge: Existing WebAgents suffer from computational cost attacks due to long reasoning processes and excessive computational cost.
Approach: They propose a framework that generates adversarial prompts and a reinforcement learning-enhanced selector to identify the most effective perturbations.
Outcome: The proposed framework exploits large language models to generate diverse adversarial prompts and a reinforcement learning–enhanced selector to identify the most effective perturbations.
LIRE: listwise reward enhancement for preference alignment (2024.findings-acl)

Copied to clipboard

Challenge: prevailing approaches to preference alignment focus on pairwise comparisons, with limited exploration into multi-response scenarios.
Approach: They propose a listwise reward enhancement approach that integrates offline rewards of multiple responses into a streamlined listwise framework.
Outcome: The proposed approach outperforms existing methods on dialogue and summarization tasks with good transferability to out-of-distribution data.
Multi-Aspect Controllable Text Generation with Disentangled Counterfactual Augmentation (2024.acl-long)

Copied to clipboard

Challenge: Existing studies neglect attribute correlations formed by the intertwining of different attributes.
Approach: They propose a multi-aspect controllable text generation method with disentangled counterfactual augmentation that alleviates imbalanced attribute correlations during training by disentanglement.
Outcome: The proposed method outperforms state-of-the-art methods in imbalanced and balanced attribute correlation scenarios.
Reasoning under Uncertainty: Efficient LLM Inference via Unsupervised Confidence Dilution and Convergent Adaptive Sampling (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models suffer from overconfidence and computational inefficiency due to fixed computation budgets and miscalibrated confidence estimates.
Approach: They propose a framework for computationally efficient, trustworthy reasoning under uncertainty using Diversity-Aware Self-Signal Dilution and Convergent Adaptive Weighted Sampling techniques.
Outcome: The proposed framework reduces inference cost by 70% while maintaining accuracy levels while reducing inference costs.
Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for detecting pre-training data in large language models rely on superficial features like prediction confidence and loss, resulting in mediocre performance.
Approach: They propose a new algorithm to analyze neuron activation patterns between training and non-training data in large language models to improve their performance.
Outcome: The proposed algorithm outperforms existing methods across three benchmarks and multiple LLMs.
ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection (2026.findings-acl)

Copied to clipboard

Challenge: Audio deepfake detection systems do not generalize well to realistic in-the-wild deepfakkes.
Approach: They propose a novel In-Context Learning paradigm with comparison-guidance for Audio Deepfake detection framework that uses audio language models for training-free generalization to unseen deepfakes.
Outcome: The proposed framework improves macro F1 over specialized detectors on in-the-wild datasets with up to 2 relative improvement over existing models.
An Unsupervised Method for Building Sentence Simplification Corpora in Multiple Languages (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to build parallel sentence simplification corpora are limited . SS is used to rephrase sentences into simpler forms for those with cognitive disabilities .
Approach: They propose to build SS corpora from large-scale bilingual translation corpors using a parallel approach.
Outcome: The proposed method outperforms the existing methods on WikiLarge and achieves state-of-the-art results.
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to RPAs focus on static role profiles, overlooking dynamic perceptual abilities inherent to humans.
Approach: They propose a framework that combines adaptive temporal sampling with dynamic and static role profiles.
Outcome: The proposed framework combines adaptive temporal sampling with dynamic and static role profiles.
Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in multimodal pre-trained models have significantly improved information extraction from visually-rich documents (VrDs).
Approach: They propose a method to predict token sequences within visually-rich documents by a simple prediction head.
Outcome: The proposed method can be used to predict token mentions as token sequences within documents.
MTAG: Modal-Temporal Attention Graph for Unaligned Human Multimodal Language Sequences (2021.naacl-main)

Copied to clipboard

Challenge: a novel graph-based neural model for multimodal sequential data is proposed . fusion is the process of blending information from multiple modalities, usually preceded by alignment .
Approach: They propose a graph-based neural model that converts unaligned data into a modal-temporal graph . they use a dynamic pruning and read-out technique to efficiently process the graph fusion operation .
Outcome: The proposed model performs state-of-the-art on multimodal sentiment analysis and emotion recognition benchmarks while utilizing significantly fewer model parameters.
MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures - A Comprehensive Framework (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs)-based Multi-Agent Systems (MAS) exhibit remarkable problem-solving and task planning capabilities across diverse domains .
Approach: They propose a security research framework for LLM-based multi-agent systems . they propose corresponding defense strategies to address MAS security risks .
Outcome: The proposed framework amplifies the severity of security risks under MAS attacks . it offers an automated construction process for different MAS setups and an interaction paradigm .
Reinforcement Learning on Pre-Training Data (2026.acl-long)

Copied to clipboard

Challenge: Recent progress in large language models is driven by scaling of training compute through pre-training with nexttoken prediction (NTP) or post-training (RL) Pre-training using NTP enables models to acquire extensive knowledge and skills from general data, but it suffers from data inefficiency and catastrophic forgetting in continual learning settings.
Approach: They propose to scale training compute through pre-training with next-token prediction (NTP) or post-training by scaling reinforcement learning (RL) to improve learning from general data.
Outcome: Experiments on multiple benchmarks and models show that the proposed approach improves continual pre-training and provides a strong foundation for post-training on Qwen3-8B-Base.
Tailoring Instructions to Student’s Learning Levels Boosts Knowledge Distillation (2023.acl-long)

Copied to clipboard

Challenge: Recent success of natural language processing (NLP) is driven by the adoption of large-scale pretrained language models.
Approach: They propose a method to determine the impact of distillation influence on student generalization ability by prioritizing samples likely to enhance the student's generalization abilities.
Outcome: The proposed method outperforms 10 common knowledge distillation baselines on 6 text classification tasks in the GLUE benchmark.
Text Augmented Spatial Aware Zero-shot Referring Image Segmentation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing zero-shot referring image segmentation methods focus on global-level alignment of image-text pairs, neglecting fine-grained matching between referring sentence and local image regions.
Approach: They propose a zero-shot referring image segmentation task that is training-free . they use a mask proposal network and a text-augmented spatial-correction score .
Outcome: The proposed method outperforms state-of-the-art zero-shot referring image segmentation methods.
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in multi-modal learning have enhanced MLLMs' ability to reason about visual content.
Approach: They propose a framework that unifies multi-step multimodal reasoning with grounded visual understanding.
Outcome: The proposed framework surpasses state-of-the-art methods by +6.5 gIoU and +9.2 cIou on ReasonSeg and achieves 49.7 mAP on SegInW under zero-shot settings.
Guiding Medical Vision-Language Models with Diverse Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations (2025.naacl-long)

Copied to clipboard

Challenge: Current vision-language models lack the ability to focus on specific areas designated by humans . a new framework that integrates medical entity extraction, visual prompt generation, and dataset adaptation is proposed to improve visual prompt-guided fine-tuning.
Approach: They propose to use visual prompts to guide and enhance formation of region-specific attention.
Outcome: The proposed framework outperforms state-of-the-art large vision-language models on medical datasets.
TranSFormer: Slow-Fast Transformer for Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Prior work has focused on treating subwords as basic units in developing such systems.
Approach: They propose a slow-fast two-stream learning model that uses a “slow” branch to deal with subword sequences and a "fast" branch to cope with longer character sequences.
Outcome: The proposed model shows consistent BLEU improvements (larger than 1 BLUE point) on several machine translation benchmarks.
ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation (2022.acl-long)

Copied to clipboard

Challenge: Residual networks are an Euler discretization of solutions to Ordinary Differential Equations (ODE).
Approach: They propose a residual block of layers in Transformer that can be described as a higher-order solution to ODE.
Outcome: The proposed architecture can gain large improvements over strong baselines at a slight cost in inference efficiency.
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for agentic repository-level code understanding overlook long tail topics and rely on memorized knowledge.
Approach: They propose a repository-level agentic code understanding benchmark that uses long-tail repositories with executable environments to enforce topical balance.
Outcome: Empirically, a Qwen3-8B model trained with the proposed benchmark outperforms GPT-4o by 2.3 points.
ParaLS: Lexical Substitution via Pretrained Paraphraser (2023.acl-long)

Copied to clipboard

Challenge: Lexical substitution (LS) is an extremely powerful technology that can be used as a backbone of various NLP applications such as writing assistance.
Approach: They propose two simple decoding strategies that focus on the variations of the target word during decoding to generate substitutes from a paraphraser.
Outcome: The proposed methods outperform state-of-the-art LS methods based on pre-trained language models on three benchmarks.
Diversity, Density, and Homogeneity: Quantitative Characteristic Metrics for Text Collections (2020.lrec-1)

Copied to clipboard

Challenge: Existing descriptive statistics are inadequate to summarize text collections by quantitative measures.
Approach: They propose a set of characteristic metrics that quantitatively measure the dispersion, sparsity, and uniformity of a text collection.
Outcome: The proposed metrics are highly correlated with text classification performance of a renowned model, which could inspire future applications.
Combining Deep Generative Models and Multi-lingual Pretraining for Semi-supervised Document Classification (2021.eacl-main)

Copied to clipboard

Challenge: Semi-supervised learning and multilingual pretraining have been shown to be effective for task-specific labelled data shortages.
Approach: They propose to combine semi-supervised deep generative models and multi-lingual pretraining to form a pipeline for document classification task.
Outcome: The proposed method outperforms state-of-the-art models in low-resource settings across several languages and outperformed existing models in English.
RelCLIP: Adapting Language-Image Pretraining for Visual Relationship Detection via Relational Contrastive Learning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing visual relationship detection models only use numeric ids of relation labels for training, but ignore semantic correlation between labels.
Approach: They propose a visual Relationship prediction framework that transfers natural language knowledge from Contrastive Language-Image Pre-training models to enhance the relationship prediction.
Outcome: The proposed framework improves visual relationship prediction by matching semantic correlations with relation triplets.
FaStFact: Faster, Stronger Long-Form Factuality Evaluations in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Prior evaluation pipelines fail to evaluate factuality of long-form LLMs due to inefficiency and costly human assessment.
Approach: They propose a fast and strong evaluation pipeline that can evaluate factuality of long-form LLMs . they propose 'faStFact' to reduce cost of web searching and inference calling .
Outcome: The proposed evaluation pipeline achieves highest alignment with human evaluation and efficiency among existing baselines.
Chinese Idiom Paraphrasing (2023.tacl-1)

Copied to clipboard

Challenge: Chinese idioms are hard to understand by children and non-native speakers due to their non-compositionality and metaphorical meaning.
Approach: They propose a task to rephrase idiom-containing sentences to non-idiomatic ones under the premise of preserving the original sentence’s meaning.
Outcome: The proposed method has better performance than baselines based on the established dataset.
Grounded Multimodal Procedural Entity Recognition for Procedural Documents: A New Dataset and Baseline (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to extract procedural knowledge from documents focus on text-only settings, which is insufficient for entity disambiguation.
Approach: They propose a model to detect the entity and the corresponding bounding box groundings in images.
Outcome: The proposed model detects the entity and the corresponding bounding box groundings in image (i.e., visual entities) it is based on a dataset of a WikiHow 1 and EHow 2 document and the results are compared with existing models.
Improved Knowledge Distillation for Pre-trained Language Models via Knowledge Selection (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on knowledge distillation have shown that not all knowledge is necessary for learning a good student model.
Approach: They propose an actor-critic approach to selecting appropriate knowledge to transfer during the process of knowledge distillation.
Outcome: The proposed method outperforms several strong knowledge distillation baselines significantly on the GLUE datasets.
Abstract then Play: A Skill-centric Reinforcement Learning Framework for Text-based Games (2023.findings-acl)

Copied to clipboard

Challenge: Existing reinforcement learning frameworks fail to decompose the task and abstract the action autonomously.
Approach: They propose a skill-centric reinforcement learning framework capable of abstracting the action in an end-to-end manner.
Outcome: Empirical experiments on the Jericho environment validate the proposed framework against state-of-the-art baselines.
Bayesian Learning for Neural Dependency Parsing (N19-1)

Copied to clipboard

Challenge: Several approaches for dependency parsing in the small data regime have been proposed.
Approach: They propose to use stochastic gradient Langevin dynamics to generate samples from the approximated posterior to overcome the computational and statistical costs of the approximate inference step.
Outcome: The proposed model outperforms the biaffine model on 6 languages with less than 5k training instances and improves across five languages.
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models of layout reading order do not convey the complete reading order information in the layout.
Approach: They propose to model layout reading order as ordering relations over layout elements . they propose a reading-order-relation-enhancing pipeline to improve model performance .
Outcome: The proposed model outperforms existing models on a visual-rich document dataset and on eight cross-domain VrD-IE/QA tasks without targeted optimization.
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for visual storytelling ignore latent topic information.
Approach: They propose a topic-aware reinforcement network for VIsual StoryTelling that takes topic information into account to generate a coherent story.
Outcome: The proposed method outperforms most of the competing models across multiple evaluation metrics.
Joint Modeling of Structure Identification and Nuclearity Recognition in Macro Chinese Discourse Treebank (C18-1)

Copied to clipboard

Challenge: Discourse parsing is a challenging task and plays a critical role in discourse analysis.
Approach: They propose a macro discourse structure presentation schema to present the macro level discourse structure analysis.
Outcome: The proposed corpus is based on two tasks of macro discourse structure analysis, including structure identification and nuclearity recognition.
Gloss-Free End-to-End Sign Language Translation (2023.acl-long)

Copied to clipboard

Challenge: a study of sign language translation without gloss annotations focuses on the problem of gloss annotation . gloss annotation is hard to acquire, especially in large quantities, and limits the domain coverage of translation datasets .
Approach: They propose a gloss-free end-to-end sign language translation framework to solve this problem . gloss annotations are hard to acquire, especially in large quantities, they argue .
Outcome: The proposed framework improves sign language translation performance on large-scale datasets . gloss annotations are hard to acquire, especially in large quantities .
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in vision-language-action models prioritize robotic action mastery . however, models trained on visual-text pairs struggle to interpret multimodal data .
Approach: They propose a framework that integrates multimodal data after initial control mastery and a Mixture-of-Experts architecture to minimize task interference.
Outcome: The proposed framework surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks and achieves six times higher performance on visual question-answering datasets.
Recognizing Conflict Opinions in Aspect-level Sentiment Classification with Dual Attention Networks (D19-1)

Copied to clipboard

Challenge: Existing models ignore conflict opinions because they are sparse in the datasets.
Approach: They propose a multi-label classification model with dual attention mechanism to address these problems by excluding conflict opinions from existing models.
Outcome: The proposed model addresses the problem of exclusion of conflict opinions from the datasets.
PARA: Parameter-Efficient Fine-tuning with Prompt-Aware Representation Adjustment (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for parameter-efficient fine-tuning excel in the context of single-backbone multi-tenant applications.
Approach: They propose to integrate a lightweight vector generator within each Transformer layer to improve prompt-aware representation adjustment.
Outcome: The proposed method surpasses current benchmarks in terms of performance despite having a similar number of adjustable parameters.
DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing ESC data entangles psychological strategies and response content, making it difficult to construct high-quality preference pairs.
Approach: They propose a Decoupled ESC framework that decomposes the ESC task into two sequential subtasks: strategy planning and empathic response generation.
Outcome: The proposed framework outperforms baselines, reducing preference bias and improving response quality.
VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in Minecraft (2024.findings-acl)

Copied to clipboard

Challenge: Multi-agent collaboration using LLMs is a challenging research topic that aims to enable multiple autonomous agents to coordinate their actions and achieve a common goal.
Approach: They propose a benchmark for multi-agent collaboration in the Minecraft environment and introduce a Directed Acyclic Graph Multi-Agent Framework to resolve complex inter-ag dependencies.
Outcome: The proposed framework outperforms existing ModelVerse, reducing hallucinations and improving task decomposition efficacy.
MR-ALIGN: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models (2026.findings-acl)

Copied to clipboard

Challenge: Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited.
Approach: They propose a Meta-Reasoning informed alignment framework that quantifies state-transition probabilities along the model’s thinking process and constructs a transition-aware implicit reward that reinforces beneficial reasoning patterns while suppressing defective ones at the atomic thinking segments.
Outcome: Empirical evaluations of four factual QA datasets and one long-form factuality benchmark show that MR-ALIGN consistently improves accuracy and truthfulness while reducing misleading reasoning.
Collaborative Document Simplification Using Multi-Agent Systems (2025.coling-main)

Copied to clipboard

Challenge: Document simplification requires complex factors such as technical terminology, metaphors, and overall coherence.
Approach: They propose a multi-agent framework for document simplification based on large language models that emulates the collaborative process of a human expert team through the roles played by multiple agents.
Outcome: The proposed framework emulates the collaborative process of a human expert team through the roles played by multiple agents, addressing the intricate demands of document simplification.
Post-Hoc Watermarking for Robust Detection in Text Generated by Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for document simplification address complex factors such as technical terminology, metaphors, and overall coherence.
Approach: They propose a multi-agent framework AgentSimp for document simplification based on large language models that simulates collaboration among agents through roles played by multiple agents.
Outcome: The proposed framework produces simplified documents that are more thoroughly simplified and more coherent across various articles and styles.
Learning the Beauty in Songs: Neural Singing Voice Beautifier (2022.acl-long)

Copied to clipboard

Challenge: Existing techniques for pitch correction are limited to intonation but ignore the overall aesthetic quality.
Approach: They propose a novel time-warping approach for pitch correction to synchronize the amateur recording with the template pitch curve.
Outcome: The proposed model improves intonation and vocal tone while keeping content and vocal timbre.
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in preference alignment have significantly improved Large Language Models' ability to generate texts that align with human preferences and values.
Approach: They propose a constrained optimization approach to detect and mitigate update regression with focal attention.
Outcome: The proposed approach detects and mitigates update regression with focal attention while maintaining excellent overall performance.
OAgents: An Empirical Study of Building Effective Agents (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that agent research practices are far from standard, rigorous . lack of a standard evaluation protocol makes previous works not reproducible, authors say .
Approach: They conduct an empirical study on the GAIA benchmark to investigate agent design choices . they find that lack of a standard evaluation protocol makes previous works not reproducible .
Outcome: The proposed framework achieves state-of-the-art performance among open-source projects.
Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decomposition (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for knowledge editing fail to work in multi-hop question answering due to 'edit skipping' edit skipping occurs due to the mismatch between the granularity of LLMs in problem-solving and the facts in the edited memory.
Approach: They propose a retrieval-augmented generation-based method that edits knowledge without modifying parameters without retraining LLMs.
Outcome: The proposed method outperforms state-of-the-art methods for KE in multi-hop question answering.
Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) suffer catastrophic forgetting when tailored to specific domains . authors present a novel approach to manage multi-domain LLM adaptation .
Approach: They propose a strategy to manage multi-domain LLM adaptation using self-distillation and role integration.
Outcome: The proposed model alleviates catastrophic forgetting and inter-domain confusion while maintaining robust general capabilities.
Seq2Path: Generating Sentiment Tuples as Paths of a Tree (2022.findings-acl)

Copied to clipboard

Challenge: Existing generative methods for extracting sentiment tuples do not have orders between the t-uples . a novel parallel generative framework for ABSA is proposed .
Approach: They propose a parallel generative framework to generate sentiment tuples as paths of a tree . they train the model with an independent target and introduce a discriminative token .
Outcome: The proposed method achieves state-of-the-art on AOPE, ASTE, TASD, UABSA, ACOS . it trains with the loss of ordinary Seq2Seq averaged over paths, and inferences automatically select valid paths.
Parsing Tweets into Universal Dependencies (N18-1)

Copied to clipboard

Challenge: a new tweet treebank for English is designed to analyze tweets with universal dependencies (UD).
Approach: They extend the universal dependencies guidelines to include special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies.
Outcome: The proposed method outperforms state-of-the-art parsers on other treebanks in accuracy and speed.
AI4Reading: Chinese Audiobook Interpretation System Based on Multi-Agent Collaboration (2025.acl-demo)

Copied to clipboard

Challenge: Interpretative audiobooks are becoming more popular, but their manual creation process remains time-consuming and resource-intensive.
Approach: They propose a multi-agent collaboration system that leverages large language models and speech synthesis technology to generate podcast-like audiobook interpretations.
Outcome: The proposed system is open source and open to the public.
FragRel: Exploiting Fragment-level Relations in the External Memory of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches to process contexts with unlimited length are limited to finite expansion length or prone to performance degradation when dealing with very long contexts.
Approach: They propose to exploit fragment-level relations in external memory to hierarchically process the long text.
Outcome: The proposed model improves story understanding, repository-level code generation, and long-term chatting.
Chinese Live-Streaming E-Commerce Morph Resolution: Datasets and Methods (2026.findings-acl)

Copied to clipboard

Challenge: Live-stream E-commerce faces significant challenges from morphs, deliberate linguistic variants used to evade real-time voice filters and amplify product claims illegally.
Approach: They propose a framework that resolves morphs and generates structured explanations . they propose morph-aware dual-output refinement framework that detects inconsistencies .
Outcome: The proposed framework improves morph resolution accuracy and interpretability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations