Papers by Wen Wu

65 papers
Evolving Beyond Snapshots: Harmonizing Structure and Sequence via Entity State Tuning for Temporal Knowledge Graph Forecasting (2026.acl-long)

Copied to clipboard

Challenge: Temporal knowledge graphs (TKGs) require predicting future facts by modeling structural dependencies within each snapshot and temporal evolution across snapshots.
Approach: They propose an encoder-agnostic framework that provides persistent entity states . EST maintains a global state buffer and aligns structural evidence with sequential signals .
Outcome: Experiments show that EST improves diverse backbones and achieves state-of-the-art performance.
Towards Self-Evolving Agents: Enabling Autonomy through Interactive Experience Refinement (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models struggle with complex, multi-step operational tasks because they remain static during inference and cannot learn from past experience.
Approach: They propose a framework that organizes cross-domain insights to facilitate orchestration of long-horizon workflows.
Outcome: The proposed framework outperforms existing methods on the TAC productivity benchmark and shows strong cross-task transferability.
From Selection to Refinement: Iterative Optimization for Instruction Data (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to optimize instruction tuning datasets face two main challenges: unreasonable pruning of potentially valuable low-quality data and the persistence of noise or semantic drift during revision.
Approach: They propose an automated iterative framework for instruction data optimization that prunes low-quality data and refines low quality data using feedback-driven iteration.
Outcome: The proposed framework outperforms state-of-the-art methods on seven public benchmark datasets with high data efficiency.
Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to predict human emotions are inconsistent due to complexity of emotion and subjectivity of perception.
Approach: They propose a Bayesian approach to estimate uncertainty in emotion attributes using a deep neural network model.
Outcome: The proposed approach estimates uncertainty in emotion attributes along with aleatoric and epistemic uncertainties.
Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language Models (2024.acl-long)

Copied to clipboard

Challenge: Despite the advances in large language models, they still face difficulties with multi-step reasoning tasks.
Approach: They propose a method that randomly masks certain tokens within the chain of thought to improve model accuracy by 5% over standard supervised fine-tuning.
Outcome: The proposed method improves accuracy and accuracy by 5% over standard fine-tuning with a few codes modified.
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmented Generation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing research focuses on single-turn RAG, leaving a gap in addressing multi-turn conversations . a new benchmark is designed to assess RAG systems in realistic multi-turned conversations based on Wikipedia .
Approach: They propose a large-scale benchmark to assess RAG systems in multi-turn contexts . CORAL includes diverse information-seeking conversations automatically derived from Wikipedia . authors propose unified framework to standardize various conversational RAG methods .
Outcome: The proposed framework supports three core tasks of conversational RAG: passage retrieval, response generation, and citation labeling.
Calibration-Aware Policy Optimization for Reasoning LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to model calibration are limited or sacrifice gains in reasoning accuracy.
Approach: They propose a method that improves calibration by 15% while boosting accuracy by 5% . they propose GRPO-style algorithms that misalign uncertainty-agnostic advantage estimation .
Outcome: The proposed approach improves calibration by 15% while achieving comparable to or better than GRPO on multiple mathematical reasoning benchmarks.
COM2SENSE: A Commonsense Reasoning Benchmark with Complementary Sentences (2021.findings-acl)

Copied to clipboard

Challenge: Recent advances in pretrained language models have shown promising results on commonsense reasoning benchmark datasets.
Approach: They propose a commonsense reasoning benchmark dataset with 4k sentence pairs . they propose 'gamified' model-in-the-loop setup to incentivize challenging samples .
Outcome: The proposed benchmarks show that the proposed model achieves 71% standard accuracy and 51% pairwise accuracy, well below human performance.
Defending against Indirect Prompt Injection by Instruction Detection (2025.findings-emnlp)

Copied to clipboard

Challenge: Indirect Prompt Injection attacks can be exploited by LLMs that are embedded with external data.
Approach: They propose a detection-based approach that leverages the behavioral states of LLMs to identify potential IPI attacks.
Outcome: The proposed approach reduces the success rate of attacks to 0.03% on the BIPIA benchmark.
ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are limited by knowledge cutoff and can generate factual hallucinations when handling time-sensitive news.
Approach: They propose a two-stage zero-shot fake news detection framework that uses a hierarchical salience and saliency-calibrated minimum margin of relevance algorithm to extract core entities accurately.
Outcome: The proposed framework outperforms existing zero-shot baselines and even most few-shot methods on two public datasets.
TOME: A Two-stage Approach for Model-based Retrieval (2023.acl-long)

Copied to clipboard

Challenge: Recent research has focused on model-based retrieval, which discards the index in the traditional retrieval model and memorizes the candidate corpora using model parameters.
Approach: They propose a model-based retrieval approach that discards the index in the traditional retrieval model and memorizes the candidate corpora using model parameters.
Outcome: The proposed approach eliminates the index in the traditional retrieval model and memorizes the candidate corpora using model parameters.
CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-Checking (2022.naacl-main)

Copied to clipboard

Challenge: CHEF dataset provides evidence retrieval over non-English claims . e-fact-checking is a time-consuming task, which can take journalists several hours or days.
Approach: They construct a dataset of 10K real-world claims that is based on annotated evidence retrieved from the Internet.
Outcome: The proposed dataset provides evidence retrieval as a latent variable and can be used to train and reason over non-English claims.
Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Language Models with Combination of Regression and Ranking (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent work on sentence prediction tasks uses shallow neural networks to learn essay representations and constrain calculated scores with regression loss or ranking loss.
Approach: They propose to use a pre-trained language model to learn text representations first and then to constrain the scores with regression loss or ranking loss.
Outcome: The proposed model outperforms state-of-the-art models on the Automated Student Assessment Prize dataset.
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting in errors in visually grounded tasks and hallucinations.
Approach: They propose a novel finetuning objective that steers the model toward capturing important visual details and aligning them with corresponding text tokens.
Outcome: The proposed method achieves up to 22% reduction in hallucinations and significant gains in vision-centric and general tasks while maintaining or improving the model's general abilities.
Two-Stage Regularization-Based Structured Pruning for LLMs (2026.acl-long)

Copied to clipboard

Challenge: Structural pruning is a promising solution for large language models . prior structured pruning methods remove unimportant parameters based on certain metrics .
Approach: They propose a structural pruning method that iteratively learns the weights of transformer layers by adding their l1-norm to the loss function.
Outcome: The proposed pruning method outperforms strong layer-wise pruning methods without requiring retraining.
FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to correct factually inaccurate outputs are lacking the semantic richness needed to properly understand its internal states of trustworthiness and honesty.
Approach: They propose a framework for factuality alignment that integrates natural-language uncertainty signals with external knowledge and computes confidence scores and semantic entropy from LLM outputs.
Outcome: Extensive experiments on four knowledge-intensive benchmarks show that FAITH improves the factual accuracy and truthfulness of Large Language Models (LLMs).
Multi-Programming Language Sandbox for LLMs (2025.acl-demo)

Copied to clipboard

Challenge: MPLSandbox is an out-of-the-box multi-programming language sandbox designed to provide unified and comprehensive feedback from compiler and analysis tools for Large Language Models (LLMs).
Approach: They propose a multi-programming language sandbox that provides unified feedback from compilers and analysis tools for Large Language Models.
Outcome: The proposed multi-language sandbox can provide comprehensive feedback from compilers and analysis tools for large language models (LLMs).
SPARK: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for training large language models waste computation budget on trivial steps while failing to guarantee sample quality.
Approach: They propose a framework that selectively branches at critical decision states for resource-efficient exploration.
Outcome: The proposed framework activates adaptive branching exploration at critical decision states to probe promising trajectories, thereby achieving precise resource allocation that prioritizes sampling quality over blind coverage.
TrendSim: Simulating Trending Topics in Social Media Under Poisoning Attacks with LLM-based Multi-agent System (2025.findings-naacl)

Copied to clipboard

Challenge: Trending topics bring in a new channel for poisoning attacks, resulting in negative impacts on society.
Approach: They propose an LLM-based multi-agent system to simulate trending topics in social media . they propose a time-aware interaction mechanism, centralized message dissemination, and an interactive system .
Outcome: The proposed system simulates trending topics under poisoning attacks on social media platforms.
Handling Ambiguity in Emotion: From Out-of-Domain Detection to Distribution Estimation (2024.acl-long)

Copied to clipboard

Challenge: Experimental results show that incorporating utterances without majority-agreed labels into an additional class reduces the classification performance of the other emotion classes.
Approach: They propose to combine utterances without majority-agreed labels into an additional class . they propose to quantify uncertainty in emotion classification using evidential deep learning .
Outcome: The proposed method retains classification accuracy while effectively detects ambiguous emotion expressions.
Virtual Data Augmentation: A Robust and General Framework for Fine-tuning Pre-trained Models (2021.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that powerful pre-trained language models can be fooled by small perturbations or intentional attacks.
Approach: They propose a framework for fine-tuning PLMs using a masked language model and Gaussian noise to augment semantically relevant examples with sufficient diversity.
Outcome: The proposed framework improves the robustness of pre-trained language models and alleviates performance degradation under adversarial attacks.
A Multi-Modal Knowledge Graph for Classical Chinese Poetry (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies in classical Chinese poetry area focus on generation and analysis of poetry.
Approach: They propose to integrate the visual information of words in classical Chinese poetry into a multi-modal knowledge graph.
Outcome: The proposed model bridges the semantic gap between two modalities and achieves state-of-the-art performance on the poetry-image retrieval task.
Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing Diffusion Language Models rely on hard binary masking and discrete token assignments, which hinder the revision of early decisions.
Approach: They propose a diffusion-based language modeling approach that replaces hard binary masks with evolving soft token distributions.
Outcome: The proposed approach outperforms existing DLMs on multiple benchmarks.
PAIR: Leveraging Passage-Centric Similarity Relation for Improving Dense Passage Retrieval (2021.findings-acl)

Copied to clipboard

Challenge: Recent studies only consider query-centric similarity relation when learning the dual-encoder retriever.
Approach: They propose a query-centric and PAssage-centric approach to capture more comprehensive similarity relations for dense passage retrieval.
Outcome: The proposed approach significantly outperforms existing models on both MSMARCO and Natural Questions datasets.
Breaking the Evaluation Paradox: Evaluating High-Entropy Search with Computationally Irreducible Constraints (2026.findings-acl)

Copied to clipboard

Challenge: a new framework for evaluation of exhaustive search capabilities is needed . high-entropy enumeration tasks make such ground truth impossible for humans to create . VERITAS is a framework built on the principle of computationally irreducible constraints .
Approach: They propose a framework that uses non-optimizable constraints to create verifiable searches . VERITAS can generate infinite number of test cases with perfect ground truth and precise difficulty control .
Outcome: a new evaluation framework for large language models is based on non-optimizable constraints . the framework can generate infinite number of test cases with perfect ground truth and precise difficulty control .
Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration (2026.acl-long)

Copied to clipboard

Challenge: Non-sequential and bidirectional nature of diffusion large language models makes direct likelihood-based self-evaluation challenging.
Approach: They propose a self-evaluation confidence quantification method for diffusion large language models that quantifies confidence by computing the probability of regenerating tokens in the entire generated sequence, given the full context.
Outcome: The proposed method is correlated with semantic coherence and answer accuracy.
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model (2025.emnlp-demos)

Copied to clipboard

Challenge: Empathetic speech models are increasingly closed off, leaving details about the architecture, data and development opaque to researchers.
Approach: They propose an open-source empathetic speech-to-text model with a streaming interleaved decoding architecture and a data pipeline to enable end-to end training.
Outcome: The proposed model is open-source and transparent, with no data or data required to build it.
RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing (2025.findings-emnlp)

Copied to clipboard

Challenge: Current routing methods are limited in exploring the connection between query and LLM characteristics.
Approach: They propose a framework for LLM routing that uses a transformer-based backbone and a radial structure to articulate the query-LLMs relationship.
Outcome: The proposed framework outperforms existing routing methods by 9.2% and 5.8% on RouterBench.
Glance-or-Gaze: Incentivizing LMMs to Adaptively Focus Search via Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing search-augmented approaches rely on indiscriminate whole-image retrieval and lack deep iterative reflection, limiting their effectiveness on complex visual queries.
Approach: They propose a fully autonomous framework that shifts from passive perception to active visual planning and introduces a Selective Gaze mechanism that dynamically chooses whether to glance at global context or gaze into high-value regions.
Outcome: Experiments across six benchmarks demonstrate state-of-the-art performance.
InferPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents (2026.findings-acl)

Copied to clipboard

Challenge: Inference attacks are important for assessing model's robustness, but their implementation and parameters are challenging for non-experts.
Approach: They propose an autonomous agent capable of conducting inference attacks without human intervention.
Outcome: The proposed agent achieves a 100.0% task completion rate and near-expert attack performance with an average token cost of only 0.627 per run.
Learning Intrinsic Dimension via Information Bottleneck for Explainable Aspect-based Sentiment Analysis (2024.lrec-main)

Copied to clipboard

Challenge: Gradient-based explanation methods are increasingly used to interpret neural models in natural language processing (NLP) however, in the context of Aspect-based Sentiment Analysis, only specific dimensions are pertinent.
Approach: They propose a Gradient-based explanation framework that leverages an information bottleneck to refine word embeddings into a concise intrinsic dimension, maintaining essential features and omitting unrelated information.
Outcome: The proposed framework improves both the models’ performance and explanations’ clarity by identifying sentiment-aware features.
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing RL methods rely on unstructured self-sampling to fit scalar rewards, resulting in inefficient rollouts.
Approach: They propose a structured template-guided RL framework that augments policy optimization with explicit template guidance.
Outcome: Experiments show that TemplateRL outperforms GRPO and GRPI by 99% on AIME and 41% on AMC with superior stability on weak models and remarkable cross-domain generalization.
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for activation sparsification do not capture the relationship between activation and model performance.
Approach: They propose a general activation sparsification approach using channel-wise thresholding and selective sparsifying to capture the relationship between activation and model performance.
Outcome: The proposed approach reduces the number of activated neurons during inference by 1.27x over eight downstream tasks while activating fewer parameters than existing methods.
ImF: Embedding an Implicit Fingerprint in Your Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Training and serving large language models (LLMs) is resource-intensive, making reliable intellectual property protection and black-box ownership verification increasingly important.
Approach: They propose a method to inject a small set of secret query–response behaviors into model fingerprinting . they encode ownership information into a natural-looking target response and derive a semantically aligned query .
Outcome: The proposed fingerprints improve stealthiness and remain verifiable under model updates and deployment-time prompt interventions.
Smart-Searcher: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are powerful but prone to hallucinations due to static knowledge. Retrieval-augmented generation (RAG) helps by injecting external information, but current methods are costly, generalize poorly, or ignore the model’s internal knowledge.
Approach: They propose a framework to train large language models to leverage both internal and external knowledge sources.
Outcome: The proposed framework outperforms existing methods and achieves efficient retrieval-augmented reasoning.
LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for long-form generation assess real-world queries with hard-to-verify metrics or use synthetic setups that overlook real-life intricacies.
Approach: They propose a new approach that balances verifiable and real-world assessment with Target-Anchored Evaluation.
Outcome: The proposed model balances real-world and verifiable assessment with Target-Anchored Evaluation (TAE) it generates queries, textual materials, and anchors based on verifier targets within real-life scenarios .
Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation (2025.acl-industry)

Copied to clipboard

Challenge: Effective content moderation is essential for video platforms to safeguard user experience and uphold community standards.
Approach: They propose a method to transform a generative MLLM into a multimodal classifier using minimal discriminative training data.
Outcome: The proposed method improves F1 score by 66.50% over traditional classifiers while requiring only 2% of the fine-tuning data.
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models (2026.acl-long)

Copied to clipboard

Challenge: Masked diffusion language models have achieved significant progress in language modeling . however, the systematic analysis and empirical validation of their alignment on general tasks remains underexplored.
Approach: They propose a framework that analyzes the bias and variance of preference optimization loss and gradient based on Direct Preference Optimization.
Outcome: The proposed model outperforms its SFT-only predecessor on general benchmarks . it consistently outperformed other strong language models and ARMs on general tasks .
RLSeek: Evidence-Grounded Reasoning for RAG Hallucination Detection (2026.acl-long)

Copied to clipboard

Challenge: Recent work addresses this problem by training span-level hallucination detectors using reinforcement learning and chain-of-thought reasoning.
Approach: They propose a framework that explicitly enforces active evidence seeking during CoT reasoning by requiring quotation of relevant source segments at each verification step.
Outcome: The proposed framework improves hallucination span detection performance with limited reasoning overhead and improved robustness in out-of-domain settings.
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs (2026.findings-acl)

Copied to clipboard

Challenge: Existing black-box instruction backdoors can detect poisoned inputs, but fail to recover correct outputs once the backdoor is activated.
Approach: They propose a soft label mechanism and key-extraction-guided CoT-based defense against instruction backdoors in APIs (SLIP) they propose KCOT-based model to extract task-relevant keywords and phrases rather than only considering the single trigger or overall text semantics.
Outcome: The proposed model reduces the average attack success rate to 25.13% and improves clean accuracy to 87.15% and outperforms state-of-the-art black-box defenses.
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm (2025.findings-acl)

Copied to clipboard

Challenge: Existing LLMs cannot comprehend the complex data flow and computation process of the attention operator and utilize low-level primitive to exploit GPU performance.
Approach: They propose an LLM-friendly Thinking Language (LLM-TL) that can decouple the generation of high-level optimization logic and low-level implementation on GPU and enhance LLMs’ understanding of attention operator.
Outcome: The proposed method outshines existing LLMs on A100, RTX8000, and T4 GPUs, achieving a speed-up of up to 35.16.
Beyond Examples: Towards Automated Thought-level In-Context Reasoning for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: In-context learning (ICL) struggles with complex reasoning due to superficial, example-level implicit imitation.
Approach: They propose an automated method that shifts from surface-level examples to more guidance-oriented thought patterns.
Outcome: The proposed method achieves 80.6% accuracy on MATH and 62.5% on AMC, surpassing GPT-4o’s 77.2% and 57.5% accuracy.
Peering Behind the Shield: Guardrail Identification in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Identifying guardrails in conversational AI agents is critical for identifying malicious content . identifying guardrail components in black-box AI agents poses security challenges .
Approach: They propose a method that leverages guard-specific adversarial prompts to detect guardrails in black-box AI agents.
Outcome: The proposed method achieves perfect classification accuracy in multiple scenarios.
Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive prowess in solving a wide range of tasks with world knowledge, but it remains unclear how well they perceive their factual knowledge boundaries.
Approach: They propose to use a retrieval augmentation approach to enhance LLMs' awareness of factual knowledge boundaries to analyze factual and factual information in open-domain question answering (QA)
Outcome: The proposed method improves LLMs’ QA and judgemental capabilities by integrating supporting documents with the questions.
RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking (2021.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that passage retrieval and passage reranking are important for achieving mutual improvement.
Approach: They propose a unified listwise training approach for passage retrieval and passage reranking that incorporates a retrieval procedure and a hybrid data augmentation strategy.
Outcome: The proposed approach improves on both MSMARCO and Natural Questions datasets.
The Agent’s First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios (2026.findings-acl)

Copied to clipboard

Challenge: Existing research mainly focuses on performance upper bounds in static environments, overlooking stochastic real-world deployment.
Approach: They propose a dynamic evaluation environment that simulates a "trainee" agent continuously exploring a novel setting.
Outcome: The proposed model evaluates agents in a dynamic evaluation environment that simulates a "trainee" agent continuously exploring a novel setting.
TEMP: Taxonomy Expansion with Dynamic Margin Loss through Taxonomy-Paths (2021.emnlp-main)

Copied to clipboard

Challenge: Existing taxonomies are unable to maintain coverage due to the rising of new concepts . TEMP uses pre-trained contextual encoders to predict the position of new ideas .
Approach: They propose a self-supervised taxonomy expansion method that ranks taxonomies by ranking them . they use pre-trained contextual encoders to train the model with dynamic margin loss .
Outcome: The proposed method outperforms state-of-the-art taxonomy expansion methods by 14.3% and 15.8% on public benchmarks.
MAPO: Boosting Large Language Model Performance with Model-Adaptive Prompt Optimization (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing research emphasizes the importance of adapting prompts to specific tasks, rather than specific LLMs.
Approach: They propose a model-adaptive prompt optimizer method that optimizes original prompts for each LLM in downstream tasks.
Outcome: The proposed method can optimize prompts for an LLM in downstream tasks.
Kill two birds with one stone: generalized and robust AI-generated text detection via dynamic perturbations (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods focus on model generalization or focus on robustness.
Approach: They propose a model-based AIGT detection method that can be generalized and robust under two adversarial attacks.
Outcome: The proposed method outperforms state-of-the-art methods for generalization and robustness under two text adversarial attacks.
RAGLAB: A Modular and Research-Oriented Unified Framework for Retrieval-Augmented Generation (2024.emnlp-demo)

Copied to clipboard

Challenge: Existing research on Retrieval Augmented Generation (RAG) does not address the problem of hallucinations and real-time updating of knowledge.
Approach: They propose a modular open-source library to equip LLMs with external knowledge.
Outcome: The proposed approach reduces the need for expensive open-source tools and lacks fair comparisons between novel RAG algorithms.
Digging out Discrimination Information from Generated Samples for Robust Visual Question Answering (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to solve this problem rely on additional annotations and generate negative samples .
Approach: They propose a method to Dig out Discrimination information from Generated samples to address these limitations.
Outcome: The proposed method improves on the visual question-answering datasets.
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to align English LLMs with human preferences rely on expensive human annotations or advanced multilingual preference alignment models.
Approach: They propose a method that captures learned preferences from English models by implicit rewards . they annotate preference relations in cross-lingual instruction-following pairs using English .
Outcome: The proposed approach captures learned preferences from well-aligned English models by implicit rewards and transfers them to other languages through iterative training.
Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Event extraction is a task in natural language processing that involves identifying and extracting event information from unstructured text.
Approach: They propose a paradigm that combines schema paraphrasing with schema retrieval-augmented generation.
Outcome: The proposed paradigm retrieves paraphrased schemas and accurately generates targeted structures.
Neuro-Symbolic Query Compiler (2025.findings-acl)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) systems are limited in their ability to process information in open-source environments.
Approach: They propose a neuro-symbolic framework inspired by linguistic grammar rules and compiler design to formalize complex queries using a minimal yet sufficient Backus-Naur Form grammar.
Outcome: The proposed framework is based on a backus-naur form grammar and compiler design that maintains completeness while minimizing redundancy.
A Thorough Examination on Zero-shot Dense Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in dense retrieval (DR) models have been shown to be not as competitive as traditional sparse retrieval models in a zero-shot retrieval setting.
Approach: They propose to examine the zero-shot capability of DR models by analyzing key factors related to source training set and potential bias from target dataset.
Outcome: The proposed model is not as competitive as sparse retrieval models in a zero-shot retrieval setting.
ATLAS: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to optimize large language models with external tools are limited.
Approach: They propose a dual-path framework for dynamic tool usage in cross-domain complex reasoning . they exploit empirical priors for domain-specific alignment and RL-based multi-step routing .
Outcome: The proposed framework outperforms closed-source models and existing methods on in-distribution and out-of-distortion tasks.
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Backdoor-based LLM fingerprinting is a promising solution for intellectual property protection . however, the vulnerability of existing LLMs for the ensemble scenario is unexplored .
Approach: They propose two new fingerprinting attack methods to assess the robustness of LLM fingerprinting by token filter attack and sentence verification attack.
Outcome: The proposed methods inhibit the fingerprint response while maintaining ensemble performance.
GROVE: A Retrieval-augmented Complex Story Generation Framework with A Forest of Evidence (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for generating stories with complex plots rely on detailed prompts, which inadvertently limit the creative potential of the generated stories.
Approach: They propose a retrieval-auGmented stoRy generation framework with a fOrest of eVidEnce to enhance stories’ complexity.
Outcome: The proposed framework enables generating more diverse plotlines from human-written stories.
BASES: Large-scale Web Search User Simulation with Large Language Model based Agents (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing research on web search rely on real-user experiments, which can be costly to scale up.
Approach: They propose a user simulation framework with LLM-based agents that can generate unique user profiles at scale.
Outcome: The proposed framework can generate unique user profiles at scale, leading to diverse search behaviors.
Implicit Discourse Relation Identification for Open-domain Dialogues (P19-1)

Copied to clipboard

Challenge: Discourse relation identification is a challenging problem in open-domain dialogue systems . previous work relies on formal text but this data is not suitable for informal dialogue .
Approach: They propose a method to automatically extract the implicit discourse relation argument pairs from dialogic turns and a pipeline to identify them.
Outcome: The proposed pipeline extracts argument pairs from dialogic turns and improves it by performing feature ablation and incorporating dialogue features.
Quantification of Large Language Model Distillation (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have revealed the robustness degra-dation caused by data distillation.
Approach: They propose a framework to evaluate and quantify model distillation . they aim to identify identity cognition contradictions and analyse multi-granularity response similarities across models to measure the extent of homogenization.
Outcome: The proposed framework addresses two key aspects: (1) Identifying identity cognition contradictions to assess discrepancies in how models perceive and represent identity-related information; (2) Analyzing multi-granularity response similarities across models to measure the extent of homogenization.
Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis (2026.acl-long)

Copied to clipboard

Challenge: Existing mitigation strategies rely on global gradient geometry to resolve alignment conflicts . however, they overlook Modular Heterogeneity within Transformers, resulting in suboptimal trade-offs . Conflict-Aware Sparse Tuning (CAST) combines head-level diagnosis with sparse fine-tuning .
Approach: They propose a framework that integrates head-level diagnosis with sparse fine-tuning to address this limitation.
Outcome: The proposed framework integrates head-level diagnosis with sparse fine-tuning to reduce alignment conflicts in LLMs.
Modelling Variability in Human Annotator Simulation (2024.findings-acl)

Copied to clipboard

Challenge: Human annotator simulation (HAS) is a cost-effective alternative to human evaluation tasks.
Approach: They propose a framework to model human annotation variability via meta-learning . conditional softmax flow model leverages diverse human annotations via meta learning . results demonstrate that method can predict aggregated behaviours of human annotators .
Outcome: The proposed method achieves state-of-the-art performance on two real-world human evaluation tasks: emotion recognition and toxic speech detection.
SSP: Self-Supervised Post-training for Conversational Search (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to search for information are limited by data scarcity and lack contextual semantic understanding.
Approach: They propose a new post-training paradigm with three self-supervised tasks to initialize the conversational search model to enhance dialogue structure and contextual semantic understanding.
Outcome: The proposed method can initialize the conversational search model to enhance dialogue structure and contextual semantic understanding.
Every Document Owns Its Structure: Inductive Text Classification via Graph Neural Networks (2020.acl-main)

Copied to clipboard

Challenge: Existing graph-based methods for text classification cannot capture contextual word relationships within each document nor can they produce inductive learning of new words.
Approach: They propose to use Graph Neural Networks to learn the local word representations and then aggregate the word nodes as the document embeddings.
Outcome: The proposed method outperforms state-of-the-art methods on four benchmark datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations