Papers by Hao Wang

386 papers
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have remarkable capabilities but are vulnerable to adversarial “jailbreak” attacks designed to bypass safety guardrails.
Approach: They propose to empower a large language model to be its own red teamer . safety self-play allows the model to act as both the Attacker and Defender .
Outcome: The proposed approach outperforms baselines trained on static adversarial datasets and establishes a new benchmark for proactive safety alignment.
AdaMix: Adaptive Mixing for Short and Long Reasoning Adapters (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for large reasoning models have improved efficiency but still face limitations such as conflicting objectives and limited adaptability.
Approach: They propose an adaptive reasoning framework that applies a uniform, computation-intensive deep reasoning strategy to all problems.
Outcome: The proposed framework reduces the average response length of DeepSeek-R1-Distill-Qwen-7B by 54.9% while improving accuracy by up to 4.8% on five mathematical datasets.
ERNIE-Doc: A Retrospective Long-Document Modeling Transformer (2021.acl-long)

Copied to clipboard

Challenge: Existing models for document-level language pretraining are not suitable for long documents due to their quadratically increasing memory and time consumption.
Approach: They propose a document-level language pretraining model based on Recurrence Transformers.
Outcome: The proposed model outperforms existing models on language understanding tasks.
Adaptive Spatial and Temporal Redundancy Optimization for Efficient Reasoning in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing research to improve CoT efficiency falls into three categories, each with distinct limitations.
Approach: They propose a training-free framework that addresses both dimensions of CoT reasoning by applying a progressive precision reduction strategy coupled with an entropy-based confidence mechanism for adaptive termination.
Outcome: Empirical results show that the proposed framework achieves 11.3 efficiency gain without compromising accuracy.
AIDA-SEAT: Towards Reliable AI Doctor Assistant via State-Evaluation-Action Tree Enhanced LLMs in Online Hospital (2026.acl-industry)

Copied to clipboard

Challenge: Existing systems rely on large language models or retrieval-augmented generation (RAG) but these methods lack the explicit logical pathways essential for multi-step reasoning.
Approach: They propose an AIDA-SEAT framework to provide reliable clinical decision-making support by transforming and modifying medical documents and doctors' state-evaluation-action trees.
Outcome: The proposed framework achieves 1.01% higher than current state-of-the-art (SOTA) baselines across five departments, including common RAG-based methods.
Knowledge Graph Entity Typing with Curriculum Contrastive Learning (2025.coling-main)

Copied to clipboard

Challenge: Existing knowledge graphs suffer from incomplete type annotations because they are manually constructed by domain experts.
Approach: They propose a CCLET model using the Curriculum Contrastive Learning strategy for KGET to fuse the entity related semantic and the structural information of the Knowledge Graph (KG) they define the difficulty of the course by controlling the level of added noise and aim to accurately learn with curriculum contrastive learning strategy from easy to difficult.
Outcome: The proposed model outperforms state-of-the-art models and is highly accurate across multiple learning environments.
Safety Sidecar: Reflection-Driven Runtime Control for Safer Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety controls fail to provide runtime intervention or cross-architecture portability for autonomous LLM agents.
Approach: They propose a model-agnostic, plug-and-play module to provide arbitrary agent safety control and auditability.
Outcome: The proposed module improves the secure-solution rate by 2.9–11.2 percentage points . it adds only 3.2s to end-to-end latency and a negligible average cost of 5.37 10-4 per scenario .
SMARTAVE: Structured Multimodal Transformer for Product Attribute Value Extraction (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for product attribute value extraction are noisy and incomplete with missing values for most retailers.
Approach: They propose a Structure Mltimodal trAnsformeR for producT Attribute Value Extraction which jointly encodes the structured product information from multiple modalities.
Outcome: The proposed method outperforms state-of-the-art methods on two multimodal product datasets.
Disentangled Multi-span Evolutionary Network against Temporal Knowledge Graph Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for temporal knowledge Graphs neglect internal structural interactions between subgraphs and ignore potential smooth features that do not lead to semantic changes.
Approach: They propose to use a disentangled multi-span evolutionary network to capture local neighbor features while perceiving historical neighbor semantic information.
Outcome: Extensive experiments show that the proposed model outperforms the state-of-the-art in TKG reasoning by 22.7%.
TC–RAG: Turing–Complete RAG’s Case study on Medical LLM Systems (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to RAG neglect system state variables, resulting in poor performance and erroneous knowledge accumulation.
Approach: They propose a framework that incorporates a Turing Complete System to manage state variables and manage retrieval halting.
Outcome: The proposed framework improves on seven real-world healthcare datasets and shows that it is more accurate than existing methods.
Entity-Aware Dependency-Based Deep Graph Attention Network for Comparative Preference Classification (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to comparative preference classification do not learn entity-aware representations well or use sequential modeling approaches that do not generalize well.
Approach: They propose a deep-level deep-graph attention network that leverages word embeddings and syntactic information to solve a comparative preference classification problem.
Outcome: The proposed model achieves state-of-the-art performance in comparative preference classification.
EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot (2024.acl-demos)

Copied to clipboard

Challenge: EmpathyEar is an open-source, avatar-based multimodal empathetic chatbot . currently, ERG systems rely on text, sound, and vision .
Approach: They propose an open-source, avatar-based multimodal empathetic chatbot to fill the gap in traditional text-only ERG systems.
Outcome: The proposed system enables users to generate emotional responses to user queries . it can also generate avatars with talking faces and synchronized speeches .
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter (2023.emnlp-main)

Copied to clipboard

Challenge: Language Models (LMs) have demonstrated impressive molecule understanding ability on 1D text-related tasks, but lack 2D graph perception, a critical ability of human professionals in comprehending molecules’ topological structures.
Approach: They propose to combine a cross-modal projector and a uni-modal adapter to enable an LM to understand both text- and graph-based molecular contents via a Q-Former.
Outcome: The proposed model outperforms the baselines on tasks such as molecule captioning, IUPAC name prediction, and molecule-text retrieval.
Pre-training Multi-task Contrastive Learning Models for Scientific Literature Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models (LMs) have shown effectiveness in literature understanding tasks, especially when tuned via contrastive learning.
Approach: They propose a multi-task contrastive learning framework that enables common knowledge sharing across different scientific literature understanding tasks while preventing task-specific skills from interfering with each other.
Outcome: The proposed framework outperforms state-of-the-art pre-trained language models on a comprehensive dataset.
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to detect large language models (LLMs) generated for plagiarism use paraphrases to rewrite them to evade detection.
Approach: They propose a training-free method that effectively fools text detectors using off-the-shelf LLMs by rewriting them to evade detection.
Outcome: The proposed method deceives text detectors using off-the-shelf LLMs by rewriting them to produce human-like sentences that are less discernible by detectors.
Enhancing Self-Attention with Knowledge-Assisted Attention Maps (2022.naacl-main)

Copied to clipboard

Challenge: Existing works of knowledge infusion depend on multi-task learning frameworks, which are inefficient and require large-scale retraining when new knowledge is considered.
Approach: They propose a method which integrates knowledge-generated attention maps into the self-attention mechanism and integrates it into the model.
Outcome: The proposed model outperforms existing methods on academic datasets and industry-scale ad relevance applications.
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models (2026.acl-long)

Copied to clipboard

Challenge: closed-ended question-based benchmarks struggle with saturation as newer models emerge . crowd-sourced leaderboards rely on costly and slow human judges .
Approach: They propose a framework that leverages collective intelligence from all large language models to evaluate each other.
Outcome: a new framework enables a democratic, pairwise evaluation of all large language models . it achieves 97% correlation with human judgements, while significantly reducing the cost.
CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods focus on detecting LLM’s confidence via statistical uncertainty.
Approach: They propose to use a representation perspective to solve adaptive RAG by enabling dynamic retrieval during generation and enabling retrieval only when the query exceeds LLM's internal knowledge.
Outcome: The proposed framework is superior to existing adaptive RAG methods on a diverse set of tasks.
On the Sentence Embeddings from Pre-trained Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Pre-trained contextual representations like BERT have been widely used for NLP tasks.
Approach: They propose to transform anisotropic sentence embedding distribution to smooth and isotropic Gaussian distribution by normalizing flows that are learned with an unsupervised objective.
Outcome: The proposed method achieves significant performance gains over state-of-the-art embeddings on a variety of semantic textual similarity tasks.
DynamicFocalPO: Adaptive Focusing Strategy for Preference Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Recent preference optimization algorithms such as Direct Preference Optimization (DPO) have become prevalent for aligning large language models with human preferences.
Approach: They propose a preference optimization algorithm that introduces a modulating factor that down-weighs misranked preference pairs and employs focusing strategy that adapts over the course of training.
Outcome: Experiments show that DynamicFocalPO surpasses both DPO and FocalPO on benchmarks including Alpaca Eval 2.0 and Arena-Hard using Mistral-Base-7B and Llama-3-Instruct-8B.
Argumentation Mining on Essays at Multi Scales (2020.coling-main)

Copied to clipboard

Challenge: Argumentation mining on essays is a new task in natural language processing.
Approach: They propose a multi-scale argumentation mining model which aims to identify the types and locations of argumentation components from essay text.
Outcome: The proposed model outperforms existing models on mining all types of argumentation components on the Persuasive Essay dataset.
An Effective Span-based Multimodal Named Entity Recognition with Consistent Cross-Modal Alignment (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to name entity recognition rely on word-based sequence labeling and align image and text at inconsistent semantic levels.
Approach: They propose a span-based method which achieves a more consistent multimodal alignment from the perspectives of information-theoretic and cross-modal interaction.
Outcome: Experiments on two datasets show that SMNER outperforms the state-of-the-art methods.
The Retrieval Bottleneck: Scaling Laws for Reinforcement Learning in RAG (2026.acl-long)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) has become the dominant paradigm for building knowledge-intensive language systems.
Approach: They propose a sigmoidal scaling law that shows that retrieval quality determines the asymptotic performance ceiling.
Outcome: The proposed model achieves strong performance on knowledge-intensive benchmarks while retaining the predictable scaling long available for pre-training but previously absent in RAG-RL.
Enhancing Extractive Question Answering in Multiparty Dialogues with Logical Inference Memory Network (2025.coling-main)

Copied to clipboard

Challenge: Existing models for multiparty dialogue question answering (QA) do not consider logical inference relations in multiparty dialogs, leading to suboptimal performance.
Approach: They propose a memory network with logical inference for extractive QA in multiparty dialogues.
Outcome: The proposed model achieves state-of-the-art on Molweni and FriendsQA benchmarks.
SKEP: Sentiment Knowledge Enhanced Pre-training for Sentiment Analysis (2020.acl-main)

Copied to clipboard

Challenge: sentiment knowledge is ignored in sentiment analysis, despite its use in pretraining.
Approach: They propose to use sentiment knowledge to learn a unified sentiment representation for multiple sentiment analysis tasks.
Outcome: The proposed method outperforms strong pre-training baseline on three kinds of sentiment tasks.
OmniEvent: A Comprehensive, Fair, and Easy-to-Use Toolkit for Event Understanding (2023.emnlp-demo)

Copied to clipboard

Challenge: Event understanding is fundamental for humans to understand the world.
Approach: They propose an event understanding toolkit called OmniEvent that is comprehensive and fair . it supports mainstream modeling paradigms and the processing of 15 widely-used datasets .
Outcome: The toolkit supports mainstream modeling paradigms and the processing of 15 widely-used English and Chinese datasets.
Towards Better Modeling Hierarchical Structure for Self-Attention with Ordered Neurons (D19-1)

Copied to clipboard

Challenge: Recent studies have shown that a hybrid of self-attention networks (SANs) and recurrent neural networks (RNNs) outperforms both individual architectures, while not much is known about why the hybrid models work.
Approach: They propose to use an advanced variant of self-attention networks (SANs) to enhance the strength of hybrid models by introducing a syntax-oriented inductive bias to perform tree-like composition.
Outcome: The proposed model outperforms both individual models and a standard hybrid model on a machine translation task.
QueryForm: A Simple Zero-shot Form Entity Query Framework (2023.findings-acl)

Copied to clipboard

Challenge: Form-like document understanding is a key yet under-investigated problem . endlessly training specialized models on new document types is not scalable in many practical scenarios.
Approach: They propose to use large-scale query-entity pairs generated from form-like webpages to pre-train QueryForm.
Outcome: The proposed framework sets state-of-the-art average F1 score on XFUND and Payment benchmarks.
FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning (2026.acl-long)

Copied to clipboard

Challenge: Continual learning (CL) for large language models (LLMs) aims to enable sequential knowledge acquisition without catastrophic forgetting.
Approach: They propose a framework that aligns replay schedules with a model-centric notion of time.
Outcome: Experiments on three benchmarks show that FOREVER consistently mitigates catastrophic forgetting.
CODE-MVP: Learning to Represent Source Code from Multiple Views with Contrastive Pre-Training (2022.findings-naacl)

Copied to clipboard

Challenge: Recent studies have focused on code representation learning, which aims to represent the semantics of source code into distributed vectors.
Approach: They propose to integrate different views with the natural-language description of source code into a unified framework with Multi-View contrastive Pre-training.
Outcome: The proposed model outperforms state-of-the-art models on three downstream tasks over five datasets.
From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent research in mechanistic interpretability has revealed that Large Language models contain disentangled, human-understandable components.
Approach: They propose a framework that first identifies causal task features through frequency recall and interventional filtering, then selects “Feature-Resonant Data” that maximally activates task features for fine-tuning.
Outcome: The proposed framework outperforms existing models on mathematical reasoning, summarization, and translation tasks while using only 50% of the data.
Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative Likelihood (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for detecting modelgenerated texts from human texts are limited by the fact that absolute likelihood values of texts are bound to certain linguistic and cognitive constraints.
Approach: They propose to use relative likelihood values instead of absolute ones to extract useful features from the spectrum-view of likelihood for the human-model text detection task.
Outcome: The proposed method can reveal subtle differences between human and model languages, which find theoretical roots in psycholinguistics studies.
Enhancing Retrieval-Augmented Generation via Evidence Tree Search (2025.acl-long)

Copied to clipboard

Challenge: Evidence retrieval is used to enhance Large Language Models (LLMs) but in real-world applications, it often returns lengthy documents with redundant or irrelevant content, confusing downstream readers.
Approach: They propose a framework that reformulates evidence retrieval as a dynamic tree expansion process.
Outcome: The proposed framework outperforms existing methods on five datasets.
PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text Modeling (2021.acl-long)

Copied to clipboard

Challenge: PhotoChat contains 12k dialogues, each of which is paired with a user photo that is shared during the conversation.
Approach: They propose to use PhotoChat to facilitate research on image-text modeling by combining a photo-sharing intent prediction task and a picture retrieval task to retrieve the most relevant photo according to the dialogue context.
Outcome: The proposed tasks achieve 10.4% recall@1 and 58.1% F1 scores, indicating that the proposed dataset presents interesting yet challenging real-world problems.
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories (2024.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of Large Language Models (LLMs).
Approach: They propose a repository-level benchmark named DevEval to evaluate LLMs' coding abilities in real-world code repositories.
Outcome: The proposed benchmarks show that the LLMs perform better in real-world code repositories than existing benchmarks.
SafetyQuizzer: Timely and Dynamic Evaluation on the Safety of LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been used to evaluate the safety of their users . however, evaluation questions in current benchmarks are too straightforward and difficult to update with practical relevance due to their lack of correlation with real-world events.
Approach: They propose a question-generation framework to evaluate the safety of LLMs in the Chinese context.
Outcome: The proposed framework reduces decline rate while maintaining similar attack success rate.
LoopCoder: Scaling Code Intelligence via Looped Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models have mastered syntax-level code generation, but complex algorithmic reasoning remains a challenge.
Approach: They propose a recurrent inductive bias that aligns with the recursive nature of programming logic.
Outcome: The proposed model achieves comparable performance to standard dense models with more parameters.
LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Currently, large language models (LLMs) train on short text segments due to the computational overhead quadratic in the input lengths of their Transformer architectures.
Approach: They propose a method that allows LLMs pre-trained with 2K or 4K-long segments to generalize to up to 200M length inputs while retaining perplexity.
Outcome: The proposed method achieves 2.7 decoding speed up and 7.5 memory saving over the original model.
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on text-to-image (T2I) models focus on text alignment, image quality, and object composition capabilities.
Approach: They propose a T2I-FactualBench benchmark to evaluate the factuality of knowledge-intensive concept generation.
Outcome: The proposed framework evaluates the factuality of knowledge-intensive concept generation tasks.
COM-MRC: A COntext-Masked Machine Reading Comprehension Framework for Aspect Sentiment Triplet Extraction (2022.emnlp-main)

Copied to clipboard

Challenge: Aspect Sentiment Triplet Extraction (ASTE) aims to extract sentiment triplets from sentences, but when faced with multiple aspect terms, the MRC-based methods could fail due to the interference from other aspect terms.
Approach: They propose a COntext-Masked MRC framework for Aspect Sentiment Triplet Extraction (ASTE) which aims to extract sentiment triplets from sentences .
Outcome: The proposed framework outperforms state-of-the-art methods on benchmark datasets and shows that it can extract sentiment triplets from multiple aspect terms.
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms.
Approach: They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks.
Outcome: The proposed models perform well on mainstream benchmarks and are compared with other models.
Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Despite advances in training Large Language Models, they remain vulnerable to jailbreak, an adversarial attack method.
Approach: They propose an adversarial jailbreak algorithm that exploits the gradient information of the suffix tokens to accelerate the optimization process.
Outcome: The proposed model achieves 1.5x speedup while maintaining high attack success rates.
TableDreamer: Progressive and Weakness-guided Data Synthesis from Scratch for Table Instruction Tuning (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for table instruction tuning are limited due to limited data diversity and lack of data quality.
Approach: They propose a weakness-guided data synthesis framework for table instruction tuning that explores the vast input space of table understanding tasks and then iterates through the input space.
Outcome: The proposed framework boosts the average accuracy of Llama3.1-8B-instruct by 11.62% with 27K GPT-4o synthetic data and outperforms state-of-the-art data synthesis baselines which use more training data.
Knowledge-augmented Financial Market Analysis and Report Generation (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing methods to generate financial market analysis text require extensive financial knowledge and skill of financial analysts.
Approach: They propose a task to generate financial market analysis reports using financial market data using a financial knowledge graph.
Outcome: The proposed framework outperforms large-scale language models and retrieval-augmented baselines in the financial market analysis generation task.
KIA: Knowledge-Guided Implicit Vision-Language Alignment for Chest X-Ray Report Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing reports on medical images and reports lack fine-grained cross-modal interaction, leading to insufficient understanding of detailed information.
Approach: They propose a framework for establishing cross-modal semantic alignment in radiology report pairs using knowledge-guided implicit vision-language alignment.
Outcome: KIA improves understanding of medical images and reports by incorporating medical knowledge to enhance pathological observation and anatomical landm.
MAVEN-ARG: Completing the Puzzle of All-in-One Event Understanding Dataset with Event Argument Annotation (2024.acl-long)

Copied to clipboard

Challenge: Existing datasets for event understanding have limited coverage due to complexity of tasks.
Approach: They propose a dataset that augments MAVEN datasets with event argument annotations . they propose 98,591 events and 290,613 arguments obtained with laborious human annotation .
Outcome: The proposed dataset is the first all-in-one dataset supporting event detection, event argument extraction, and event relation extraction.
DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection (2026.acl-long)

Copied to clipboard

Challenge: Existing detectors are limited in their ability to detect large language models generated content in multilingual environments.
Approach: They propose a multilingual benchmark to evaluate advanced detectors across 8 dimensions to better align with real-world applications.
Outcome: The proposed benchmark encompasses 8 languages commonly used in commercial contexts and collects human-written texts from 6 domains highly susceptible to LLM misuse.
USSA: A Unified Table Filling Scheme for Structured Sentiment Analysis (2023.acl-long)

Copied to clipboard

Challenge: Structured Sentiment Analysis (SSA) is a problem of bi-lexical dependency parsing . previous studies have cast it as a bottleneck because of overlap and discontinuity issues .
Approach: They propose a bi-lexical dependency parsing graph and a table-filling scheme that addresses overlap and discontinuity issues.
Outcome: The proposed framework outperforms state-of-the-art methods on benchmark datasets.
SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are shifting the focus from single verifiable tasks toward complex, open-ended real-world scenarios.
Approach: They propose a framework that automatically adjusts reward weights and data importance to synchronize learning intent with data utility for optimal performance.
Outcome: The proposed framework improves model capabilities across all domains and scales.
A Self-supervised Joint Training Framework for Document Reranking (2022.findings-naacl)

Copied to clipboard

Challenge: Pretrained language models have been successfully applied to a wide range of tasks . however, the pretraining tasks were based on the context of documents .
Approach: They propose a self-supervised joint training framework with a method called Masked Query Prediction to establish semantic relations between given queries and positive documents.
Outcome: The proposed framework outperforms existing models on document reranking tasks without further pre-training . it uses a self-supervised method to establish semantic relations between given queries and positive documents.
COPEN: Probing Conceptual Knowledge in Pre-trained Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge probing studies focus on evaluating factual knowledge of pre-trained language models (PLMs) but ignore conceptual knowledge.
Approach: They evaluate conceptual knowledge of pre-trained language models by annotating 24k data instances covering 393 concepts.
Outcome: The proposed tasks evaluate pre-trained language models' conceptual knowledge of entities, learn conceptual properties, and conceptualize entities in contexts.
A Benchmark Suite of Japanese Natural Questions (2024.starsem-1)

Copied to clipboard

Challenge: Existing studies to solve QA tasks in an integrated manner are not available in other languages because of the lack of QA datasets.
Approach: They build a Japanese version of Natural Questions using natural questions from query logs of a search engine and crowdsource it using crowdsourcing.
Outcome: The proposed datasets are based on natural questions from Japanese search engines and crowdsourced.
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Multimodal Large Language Models have shown significant promise in various applications, but a comprehensive evaluation of their long-context capabilities remains underexplored.
Approach: They propose a benchmark to assess the long-context capabilities of multimodal large language models.
Outcome: The proposed benchmark compared MLLMs with API-based and open-source models in a long-context scenario.
ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, research on complex chart understanding tasks is limited . a pipeline for visual reasoning datasets addresses these limitations .
Approach: They propose a code-driven pipeline for generating visual reasoning datasets . pipeline integrates retrieval-augmented generation to retrieve professional chart templates .
Outcome: The proposed pipeline enhances chart diversity and data quality through model-based evaluation.
From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Structured pruning is a practical approach to deploying large language models (LLMs) but it fails to capitalize on modest task-specific calibration signals, causing limited downstream gains.
Approach: They propose a method that removes attention heads and MLP channels using loss-based important scores . they use perplexity for language modeling and a margin-based objective for decision-style tasks .
Outcome: The proposed method lowers perplexity and improves accuracy at higher sparsity . it also stabilizes accuracy and mitigates perxity collapse without fine-tuning .
Reasoning with Language Model is Planning with World Model (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown remarkable reasoning capabilities, particularly with Chain-of-Thought-style prompts.
Approach: They propose a framework that repurposes the LLM as both a world model and a reasoning agent and incorporates a principled planning algorithm (based on Monte Carlo Tree Search)
Outcome: The proposed framework repurposes the LLM as both a world model and a reasoning agent and incorporates a principled planning algorithm (based on Monte Carlo Tree Search) it achieves optimum balance between exploration and exploitation, while achieving high-reward reasoning paths efficiently.
TRACE: Traversal Retrieval-Augmented Chain of Evidence for Document Understanding (2026.acl-long)

Copied to clipboard

Challenge: Long-context Document Visual Question Answering (DocVQA) methods struggle with visual semantics or handling finite context windows.
Approach: They propose a new approach to longcontext document visual question answering that transforms retrieval into adaptive evidence chain construction using a Bi-Layered Graph.
Outcome: The proposed approach achieves an average accuracy improvement of 14.07% on M5BookVQA and exhibits robust generalization with a 13.38% gain across four established benchmarks.
Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for hyperbole and metaphor detection focus on superficial text features, ignoring the associations of hyperbola and metaphor . Existing frameworks focus on identifying superficial text, focusing on superficial features .
Approach: They propose an emotion-guided hyperbole and metaphor detection framework based on bidirectional dynamic interaction.
Outcome: The proposed framework outperforms baseline methods on four datasets.
ReCreate: Reasoning and Creating Domain Agents Driven by Experience (2026.acl-long)

Copied to clipboard

Challenge: Large Language Model (LLM) agents are reshaping the industrial landscape, but tasks differ widely, making them labor-intensive to build.
Approach: They propose an experience-driven framework for the automatic creation of domain agents . they leverage agent interaction histories to provide rich concrete signals on success or failure .
Outcome: The proposed framework outperforms human-designed agents and existing methods in experiments across diverse domains.
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) show promising results in language generation but often “hallucinate”, making their outputs less reliable.
Approach: They propose to shift attention to more relevant components at token- and sentence-levels for better UQ.
Outcome: The proposed approach improves the performance of a range of popular “off-the-shelf” LLMs with model sizes extending up to 33B parameters.
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities across a wide range of domains, but their generalpurpose pre-training objectives often leave them illsuited for specialized applications such as healthcare.
Approach: They propose a perplexity-aware data scaling law that establishes a predictive relationship between the perplexities of domain-specific data and the test loss.
Outcome: Experiments on medical and general-domain benchmarks show that the proposed scaling law consistently identifies near-optimal training subsets with significantly reduced data consumption.
Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding (2026.findings-acl)

Copied to clipboard

Challenge: Existing Large Language Models (LLMs) mainly address isolated tasks such as emotion analysis or stance detection.
Approach: They propose a large-scale model that combines large-level annotations with hyperbolic space to model human cognitive states.
Outcome: The proposed model outperforms baseline models on cognitive dimensions on single dimension tasks while retaining strong hierarchical structure.
An Evaluation Resource for Grounding Translation Errors (2025.findings-emnlp)

Copied to clipboard

Challenge: Current fine-grained error analyses do not ground the errors to the reasons why the annotated text spans are erroneous.
Approach: They use a bi-directional grounding scheme to ground erroneous text in two directions . if the error spans of both directions are consistent, the explanation is valid .
Outcome: The proposed grounding process improves translation error detection significantly.
SOTOPIA-π: Interactive Learning of Socially Intelligent Language Agents (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on building language agents have not addressed this social learning gap.
Approach: They propose an interactive learning method that improves the social intelligence of language agents by using behavior cloning and self-reinforcement based training on filtered social interaction data.
Outcome: The proposed method allows a 7B LLM to reach the social goal completion ability of an expert model (GPT-4-based agent) without the loss of more generic abilities, such as the ability to answer knowledge-based questions.
The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding (2020.acl-demos)

Copied to clipboard

Challenge: MT-DNN is an open-source natural language understanding toolkit . it allows researchers and developers to train customized deep learning models .
Approach: They present MT-DNN, an open-source natural language understanding toolkit . it is designed to facilitate rapid customization for a broad spectrum of NLU tasks . MT supports multi-task knowledge distillation, which can substantially compress a deep neural model without significant performance drop.
Outcome: The proposed model can significantly compress a large model without significant performance drop.
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing tools for detecting safety issues in LLMs are expensive and inefficient.
Approach: They propose an LLM-based safety detector which annotates the safety of queries and provides explanations for its decisions.
Outcome: The proposed detector outperforms baselines on four sets of query-response pairs and is effective as a safety evaluator for advanced LLMs.
Towards Rationality in Language and Multimodal Agents: A Survey (2025.naacl-long)

Copied to clipboard

Challenge: despite advances in language and multimodal agents, large language models lack rationality . despite their progress, large-scale models lack real-world grounding and feedback mechanisms .
Approach: They propose to build more rational language and multimodal agents . they also examine what criteria define rationality in intelligent systems .
Outcome: This paper assesses the state-of-the-art in language and multimodal agents . it also outlines open challenges and future research directions .
Analyzing Chain-of-thought Prompting in Black-Box Large Language Models via Estimated V-information (2024.lrec-main)

Copied to clipboard

Challenge: Chain-of-Thought (CoT) prompting and large language models (LLMs) have shown great potential in improving performance on challenging reasoning tasks.
Approach: They propose a new metric which extends the concept of pointwise V-information to black-box models and quantifies label-relevant new information introduced by CoT prompting.
Outcome: The proposed metric extends the concept of pointwise V-information to black-box models, quantifying label-relevant new information introduced by CoT prompting beyond pre-existing label information.
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing data synthesis methods suffer from limited diversity and lack precise control over problem difficulty, making them insufficient for efficient training paradigms such as curriculum learning.
Approach: They propose a data synthesis paradigm that generates high-quality, difficulty-controllable mathematical reasoning problems through hybrid and decomposed strategies.
Outcome: The proposed paradigm outperforms existing methods and improves mathematical reasoning abilities.
ProtT3: Protein-to-Text Generation for Text-based Protein Understanding (2024.acl-long)

Copied to clipboard

Challenge: Language Models excel in understanding textual descriptions of proteins, but struggle to process texts.
Approach: They propose a framework for Protein-to-Text Generation for Text-based Protein Understanding that integrates a PLM as its protein understanding module.
Outcome: The proposed framework surpasses existing baselines and is highly efficient in protein-to-text generation.
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook (2026.acl-long)

Copied to clipboard

Challenge: Recent sparsity-aware binarization approaches can achieve sub-1-bit compression, but they face performance degradation, mask-management overhead, and limited hardware compatibility.
Approach: They propose a binary quantization framework that leverages binary pattern clustering and weight transformation to overcome performance degradation and mask-management overhead.
Outcome: The proposed framework achieves state-of-the-art compression (1.11–0.7 bits) it maintains high performance with only a 3.1% accuracy drop in zero-shot benchmarks while delivering a 1.6 speedup over FP16.
Modeling Recurrence for Transformer (N19-1)

Copied to clipboard

Challenge: Existing studies show that the lack of recurrence modeling hinders the development of a translation model.
Approach: They propose to model recurrence for Transformer with an additional recurrent encoder.
Outcome: The proposed model outperforms the deep model on EnglishGerman and ChineseEnglish translation tasks.
R2F: A General Retrieval, Reading and Fusion Framework for Document-level Natural Language Inference (2022.emnlp-main)

Copied to clipboard

Challenge: Document-level natural language inference (DOCNLI) is a new task in natural language processing.
Approach: They propose a document-level natural language inference framework that fuses sentence-level tasks into a set of sentence-based tasks.
Outcome: The proposed framework improves interpretability and performance with evidence.
Dual Graph Convolutional Networks for Aspect-based Sentiment Analysis (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to model relationships between aspects and opinion words are inefficient due to informal expressions and complexity of online reviews.
Approach: They propose a dual graph convolutional networks model that considers complementarity of syntax structures and semantic correlations simultaneously.
Outcome: The proposed model outperforms state-of-the-art methods on three public datasets and validates it.
Train Once, and Decode As You Like (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to machine translation support autoregressive, semi-autoregressive and refinement-based non-auto-regressives.
Approach: They propose a unified approach for supporting different generation manners of machine translation including autoregressive, semi-autoregressive and refinement-based non-auto-regressives.
Outcome: The proposed approach achieves better or competitive translation performance compared with strong baseline models in all the settings.
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM𝛥 Integration into Upcycled MoE (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are expensive and require extensive Continued Pre-Training and data-intensive alignment to expand.
Approach: They propose a method which upcycles a dense model into a Mixture-of-Experts architecture, allocating different experts to different languages.
Outcome: Experiments show that the proposed model upcycles a dense model into a Mixture-of-Experts(MoE) architecture, allocating different experts to different languages.
Knowledge Distillation based Contextual Relevance Matching for E-commerce Product Search (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing approaches to e-commerce relevance matching ignore bipartite graphs in logs . experimental results show that proposed method improves human relevance judgment .
Approach: They propose an efficient knowledge distillation framework for e-commerce relevance matching to exploit the advantages of Transformer-style and classical relevance matching models.
Outcome: The proposed method significantly improves human relevance judgment on large-scale real-world data.
Multi-Task Learning with Shared Encoder for Non-Autoregressive Machine Translation (2021.naacl-main)

Copied to clipboard

Challenge: Existing non-autoregressive machine translation models have shown significant inference speedup but suffer from inferior translation accuracy.
Approach: They propose to use AT as an auxiliary task to transfer AT knowledge to NAT models by knowledge distillation.
Outcome: The proposed method achieves significant improvements over baseline non-Autoregressive machine translation models on WMT14 En-De and WMT16 En-Ro datasets.
DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling (2025.emnlp-main)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) systems have emerged as a pivotal methodology for enhancing Large Language Models (LLMs).
Approach: They propose a framework that decouples planning and search processes using dual value models, enabling independent optimization of plan reasoning and search grounding.
Outcome: The proposed framework decouples planning and search processes using dual value models, enabling independent optimization of plan reasoning and search grounding.
Xiaomingbot: A Multilingual Robot News Reporter (2020.acl-demos)

Copied to clipboard

Challenge: Xiaomingbot is a multilingual and multimodal software robot with four capabilities: news generation, news translation, news reading and avatar animation.
Approach: They propose to build a multilingual and multimodal software robot with four inte- gal capabilities: news generation, news translation, news reading and avatar animation.
Outcome: The proposed system generates Chinese news, then reads it in multiple languages and generates an animated avatar reading it.
UniRE: A Unified Label Space for Entity Relation Extraction (2021.acl-long)

Copied to clipboard

Challenge: Existing joint entity relation extraction models setup two separate label spaces for the two sub-tasks .
Approach: They propose to eliminate the different treatment on the two sub-tasks’ label spaces by applying a unified classifier to predict each cell’s label.
Outcome: The proposed model achieves competitive accuracy with the best extractor and is faster.
Blockwise Self-Attention for Long Document Understanding (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in pre-training and fine-tuning methods have drastically reshaped the landscape of natural language processing research.
Approach: They propose a lightweight BERT model that introduces sparse block structures into the attention matrix to reduce memory consumption and training/inference time.
Outcome: The proposed model uses 18.7-36.1% less memory and 12.0-25.1% more time to learn compared to an advanced BERT-based model, RoBERTa.
Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures (2026.acl-long)

Copied to clipboard

Challenge: Recent research has shown that reinforcement learning can elicit intriguing emergent reasoning behaviors.
Approach: They propose a comprehensive survey of the mechanistic understanding of large reasoning models . they organize findings into three core dimensions: 1) training dynamics, 2) reasoning mechanisms, and 3) unintended behaviors.
Outcome: This paper synthesizes the mechanistic understanding of large reasoning models into three dimensions . authors outline a roadmap for future studies including improved interpretability and methodologies .
Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation (2025.acl-short)

Copied to clipboard

Challenge: Existing methods for detecting hallucination in long-form tasks focus on limited domains or rely heavily on external fact-checking tools, which may not always be available.
Approach: They propose a new paradigm that augments fine-tuning with an auxiliary task for the model to jointly learn with the main task of hallucination detection.
Outcome: The proposed method outperforms existing methods for detecting hallucination in open-domain long-form generation and is more accurate than random guessing.
WildReward: Learning Reward Models from In-the-Wild Human Interactions (2026.acl-long)

Copied to clipboard

Challenge: Prior work focused on collecting preference pairs, requiring substantial annotation efforts.
Approach: They propose a pipeline to extract reliable human feedback from in-the-wild interactions . they propose to use WildChat as an interaction source to train the model .
Outcome: The proposed model achieves comparable or even superior performance compared to conventional models with improved calibration and cross-sample consistency.
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Text-Centric Visual Question Answering (TEC-VQA) is a text-centric visual task understanding tool.
Approach: They introduce a benchmark that features human expert annotations across 9 languages . they prioritize the text in question-answer pairs while disregarding visual text in images .
Outcome: The proposed benchmarks prioritize the text in question-answer pairs while disregarding visual text in images.
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue (2025.naacl-long)

Copied to clipboard

Challenge: Existing dialogue systems focus on brief single-session interactions, neglecting real-world needs for long-term companionship and personalized interactions.
Approach: They propose a model-agnostic framework for long-term dialogue agents . they use event summary and persona management to enable reasoning .
Outcome: The proposed framework incorporates three independently tunable modules dedicated to event perception, persona extraction, and response generation.
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference (2026.acl-long)

Copied to clipboard

Challenge: Speculative decoding (SD) is a powerful and efficient way to accelerate autoregressive generation.
Approach: They propose a training-free framework that recovers valid tokens discarded by standard verification . they use online correction memory and Semantic Consistency Gating to analyze rejections .
Outcome: The proposed framework outperforms existing methods and achieves peak throughput speedup of 2.33x.
FuzzAug: Data Augmentation by Coverage-guided Fuzzing for Neural Test Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Using large language models to generate meaningful tests is expensive and time-consuming .
Approach: They propose a data augmentation technique that incorporates valid testing semantics and diverse coverage-guided inputs into large language models.
Outcome: The proposed technique improves performance over the baselines by incorporating valid testing semantics and providing diverse coverage-guided inputs.
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for enhancing performance through increased use of expert knowledge often result in diminishing sparsity during expert selection.
Approach: They propose a framework that integrates the computational processes of MoE with the concept of knowledge transferring in multi-task learning.
Outcome: The proposed framework outperforms existing methods under identical conditions concerning the number of experts.
UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning (2021.acl-long)

Copied to clipboard

Challenge: Existing pre-training methods focus on single-modal tasks or multi-modal ones . large-scale pre- training has drawn much attention in both the community of Compute Vision (CV) and Natural Language Processing (NLP).
Approach: They propose a UNIfied-MOdal pre-training architecture which can adapt to both single-modal and multi-modal understanding and generation tasks.
Outcome: The proposed model can learn more generalizable representations with rich non-paired single-modal data.
VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown significant promise in embodied decision-making tasks within virtual open-world environments, but lack domain-specific knowledge.
Approach: They propose a cost-effective agent framework that integrates cross-modal domain knowledge and finetunes a dedicated object detection model for visual analysis.
Outcome: The proposed framework reduces the requirement for domain-specific training data from millions of samples to a few hundred.
Offline Reinforcement Learning for LLM Multi-step Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly applied to complex tasks requiring multi-step reasoning.
Approach: They propose an offline method for enhancing multi-step reasoning by optimizing the soft Bellman Equation by combining a policy model and a value function.
Outcome: The proposed method surpasses existing methods on multi-step reasoning benchmarks and can be extended to multi-iteration frameworks when additional resources are available.
DeepMed: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & Inference (2026.findings-acl)

Copied to clipboard

Challenge: Medical reasoning models are constrained by parametric knowledge and can induce hallucinations and spurious attributions.
Approach: They propose a model that uses a multi-hop med-search QA synthesis method to apply the DR paradigm in medical contexts.
Outcome: The proposed model outperforms larger medical reasoning models on medical benchmarks.
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack (2026.findings-acl)

Copied to clipboard

Challenge: Existing watermarking algorithms focus on defending against paraphrase and piggyback spoofing attacks, which can inject harmful content, compromise reliability, and undermine trust in attribution.
Approach: They propose an algorithm capable of defending against paraphrase and spoofing attacks.
Outcome: Experiments on large language models and language models show that DualGuard is the first watermarking algorithm capable of defending against both paraphrase and spoofing attacks.
Active Sentence Learning by Adversarial Uncertainty Sampling in Discrete Space (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing uncertainty sampling methods are time-consuming and can't be executed frequently.
Approach: They propose adversarial uncertainty sampling in discrete space to find informative unlabeled text samples for annotation using adversarials.
Outcome: The proposed approach outperforms baselines on effectiveness on five datasets.
CLEAN–EVAL: Clean Evaluation on Contaminated Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods to evaluate large language models are prone to data contamination.
Approach: They propose a method which parses contaminated data and back-translates it into a candidate set.
Outcome: The proposed method reduces data contamination and evaluates the LLMs more cleanly.
MILL: Mutual Verification with Large Language Models for Zero-Shot Query Expansion (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for query expansion lack corpus-specific knowledge and cost.
Approach: They propose a query-query-document generation method that leverages large language models for mutual verification to produce diverse sub-queries and corresponding documents.
Outcome: The proposed method is fully zero-shot and extensive experiments on three public benchmark datasets demonstrate its effectiveness over existing methods.
Logical Consistency as a Bridge: Improving LLM Hallucination Detection via Label Constraint Modeling between Responses and Self-Judgments (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for hallucination detection focus on implicit neural uncertainty or explicit symbolic reasoning, ignoring factual hallucinosities.
Approach: They propose a framework that bridges neural features and symbolic judgments for hallucination detection by leveraging a "meta-judgment" process to map symbolic labels back into the feature space.
Outcome: Extensive experiments on 4 public datasets, across 4 LLMs, against 8 baselines demonstrate the superiority of LaaB.
To Copy Rather Than Memorize: A Vertical Learning Paradigm for Knowledge Graph Completion (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for embedding knowledge graphs implicitly memorize relation rules to infer missing links, but they are difficult to memorize due to the inherent deficiencies of such implicit memorization strategy.
Approach: They propose a vertical learning paradigm that allows to explicitly copy target information from related factual triples for more accurate prediction.
Outcome: The proposed model improves generalization ability and makes distant link prediction significantly easier.
Derailer-Rerailer: Adaptive Verification for Efficient and Reliable Language Model Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing prompting methods struggle with complex tasks and reasoning stability, limiting their practical deployment.
Approach: They propose a framework that adaptively balances reasoning accuracy and computational efficiency by employing a lightweight Derailer mechanism to assess reasoning stability and selectively triggers an advanced Rerailer verification process only when necessary.
Outcome: The proposed framework achieves significant accuracy improvements (8-11%) while maintaining 2-3 times better efficiency than existing verification methods.
Reliable Use of Lemmas via Eligibility Reasoning and Section-Aware Reinforcement Learning (2026.acl-short)

Copied to clipboard

Challenge: Recent large language models (LLMs) perform strongly on mathematical benchmarks but often import conclusions without validating assumptions.
Approach: They propose a model that encodes a lemma specification and trains with reinforcement learning and section-aware loss masking to assign penalty to the section responsible for errors.
Outcome: The proposed model performs well on benchmarks but often misapplyes lemmas . the model is able to encode the specification and train with reinforcement learning .
Reasoning-Enhanced Domain-Adaptive Pretraining of Multimodal Large Language Models for Short Video Content Governance (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing approaches to identifying inappropriate content require extensive human-labeled data and lack cross-issue generalization.
Approach: They propose a reasoning-enhanced multimodal large language model (MLLM) pretraining paradigm for unified inappropriate content detection.
Outcome: The proposed model improves the MLLM's performance in both zero-shot and supervised fine-tuning settings and shows strong generalization capabilities to emergent, previously unseen issues.
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks (2026.acl-long)

Copied to clipboard

Challenge: Generating synthetic datasets via large language models (LLMs) has emerged as promising approach to improve LLM performance.
Approach: They propose three mitigation strategies to mitigate bias inheritance in LLMs by analyzing real and LLM-augmented data.
Outcome: The proposed methods can work differently on different tasks and biases.
Path-enhanced Pre-trained Language Model for Knowledge Graph Completion (2025.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models have achieved remarkable knowledge graph completion (KGC) success.
Approach: They propose a path-enhanced pre-trained language model-based knowledge graph completion method which uses multi-view generation to infer missing facts in triple-level and path-level simultaneously.
Outcome: The proposed method significantly improves the performance of the knowledge graph completion task.
M2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation (2026.acl-long)

Copied to clipboard

Challenge: prevailing methods for machine translation are often hindered by misleading reward signals.
Approach: They propose a framework that aligns large language models to human preferences . they propose 'M2PO' to correct the bias towards partial errors .
Outcome: The proposed framework outperforms open-source models and achieves parity with proprietary models.
UKP-SQuARE v2: Explainability and Adversarial Attacks for Trustworthy QA (2022.aacl-demo)

Copied to clipboard

Challenge: Question Answering (QA) systems rely on deep neural networks, which are difficult to interpret by humans.
Approach: They propose an interpretable model that provides an explanation infrastructure for comparing models based on saliency maps and graph-based explanations.
Outcome: The proposed methods can be used to compare models based on saliency maps and graph-based explanations.
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks on longcontext large language models fail to reflect their deep understanding capabilities across diverse tasks.
Approach: They propose a benchmark to assess the ability of long-context large language models to handle long-text problems.
Outcome: The proposed model achieves 50.1% accuracy when directly answering the questions . human experts achieve only 53.7% accuracy under a 15-minute time constraint .
On Safety Risks in Experience-Driven Self-Evolving Agents (2026.findings-acl)

Copied to clipboard

Challenge: Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduces underexplored safety risks.
Approach: They investigate how experience accumulation and utilization in self-evolving agents affect safety performance across web-based and embodied environments.
Outcome: The findings expose inherent limitations of current self-evolving agents and call for more principled strategies to ensure safe and reliable adaptation.
STACL: Simultaneous Translation with Implicit Anticipation and Controllable Latency using Prefix-to-Prefix Framework (P19-1)

Copied to clipboard

Challenge: Simultaneous translation is notoriously dif- ficult due to word-order differences.
Approach: They propose a prefix-to-prefix framework that implicitly learns to anticipate in a single translation model.
Outcome: The proposed framework achieves low latency and reasonable qual- ity on 4 directions.
Which Side Are You On? A Multi-task Dataset for End-to-End Argument Summarisation and Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have made it difficult to build an automated debate system that helps people to synthesise persuasive arguments.
Approach: They propose to use an argument mining dataset to capture the end-to-end process of preparing an argumentative essay for a debate.
Outcome: The proposed dataset shows that it performs better on individual tasks than on human-centred evaluations.
Constructing Highly Inductive Contexts for Dialogue Safety through Controllable Reverse Generation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to detect toxic generation of pretrained language models rely on templates, data extraction, crowdsourcing workers or automatic generation.
Approach: They propose a method to construct adversarial contexts conditioned on a given response . they augment existing dataset BAD+ and construct a new dataset B AD+ .
Outcome: The proposed method can detect toxic or biased content in large pretrained language models.
MTG: A Benchmark Suite for Multilingual Text Generation (2022.findings-naacl)

Copied to clipboard

Challenge: Using MTG, we train and evaluate multilingual text generation models using human-annotated data.
Approach: They propose a multilingual multiway text generation dataset with 400k human-annotated data that includes four generation tasks across five languages.
Outcome: The proposed dataset includes four generation tasks across five languages (English, German, French, Spanish and Chinese) it provides comprehensive evaluations with diverse generation scenarios.
DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to multi-hop question answering struggle to identify and organize dynamic knowledge . et al., 2023; Liu e.t. al. 2023) suggest a dual-process framework for multi-step reasoning .
Approach: They propose a synergistic dual-process framework that integrates reasoning and retrieval.
Outcome: The proposed framework improves answer accuracy and coherence even in smaller-scale models.
Contrastive Aligned Joint Learning for Multilingual Summarization (2021.findings-acl)

Copied to clipboard

Challenge: Existing summarization systems for multilingual text summarizing are limited due to the lack of large-scale data in multiple languages.
Approach: They propose a multilingual summarization system that can understand documents in multiple languages and generate summaries in the corresponding language.
Outcome: The proposed model improves over monolingual models in all languages and transferable to other languages.
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving (2025.findings-emnlp)

Copied to clipboard

Challenge: Vision-Language Models struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning.
Approach: AgentThink integrates Chain-of-Thought reasoning with dynamic, agent-style tool invocation for autonomous driving tasks.
Outcome: Experiments on the DriveLMM-o1 benchmark show AgentThink significantly boosts overall reasoning scores by 53.91% and enhances answer accuracy by 33.54% .
A Framework of Knowledge Graph-Enhanced Large Language Model Based on Question Decomposition and Atomic Retrieval (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to enhance LLMs with knowledge graphs have limited results . knowledge graph question answering (KGQA) provides interpretable reasoning for large language models .
Approach: They propose a framework for KG-enhanced LLM based on question decomposition and atomic retrieval . they propose question decomposing tree as framework for LLM reasoning .
Outcome: The proposed framework outperforms existing reasoning-based baselines on KGQA datasets.
When Allies Turn Foes: Exploring Group Characteristics of LLM-Based Multi-Agent Collaborative Systems Under Adversarial Attacks (2025.findings-emnlp)

Copied to clipboard

Challenge: a new study examines the group characteristics of adversarial agents in multi-agent collaborative systems . collaborative agents are tasked with generating counterfactual answers to a given collaborative problem .
Approach: They evaluate collaborative systems under adversarial attacks and propose methods to mitigate them . they also introduce a new metric to quantify the robustness of collaborative systems against such attacks .
Outcome: The proposed method has been proven effective against adversarial attacks.
DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to integrate lexical knowledge into deep learning models are limited by large-scale dynamic lexicons.
Approach: They propose a plug-in lexicon incorporation approach for BERT based sequence labeling tasks . they adopt word-agnostic tag embeddings to avoid re-training the representation .
Outcome: The proposed framework achieves new SOTA even with large scale lexicons, the authors show . they adopt word-agnostic tag embeddings to avoid re-training the representation .
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to optimize RAG generators fail to align with RAG requirements thoroughly.
Approach: They propose a method for optimizing the RAG generator from multiple preference perspectives to align with RAG requirements comprehensively.
Outcome: The proposed method improves the performance of RAG generators by incorporating retrieved documents into the prompt.
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models (2025.acl-long)

Copied to clipboard

Challenge: Existing methods and benchmarks for information retrieval are inadequately representing the diversity of code in various domains and tasks.
Approach: They propose a benchmark specifically designed to assess code retrieval capabilities.
Outcome: The proposed benchmark aims to invigorate research in the code retrieval domain . it shares the same data schema as other popular benchmarks like MTEB and BEIR .
Leibniz: Theory-of-Mind Driven Neuro-Symbolic Logical Reasoning via Multi-Agent Collaboration (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for logical reasoning with large language models suffer from insufficient rule semantic grounding and weak rule application mechanisms.
Approach: They propose a theory-of-mind driven neuro-symbolic reasoning framework that integrates natural language and symbolic representations throughout the reasoning process.
Outcome: The proposed model surpasses state-of-the-art models in reasoning accuracy and flexibility.
Modeling Content Importance for Summarization with Pre-trained Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies on content importance do not consider semantics and context when evaluating importance.
Approach: They apply information theory to pre-trained language models to define the concept of importance from the perspective of information amount.
Outcome: Experiments on CNN/Daily Mail and New York Times show that the proposed model can model the importance of content better than previous methods based on F1 and ROUGE scores.
Friendly Topic Assistant for Transformer Based Abstractive Summarization (2020.emnlp-main)

Copied to clipboard

Challenge: Abstractive document summarization is a comprehensive task in natural language processing.
Approach: They propose a topic assistant that rearranges and learns document semantics . they propose TA that is compatible with Transformer-based models and user-friendly .
Outcome: The proposed model is compatible with Transformer-based models and user-friendly.
TP-RAG: Benchmarking Retrieval-Augmented Large Language Model Agents for Spatiotemporal-Aware Travel Planning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models (LLMs) focus on basic plan validity, but neglect critical aspects such as route efficiency, POI appeal, and real-time adaptability.
Approach: They propose a benchmark for retrieval-augmented, spatiotemporal-aware travel planning that integrates retrieved trajectories with LLMs’ intrinsic reasoning.
Outcome: The proposed framework improves spatial efficiency and POI rationality while challenging universality and robustness due to conflicting references and noisy data.
PSSAT: A Perturbed Semantic Structure Awareness Transferring Method for Perturbation-Robust Slot Filling (2022.coling-1)

Copied to clipboard

Challenge: Existing slot filling models memorize inherent patterns of entities and contexts from training data.
Approach: They propose a perturbed semantic structure awareness transferring method for slot filling models . they use two MLM-based training strategies to learn contextual semantic structure and word distribution .
Outcome: The proposed method outperforms existing methods and gains strong generalization while preventing model from memorizing inherent patterns of entities and contexts.
Separation and Fusion: A Novel Multiple Token Linking Model for Event Argument Extraction (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for event argument extraction (EAE) lack cross-event information and require longer role sequences . et al. (2017): outperforms state-of-the-art methods for EE.
Approach: They propose a separation-and-fusion paradigm to separate the acquisition of cross-event information and fuse it into the argument extraction of a target event.
Outcome: The proposed model outperforms the state-of-the-art models on four widely used datasets.
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks (2025.naacl-long)

Copied to clipboard

Challenge: Large visionlanguage models (LVLMs) are a powerful visual-language reasoning tool.
Approach: They propose to integrate attention analysis with LLaVA-CAM to determine interactions between visual representations.
Outcome: The proposed approach can be used to determine interactions between visual representations.
Loki: An Open-Source Tool for Fact Verification (2025.coling-demos)

Copied to clipboard

Challenge: Loki is an open-source fact-checking tool designed to address the growing problem of misinformation.
Approach: They propose a tool that breaks down the fact-checking task into five steps . they propose LOKI, which offers a semiautomated, human-in-the-loop approach .
Outcome: a new open-source tool is designed to address the growing problem of misinformation . the tool breaks down the fact-checking task into five steps to assist human judgment .
GAPO: Robust Advantage Estimation for Real-World Code LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement learning (RL) is widely used for post-training large language models (LLMs) in code editing, but in real-world code editing scenarios, reward distributions are often skewed with unpredictable noise, leading to distorted advantage computation and increased rollout outliers.
Approach: They propose a group-relative method that finds an interval with the highest SNR and uses the median of that interval as an adaptive Q to replace the group mean in advantage calculation.
Outcome: The proposed method improves on nine instruction-tuned LLMs while remaining plug-and-play and efficient.
Clip-Tuning: Towards Derivative-free Prompt Learning with a Mixture of Rewards (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing work does not take full advantage of over-parameterized characteristics of large pre-trained language models.
Approach: They propose a method that uses frozen "thinned" networks to obtain a mixture of rewards and advance the derivative-free prompt learning.
Outcome: The proposed method outperforms previous gradient-free prompt learning methods and achieves parity with gradient-based counterparts on seven language understanding benchmarks under few-shot settings.
AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models have been remarkable . users face a choice between using cloud-based LLMs for generation quality or local-based ones for lower computational cost .
Approach: They propose a new LLM utilization paradigm that facilitates collaborative operation . they evaluate AdaSwitch across 7 benchmarks and compare it to other LLMs .
Outcome: The proposed model improves performance of local and cloud agents across 7 benchmarks . it achieves competitive results compared to the cloud agent while utilizing less computational overhead.
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Recent advances establish "SFT-then-RL" as the defacto paradigm for enhancing large reasoning mod- els on automatically verifiable tasks.
Approach: They propose an entropy-preserving SFT method to enhance exploration capabilities through intrinsic curiosity.
Outcome: The proposed method outperforms the vanilla method on reasoning tasks by 2.5 points . it also outperformed the vanilla SFT by 2.9 points on out-of-distribution tasks .
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to training LLMs at ultra-low precisions suffer from convergence instability and substantial training costs.
Approach: They propose a progressive QAT framework with outlier channel splitting to address these issues . they use nested structure of integer quantization grids to enable a "train once, deploy any precision" paradigm .
Outcome: The proposed framework outperforms baselines on both Llama2/3 and W2A16, with an 11 speedup over BF16.
RealMedDial: A Real Telemedical Dialogue Dataset Collected from Online Chinese Short-Video Clips (2022.coling-1)

Copied to clipboard

Challenge: Existing medical dialogue systems are limited by the lack of corpora and data from real scenarios.
Approach: They construct a Chinese medical dialogue dataset based on real medical consultations.
Outcome: The proposed dataset is applicable to a wide range of NLP tasks with respect to medical dialogue.
A Bounding Box is Worth One Token - Interleaving Layout and Text in a Large Language Model for Document Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for integrating spatial layouts with text have limitations . existing methods produce overly long text sequences or lack autoregressive traits of LLMs .
Approach: They introduce Interleaving Layout and Text in a Large Language Model (LayTextLLM) they use OCR-derived text and spatial layouts to integrate with LLMs for document understanding .
Outcome: The proposed model shows an increase in performance in KIE and VQA tasks.
Fine-Grained Data Ordering Improves Fine-Tuning for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Prior work focused on data preprocessing, focusing on filtering and cleaning data . a study aimed to improve fine-grained scheduling of data order in epochs .
Approach: They propose a fine-grained scheduling method of data order in epochs to fill this gap . they define data difficulty based on relevance between data and model .
Outcome: The proposed method improves on pre-training and small-scale fine-tuning experiments 2.4% over baselines.
LatticeGen: Hiding Generated Text in a Lattice for Privacy-Aware Large Language Model Generation on Cloud (2024.findings-naacl)

Copied to clipboard

Challenge: Currently, the server controls the generated text, but users can't keep it private . prompted generation is a common interaction paradigm for large language models on cloud .
Approach: They propose a protocol where the server handles most of the computation while the client controls the sampling operation.
Outcome: The proposed protocol protects both prompt and generation under strong attacks.
Vision-Enhanced Semantic Entity Recognition in Document Images via Visually-Asymmetric Consistency Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models train a visual encoder with weak cross-modal supervision signals, resulting in a limited capacity to capture non-textual features and suboptimal performance.
Approach: They propose a Visually-Asymmetric coNsistenCy Learning approach that enhances the model’s ability to capture fine-grained visual and layout features through the incorporation of color priors.
Outcome: The proposed approach outperforms the strong LayoutLM series baseline on benchmark datasets and provides insights for optimizing model performance.
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for document understanding in the wild are based on scanned or digital documents . however, these benchmarks fail to capture the challenges posed by documents in the real world .
Approach: They propose a new benchmark that incorporates a diverse set of manually captured document images reflecting real-world conditions.
Outcome: The proposed model is based on a set of manually captured document images reflecting real-world conditions and is compared with digital or scanned documents.
Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data.
Approach: They investigate the existence of code-switching in the pre-training corpus and categorize it into four types within two quadrants.
Outcome: The proposed approach improves performance across benchmarks and representation space.
MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction (2023.acl-long)

Copied to clipboard

Challenge: Natural language video localization (NLVL) aims to localize a temporal moment from an untrimmed video that semantically corresponds to a given text query.
Approach: They propose a proposal-based solution that generates proposals and selects the best matching proposal.
Outcome: The proposed solution is faster than existing approaches on three public datasets.
Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing Diffusion Language Models rely on hard binary masking and discrete token assignments, which hinder the revision of early decisions.
Approach: They propose a diffusion-based language modeling approach that replaces hard binary masks with evolving soft token distributions.
Outcome: The proposed approach outperforms existing DLMs on multiple benchmarks.
FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing non-factuality detection methods require response generation, which incurs significant computational overhead.
Approach: They propose a lightweight model called Factuality Lens which effectively probes hidden representations of fact-seeking questions for the NFP task.
Outcome: The proposed model is able to probe hidden representations of fact-seeking questions and reduce development costs.
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches focus on syntactic correctness through synthetic micro-benchmarks or subjective human ratings, despite semantic fidelity and usability.
Approach: They propose a framework that enables effective evaluation of decompilers in reverse engineering workflows . they compare six industrial-strength decompils and six recent LLM-powered approaches .
Outcome: The proposed framework outperforms commercial tools in code understandability despite lower functionality correctness . it shows that it can transform human-centric reverse engineering workflows .
FewRel: A Large-Scale Supervised Few-Shot Relation Classification Dataset with State-of-the-Art Evaluation (D18-1)

Copied to clipboard

Challenge: Empirical results show that even the most competitive few-shot learning models struggle on this task, especially as compared with humans.
Approach: They propose a Few-Shot Relation Classification Dataset consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers.
Outcome: The proposed methods perform well on the most competitive few-shot learning models, especially as compared with humans.
KeFVP: Knowledge-enhanced Financial Volatility Prediction (2023.findings-emnlp)

Copied to clipboard

Challenge: Current studies ignore the role of financial metrics knowledge in earnings calls and little consideration is given to integrating text and price information.
Approach: They propose to integrate financial metrics knowledge into text comprehension by knowledge-enhanced adaptive pre-training and effectively incorporating text and price information by introducing a conditional time series prediction module.
Outcome: The proposed method outperforms state-of-the-art methods on three real-world datasets and is effective and reliable.
Reinforcement Learning for Edit-Based Non-Autoregressive Neural Machine Translation (2024.naacl-srw)

Copied to clipboard

Challenge: Non-autoregressive (NAR) language models have a performance gap due to the large decoding space and difficulty in capturing dependency between target words accurately.
Approach: They propose to use reinforcement learning to enhance the performance of edit-based NAR models by using stepwise reward maximization and episodic reward maximisation.
Outcome: The proposed model outperforms autoregressive models in the evaluation of an edit-based model.
SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing solutions for supervised fine-tuning often lead to catastrophic forgetting, where models lose their previously acquired knowledge and general capabilities.
Approach: They propose a self-distribution alignment method that aligns input sequence logits to preserve the model’s semantic distribution, thereby mitigating catastrophic forgetting and improving downstream performance.
Outcome: The proposed method achieves a superior balance between downstream learning and general capability retention.
On the Emotion Understanding of Synthesized Speech (2026.acl-long)

Copied to clipboard

Challenge: Existing models for emotion understanding do not capture fundamental features of synthesized speech.
Approach: They evaluate emotion recognition models on synthesized speech using SER models and generative models.
Outcome: The proposed model can't generalize to synthesized speech because of speech token prediction . generative models tend to infer emotion from textual semantics while ignoring paralinguistic cues.
Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework (2025.acl-long)

Copied to clipboard

Challenge: Existing systems fail to fully leverage the structure of logical tasks throughout the reasoning process, causing bottlenecks in efficiency and efficacy.
Approach: They propose a logic-complete reasoning framework, Aristotle, which integrates symbolic expressions and logical rules into the entire reasoning process.
Outcome: The proposed framework outperforms state-of-the-art reasoning frameworks in accuracy and efficiency.
From Observation to Understanding: Front-Door Adjustments with Uncertainty Calibration for Enhancing Egocentric Reasoning in LVLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods that adapt LVLMs to egocentric tasks overlook critical agent-environment interactions, limiting their ability to perform egoic reasoning.
Approach: They propose a zero-shot paradigm to enhance egocentric reasoning by simulating human causal reasoning by formalizing ego-centric reasoning using a structural causal model.
Outcome: The proposed method improves egocentric reasoning abilities on six tasks.
Contextual Rephrase Detection for Reducing Friction in Dialogue Systems (2021.emnlp-main)

Copied to clipboard

Challenge: Large-scale conversational AI based dialogue systems like Alexa, Siri, and Google Assistant, are getting more and more prevalent in real-world applications to help users across the globe.
Approach: They propose a contextual rephrase detection model ContReph to automatically identify rephrasings from multi-turn dialogues using contextual information and user-agent interaction signals.
Outcome: The proposed model outperforms the pairwise rephrase detection models by leveraging the context and user-agent interaction signals.
TabPrompt: Graph-based Pre-training and Prompting for Few-shot Table Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods of Table Understanding (TU) focus on the textual content within the tabular data, disregarding the topological information of the table.
Approach: They propose a framework that uses tabs to understand tabular data without ignoring the topological information of the table.
Outcome: The proposed framework outperforms baselines in few-shot table understanding tasks.
Can Large Language Models Effectively Support Decision-Making in Sudden Emergencies? (2026.findings-acl)

Copied to clipboard

Challenge: Existing research has focused on the earlier stages of emergency response . lack of suitable datasets for reliable and compliance-aware decision-oriented modeling and evaluation is limiting current research .
Approach: They propose a first real-world emergency decision-making dataset EDM-Bench . they propose 'rule-enhanced reasoning framework' that integrates external regulatory knowledge with constrained inference mechanisms to improve both decision safety and interpretability.
Outcome: The proposed framework improves decision safety and interpretability by integrating regulatory knowledge with constrained inference mechanisms.
ClinAlign: Scaling Healthcare Alignment from Clinician Preference (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for aligning open-ended outputs with fine-grained clinician preferences are weakly grounded in professional guidelines.
Approach: They propose a framework to align large language models' outputs with fine-grained clinician preferences . they propose 119 broadly reusable, clinically grounded principles organized by clinical dimensions .
Outcome: The proposed framework outperforms existing models on HealthBench-Hard and Deepseek-R1 and o3.
Capture Human Disagreement Distributions by Calibrated Networks for Natural Language Inference (2022.findings-acl)

Copied to clipboard

Challenge: Previously, it's common to disregard it as noise or as a sign of poor-quality data, as their annotations are heavily based on personal experience and opinions.
Approach: They propose to capture the human disagreement distribution from the perspective of model calibration.
Outcome: The proposed model can achieve competitive performance when well-calibrated, on divergence scores between predictive probability and the true human opinion distribution, and the accuracy.
BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on narrow tasks and leave a fundamental question unanswered . Existing models only focus on specific tasks, requiring rigorous reasoning and knowledge .
Approach: They propose a benchmark to connect theoretical foundations with practical business knowledge and applications.
Outcome: The benchmark systematically evaluates both open-source and commercial LLMs . it reveals how theoretical knowledge translates into practical performance in business .
Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration (2026.acl-long)

Copied to clipboard

Challenge: Non-sequential and bidirectional nature of diffusion large language models makes direct likelihood-based self-evaluation challenging.
Approach: They propose a self-evaluation confidence quantification method for diffusion large language models that quantifies confidence by computing the probability of regenerating tokens in the entire generated sequence, given the full context.
Outcome: The proposed method is correlated with semantic coherence and answer accuracy.
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a recent study shows that large language models have limited generalization in low-resource languages like Chinese.
Approach: They propose to evaluate the zero-shot generalizability of large language models to the Chinese language . they release only half of the dataset publicly, with the remainder kept private .
Outcome: The Chinese Instruction-Following Benchmark evaluates the generalizability of LLMs to the Chinese language.
Mixture of Multimodal Adapters for Sentiment Analysis (2025.naacl-long)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) have been used for text sentiment analysis but sentiment is hidden in other modalities.
Approach: They propose to fuse emotions from different data to analyze sentiments . they use compression parameter for each expert to reduce training burden .
Outcome: The proposed method achieves state-of-the-art with a tiny trainable parameter count compared to current methods . emotions hidden in body movements or vocal timbres eclipse traditional methods compared with text sentiment analysis .
Rethinking Document-level Neural Machine Translation (2022.findings-acl)

Copied to clipboard

Challenge: Neural machine translation models are weak enough for document-level translation . current models only translate sentences individually, resulting in poor document coherence .
Approach: They propose to use the original Transformer model to test document-level neural machine translation . they find that the original transformer models can achieve strong results for document translation if trained properly .
Outcome: The proposed model outperforms sentence-level models on nine datasets and two sentence- level datasets across six languages.
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following (2025.emnlp-main)

Copied to clipboard

Challenge: Best practices for RL in instruction following remain underexplored.
Approach: They propose a verification method that combines rule-based code verification with LLM-based verification from a large reasoning model.
Outcome: The proposed method achieves state-of-the-art performance among models of comparable size and generalizes well to unseen constraints.
IMCI: Integrate Multi-view Contextual Information for Fact Extraction and Verification (2022.coling-1)

Copied to clipboard

Challenge: Existing models for fact extraction and verification fail to utilize multi-view contextual information.
Approach: They propose to integrate multi-view contextual information (IMCI) for fact extraction and verification by combining contextual information with inter-document context.
Outcome: The proposed framework achieves state-of-the-art performance on the open-domain Wikipedia task with a winning FEVER score of 73.96% and label accuracy of 77.25% on the online blind test set.
Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries (2025.acl-long)

Copied to clipboard

Challenge: Existing adversarial text attacks rely on abundant access to shared internal features and numerous queries, limited to a single task type.
Approach: They propose a black-box attack that exploits the transferability of adversarial texts . they use a deep-level substitute model trained in a plug-and-play manner for text classification .
Outcome: The proposed attack can target multiple tasks with minimal perturbations . it can target commercial APIs, large language models, and image-generation models .
LLM-Based Agent Society Investigation: Collaboration and Confrontation in Avalon Gameplay (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies on LLM agents' social behaviors are lacking . previous studies focused on positive social behaviors, leaving research on negative social behaviors relatively scarce.
Approach: They propose a framework that features a multi-agent system facilitating efficient communication and interaction with LLM agents.
Outcome: The proposed framework is based on Avalon and evaluates on game success and analyzes agents’ social behaviors.
Sensitivity-LoRA : Low-Load Sensitivity-Based Fine-Tuning for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) is a promising approach to adapting LLMs to specialized tasks . existing rank allocation techniques remain computationally inefficient and unstable .
Approach: They propose a low-rank adapted model that approximates model weight updates using low-ranked decomposition.
Outcome: The proposed method is limited by its uniform rank allocation to each incremental matrix . it leverages the second-order derivatives of the loss function to capture weight sensitivity .
SimPBL: A Multi-Agent Framework for Project-Based Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing LLMs provide partial assistance without modeling these roles, and overly comprehensive help can reduce learner autonomy.
Approach: They propose a multi-agent framework with an orchestrator agent that provides adaptive scaffolding from interaction logs and collaborator agents that support project work through boundary-aware collaboration.
Outcome: The proposed framework improves learner examination scores by 14% . it is based on a multi-agent framework with an orchestrator agent .
Mixture-of-LoRAs: An Efficient Multitask Tuning Method for Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Instruction Tuning has the potential to stimulate or enhance specific capabilities of large language models.
Approach: They propose a mixture-of-LoRAs architecture which is a parameter-efficient tuning method designed for multi-task learning with LLMs.
Outcome: The proposed method can be iteratively adapted to a new domain, enabling quick domain-specific adaptation.
Incorporating Instructional Prompts into a Unified Generative Framework for Joint Multiple Intent Detection and Slot Filling (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to multiple intent detection and slot filling focus on task-specific components to capture the relationships between intents and slots.
Approach: They propose a Unified Generative framework that captures the relationships between intents and slots in an utterance and formulates the task as a question-answering problem.
Outcome: The proposed framework surpasses baselines on full-data and multi-intent benchmarks on 5-shot and 10-shot scenarios.
ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual Corpora (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for pretraining cross-lingual models are limited in their size due to the limited amount of parallel corpora.
Approach: They propose a method that encourages the model to align multiple languages with monolingual corpora to overcome the constraint of the parallel corpus size.
Outcome: The proposed method outperforms existing cross-lingual models and delivers new state-of-the-art results in various cross-linguistic downstream tasks.
Composition-based Heterogeneous Graph Multi-channel Attention Network for Multi-aspect Multi-sentiment Classification (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for Aspect-based sentiment analysis (ABSA) focus on aspect terms with the same sentiment polarity . current methods focus on sentences with only one aspect term or multiple aspect terms .
Approach: They propose a novel method to model inter-aspect relationships and aspect-context relationships simultaneously using a heterogeneous graph.
Outcome: The proposed method can predict sentiments towards the given aspect term in a sentence . it can provide more detailed predictions compared with sentence-level sentiment analysis.
FAER: Benchmarking VLMs for Failure-Aware Embodied Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Visual-language models (VLMs) are the core component of embodied agents in perceiving the environment and making decisions.
Approach: They propose a failure-aware benchmark to evaluate the performance of visual language models (VLMs) in long-horizon tasks.
Outcome: The proposed benchmark evaluates the performance of 16 widely utilized VLMs and 4 LLMs for FAER tasks.
INarIG: Iterative Non-autoregressive Instruct Generation Model For Word-Level Auto Completion (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models for word-level autocompletion (WLAC) only use human typed sequences as prefixes in decoding module.
Approach: They propose a novel iterative nonautoregressive instruct generation model for WLAC task . it uses human typed sequences and iterating decoding with subwords to fully utilize input information.
Outcome: The proposed model is more competent in dealing with low-frequency words, and achieves state-of-the-art results on the WMT22 and benchmark datasets.
TLM: Token-Level Masking for Transformers (2023.emnlp-main)

Copied to clipboard

Challenge: Structured dropout approaches have been investigated to regularize the multi-head attention mechanism in Transformers.
Approach: They propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting by manipulating the connections between tokens in the multi-head attention via masking.
Outcome: The proposed regularization scheme outperforms attention dropout and DropHead on 18 datasets and can establish a new record on the data-to-text benchmark Rotowire (18.93 BLEU).
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: a framework that leverages the visual-language model to select key knowledge retrieved by DPR and answer questions improves performance of the baseline on the open-domain Knowledge-based VQA benchmark, OK-VQA.
Approach: They propose a framework that leverages visual-language models to retrieve related knowledge . they use dense passage retrieval to retrieve knowledge related to visual-linguistics .
Outcome: The proposed framework significantly improves the baseline on the open-domain Knowledge-based VQA benchmark, OK-VQA.
Adaptive Structure Induction for Aspect-based Sentiment Analysis with Spectral Perspective (2023.findings-emnlp)

Copied to clipboard

Challenge: incorporating structure information can enhance the performance of aspect-based sentiment analysis.
Approach: They propose to use pre-trained language models to induct latent structures from a spectrum perspective.
Outcome: The proposed model shortens Aspects-sentiment Distance and improves structure induction ability.
Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for parameter pruning fail to utilize the knowledge from pruned parameters.
Approach: They propose a method that uses manifold learning and the Information Bottleneck measure to merge similar layers to preserve model performance.
Outcome: The proposed method outperforms pruning methods on multiple datasets and LLMs with quantization and achieves substantial compression ratios.
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs (2025.findings-acl)

Copied to clipboard

Challenge: Existing tool environments face challenges in balancing stability, scale, and realism, especially for benchmarking purposes.
Approach: They propose a framework that trains specialized LLMs to accurately simulate real API responses by supervised fine-tuning and chain-of-thought reasoning.
Outcome: The proposed framework achieves superior accuracy and stability compared to state-of-the-art methods on the newly constructed MirrorAPI-Bench and its integration into StableToolBench.
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) based agent systems have made great strides in real-world applications beyond traditional NLP tasks.
Approach: They propose a new LLM-based Multi-Agent System benchmark, Collab-Overcooked, built on the popular Overcooked-AI game with more applicable and challenging tasks in interactive environments.
Outcome: The proposed benchmark provides a multi-agent framework supporting diverse tasks and objectives and encourages collaboration through natural language communication.
MdEval: Massively Multilingual Code Debugging (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks primarily focus on Python and are limited in terms of language diversity.
Approach: They propose a multilingual debugging benchmark that includes 3.9K test samples of 20 programming languages and introduces the debug instruction corpora MdEval-Instruct by injecting bugs into the correct multilingual queries and solutions.
Outcome: The proposed benchmark includes 3.9K test samples of 20 programming languages and covers the automated program repair task, bug localization task, and bug identification task.
Diffusion Glancing Transformer for Parallel Sequence-to-Sequence Learning (2024.naacl-long)

Copied to clipboard

Challenge: Experimental results show that non-autoregressive generation models are superior in generation efficiency but inferior in generation quality.
Approach: They propose a diffusion glancing transformer which employs a modality diffusion process and residual glancy sampling to improve multi-modality modeling.
Outcome: The proposed model outperforms autoregressive and non-autoregressive models on machine translation and text generation benchmarks.
HyperAdaLoRA: Accelerating LoRA Rank Allocation During Training via Hypernetworks without Sacrificing Performance (2026.findings-acl)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) assumes a uniform rank r for each incremental matrix, not accounting for the varying significance of weight matrices across modules and layers.
Approach: They propose a framework that allows for faster convergence of low-rank adaptive models . they use a hypernetwork to prune the outputs of the hypernetworks to generate parameters .
Outcome: The proposed framework accelerates convergence of AdaLoRA by leveraging a hypernetwork.
LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing (2026.findings-acl)

Copied to clipboard

Challenge: Large language model (LLM) routing assigns each query to the best suitable model from an ensemble.
Approach: They introduce a large-scale benchmark and unified framework for LLM routing . they find that many routing methods exhibit similar performance under unified evaluation .
Outcome: The proposed benchmark provides comprehensive metrics for both performance-oriented and performance-cost trade-off routing.
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling (2025.acl-long)

Copied to clipboard

Challenge: Speculative sampling is an efficient way to accelerate the auto-regressive generation process of large language models.
Approach: They propose a frequency-ranked speculative sampling framework that optimizes draft candidate selection through vocabulary space compression.
Outcome: Experiments show that FR-Spec reduces LM Head computation overhead by 75% while ensuring the equivalence of the final output distribution.
AHVE-CNER: Aligned Hanzi Visual Encoding Enhance Chinese Named Entity Recognition with Multi-Information (2025.coling-main)

Copied to clipboard

Challenge: Existing glyph-based models neglect the relationship between pictorial elements and radicals for Named Entity Recognition (NER) tasks.
Approach: They propose a model that integrates multi-source visual and phonetic information of Hanzi . they propose combining pictographic features with radicals to facilitate integration .
Outcome: The proposed model improves performance on benchmark datasets.
Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Experimental results demonstrate the efficacy of our approach in generating high-quality sentences resembling human output.
Approach: They propose a forgetting mechanism that disregards distant tokens, reducing the burden of penalty selection.
Outcome: The proposed approach generates high-quality sentences resembling human output.
Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation Dataset (2021.emnlp-main)

Copied to clipboard

Challenge: Existing evaluations on the population task are either not accurate (automatic evaluation with randomly sampled negative examples) or of small scale (human annotation).
Approach: They propose a reasoning over commonsense knowledge bases (CSKBs) that are free-text and have a human annotation set to probe commonsensical reasoning.
Outcome: The proposed model is based on a human-annotated evaluation set and is compared with existing models on the population task.
MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets only cover limited relation types at once, which prevents models from taking full advantage of relation interactions.
Approach: They construct a large-scale human-annotated ERE dataset with improved annotation schemes to address these drawbacks.
Outcome: The proposed dataset is larger than existing datasets of all the ERE tasks by at least an order of magnitude.
ALPS: Attention Localization and Pruning Strategy for Efficient Adaptation of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Prior research has focused on optimizing general-purpose large language models to downstream tasks . however, these approaches inherently introduce data dependency, which hinders generalization and reusability.
Approach: They propose an algorithm that localizes the most task-sensitive attention heads and prunes by restricting attention training updates to these heads, thereby reducing alignment costs.
Outcome: The proposed algorithm achieves 2% performance improvement over baselines on three tasks while localizing the most task-sensitive attention heads.
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters (2026.acl-long)

Copied to clipboard

Challenge: Existing research lacks systematic analysis of the applicability and methodology of cross-modal skill injection.
Approach: They investigate the applicability and methodology of cross-modal skill injection by integrating a domain-expert LLM into a VLM.
Outcome: The proposed method enables transfer of domain-specific expertise from Large Language Models (LLMs) to VLMs without incurring additional training data requirements or significant computational overhead.
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks designed to evaluate the reasoning capabilities of large models are limited in scope and lack flexibility to adapt difficulty according to evolving reasoning capacities of models.
Approach: They propose a benchmark that incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models.
Outcome: The proposed benchmark incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models.
Calibrating Inference Time Alignment with Sequence-level Risk Accumulation (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to decode large language models (LLMs) often over-reject benign information, limiting their generalizability in real-world scenarios where harmful and benign information coexist.
Approach: They propose a framework to regulate decoding alignments for Large Language Models (LLMs) they employ a reward-guided branch decoding paradigm to incorporate safety awareness during generation.
Outcome: The proposed framework achieves superior performance on four attack benchmarks and two neutral datasets.
Pre-training Multilingual Neural Machine Translation by Leveraging Alignment Information (2020.emnlp-main)

Copied to clipboard

Challenge: Existing pre-training methods are not effective for machine translation tasks.
Approach: They propose a method to pre-train a universal multilingual neural machine translation model . they use random aligned substitution technique to bring words and phrases with similar meanings closer in the representation space.
Outcome: The proposed approach improves translation quality on low, medium, rich resource languages.
R3: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Experimental results show that R3 is a superior alternative to traditional search algorithms for multistep retrosynthesis planning.
Approach: They propose a framework that reformulates multistep retrosynthetic planning as a generative reasoning task.
Outcome: The proposed framework achieves state-of-the-art Top-1 accuracy of 43.7% on retrobench . it leverages Large Language Models to reformulate multistep retrosynthesis as a generative reasoning task.
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark (2026.findings-acl)

Copied to clipboard

Challenge: Mobile GUI agents show promise in automating tasks but face significant generalization challenges in long-tail scenarios.
Approach: They propose a benchmark framework for mobile GUI agents that measures the performance of GUI agents by analyzing their performance.
Outcome: The LearnGUI benchmark outperforms existing methods in offline and online evaluations and demonstrates consistent gains across model architectures.
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing multi-objective preference alignment methods for large language models face limitations such as auxiliary reward/reference models and computational complexity.
Approach: They propose a framework that achieves dynamic balance across preference dimensions by using dimension-aware generation metrics as implicit rewards.
Outcome: Empirical results show that AMoPO outperforms state-of-the-art methods by 28.5% .
Towards Verifiable Text Generation with Evolving Memory and Self-Reflection (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) often produce factually incorrect information, also known as hallucination.
Approach: They propose a framework for verifiable text generation with evolving memory and self-reflection that incorporates long-term memory to retain documents and recent documents.
Outcome: The proposed framework outperforms baselines on five datasets across three knowledge-intensive tasks.
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems (2025.findings-acl)

Copied to clipboard

Challenge: SciVerse is a multi-modal scientific evaluation benchmark to assess large multi-models . it examines the scientific knowledge comprehension, multi-mod content interpretation and Chain-of-Thought reasoning . authors examine the scientific proficiency of LMMs in scientific domains based on their work .
Approach: They propose a multi-modal scientific evaluation benchmark to thoroughly assess Large Multi-modal Models across 5,735 test instances in five different versions.
Outcome: The proposed evaluation reveals critical limitations in LMMs' scientific proficiency and provides new insights into future developments.
Virtual Compiler Is All You Need For Assembly Code Search (2024.acl-long)

Copied to clipboard

Challenge: Using a large dataset, we find that assembly code search is a significant task for reverse engineers.
Approach: They propose to train a Large Language Model (LLM) to emulate a general compiler.
Outcome: The proposed model surpasses the baseline by 26%.
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation (2026.acl-demo)

Copied to clipboard

Challenge: Existing document translation pipelines face a tension between linguistic processing and layout preservation.
Approach: They propose a framework for layout-preserving PDF translation that decouples visual layout metadata from semantic content.
Outcome: The proposed framework improves layout fidelity, visual aesthetics, and terminology consistency over representative baselines while maintaining competitive translation precision.
Demystifying Data Organization for Enhanced LLM Training (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized various fields, yet their training efficiency is heavily reliant on effective data curation.
Approach: They propose to reuse pre-computed sample-level scores originally generated for data efficiency and introduce two new data ordering methods to improve LLM training.
Outcome: The proposed methods improve the stability and performance of LLM training.
Multi-agent Learning for Neural Machine Translation (D19-1)

Copied to clipboard

Challenge: Experimental results show that training with more than one agent improves translation quality and improves accuracy.
Approach: They propose to introduce diverse agents in an in- teractive updating process to train NMT models with an additional agent.
Outcome: The proposed approach improves on NIST Chinese-English, IWSLT 2014 German- English, WMT 2014 English-German translation tasks and shows competitive performance on all tasks.
LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval (2026.findings-acl)

Copied to clipboard

Challenge: Document editing requires full-context awareness of dependencies, but processing entire documents for each edit incurs prohibitive token costs and latency.
Approach: a framework that constructs lightweight dependency graphs captures semantic relationships and structural hierarchies across document elements is proposed for agentic document editing . a scaLing agentic agentic framework is based on a dependency graph framework that captures dependencies and refactors function dependencies.
Outcome: a new framework achieves 76 consistency versus 56 baseline while reducing token usage by 85 . the framework is based on a framework that captures semantic relationships and structural hierarchies across document elements . it can be used to improve document consistency, but it also reduces token costs and latency .
Stable Language Guidance for Vision–Language–Action Models (2026.acl-long)

Copied to clipboard

Challenge: Existing vision-Language-Action models are notoriously brittle to linguistic perturbations.
Approach: They propose a probabilistic framework that disentangles physical affordance from semantic execution.
Outcome: The proposed framework disentangles physical affordance from semantic execution.
Better Literary Translation: A Multi-Aspect Data Generation and LLM Training Approach (2026.acl-industry)

Copied to clipboard

Challenge: Literary translation requires balancing expression fluency with literary effect due to the scarcity of high-quality training data and the difficulty of capturing nuanced quality trade-offs.
Approach: They propose a multi-aspect iterative refinement framework that generates high-quality translation references and preference data through specialized LLM translators.
Outcome: The proposed models outperform the ground truth for SFT by 8.65 CEA100 points while leveraging an explicit reward model for GRPO yields an additional 1.51 point improvement.
pFedGPT: Hierarchically Optimizing LoRA Aggregation Weights for Personalized Federated GPT Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for fine-tuning Large Language Models (LLMs) struggle with data heterogeneity and adapt shared global knowledge to individual client needs.
Approach: They propose a framework that leverages Hierarchical Bayesian Optimization (HBO) for fine-grained, personalized LoRA aggregation.
Outcome: The proposed framework achieves state-of-the-art (SOTA) performance on personalized FL benchmarks while introducing only minimal (approx. 4%) additional optimization overhead.
Variational Language Concepts for Interpreting Foundation Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Foundation Language Models (FLMs) have achieved remarkable success in natural language processing.
Approach: They propose a variational Bayesian framework to provide word-level interpretations for FLMs . they propose valc to find optimal language concepts to interpret FLM predictions .
Outcome: Empirical results show that the proposed framework can provide conceptual interpretations for foundation language models.
Kanbun-LM: Reading and Translating Classical Chinese in Japanese Methods by Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Classical Chinese was introduced to Japan approximately 2,000 years ago . it was gradually adapted to a Japanese form called Kanbun-Kundoku (Kanbun) in Japanese reading and translating methods .
Approach: They construct a dataset that compares Classical Chinese and Kanbun in Japan using character reordering and machine translation tasks.
Outcome: The proposed dataset compares the current language models with human scores and compared them with human-level models.
CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing work on multi-agent collaborative tasks in Minecraft is limited due to inefficiency and limited fault tolerance.
Approach: They propose a framework that incorporates causality to manage dependencies among subtasks.
Outcome: The proposed framework achieves state-of-the-art performance in multi-agent cooperative tasks of Minecraft.
UCL-Bench: A Chinese User-Centric Legal Benchmark for Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing legal benchmarks focusing on knowledge and logic evaluate LLMs on various tasks in legal domain, but few have explored the practical application of LLM by actual users.
Approach: They propose a Chinese user-centric legal benchmark that aims to assess the practical application of LLMs by real users.
Outcome: The proposed model outperforms existing models on various tasks in legal domain but does not outperfect ChatGPT.
DUET: Joint Exploration of User–Item Profiles in Recommendation System (2026.findings-acl)

Copied to clipboard

Challenge: Existing LLMs are opaque and difficult to interpret, resulting in limited interpretability.
Approach: They propose an interaction-aware profile generator that jointly produces user and item profiles conditioned on both user history and item evidence.
Outcome: The proposed model outperforms baselines on three real-world datasets.
Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation (2026.acl-long)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) is a key parameter-efficient fine-tuning method . however, its effectiveness is hampered by semantic drift and structural incoherence .
Approach: They propose a low-rank Adaptation framework that tackles semantic drift and structural incoherence by pruning task-irrelevant directions.
Outcome: Experiments on large language models, vision models, and vision models show that the proposed framework outperforms LoRA and advanced dynamic rank allocation and sparsity-based methods.
ADELIE: Aligning Large Language Models on Information Extraction (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) struggle to follow complex instructions of IE tasks due to not being aligned with humans.
Approach: They propose an aligned large language moDEL that effectively solves various IE tasks including closed IE, open IE and on-demand IE.
Outcome: The proposed model achieves state-of-the-art (SoTA) performance among open-source models.
Communication-Efficient Desire Alignment for Proactive Embodied Human–Agent Interaction (2026.acl-long)

Copied to clipboard

Challenge: Effective real-world human–agent interactions are long-term and repeated.
Approach: They propose a simulation that uses a proxy user with value-driven preferences and natural language behavior to evaluate how agents adapt to users across interactions and satisfy their desires.
Outcome: HA-Desire, a home assistance simulation, shows that agents can adapt to user needs and provide proactive assistance within limited communication.
JanusMM: A Benchmark for Self-Deprecation Understanding in Real-World Multimodal Conversations (2026.acl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) are a common communicative strategy in human society, often using image-text interplay to express emotions and intentions.
Approach: They propose to evaluate multimodal large language models (MLLMs)' understanding of self-deprecation in real-world conversations using 2,016 bilingual memes.
Outcome: The proposed framework evaluates MLLMs' understanding of self-deprecation in real-world conversations.
Beyond the Surface: Measuring Self-Preference in LLM Judgments (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods measure self-preference bias by comparing the scores a judge model assigns to its own responses with those assigned to other models.
Approach: They propose to use gold judgments as proxies for the actual quality of responses . they propose to measure self-preference bias as the difference between the judge model's own and other models' scores .
Outcome: The proposed method can assess self-preference bias across large language models . it uses gold judgments as proxies for the ground truth scores of the judge model .
DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal large language models may deviate from this pattern due to attention drift and underutilization of visual evidence.
Approach: They propose a Dual-Indicator Guided Contrastive Alignment (DICA) that tracks visual attention and output image correlations to improve visual grounding.
Outcome: The proposed model outperforms existing approaches and significantly reduces hallucinations.
Learning with Noisy Labels for Sentence-level Sentiment Classification (D19-1)

Copied to clipboard

Challenge: Existing research on learning with noisy labels dates back to the 1980s, but it is still vibrant today.
Approach: They propose a novel DNN model called NetAb to deal with noisy labels during training and train the networks using their respective loss functions in mutual reinforcement.
Outcome: The proposed model can fit training data with noisy labels and predict clean labels.
Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching (2025.coling-main)

Copied to clipboard

Challenge: Entity matching (EM) is a critical step in entity resolution (ER).
Approach: They propose a method that incorporates record interactions from different perspectives.
Outcome: The proposed framework improves on 8 ER datasets and 10 LLMs and achieves higher efficiency and effectiveness.
Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration (2026.acl-short)

Copied to clipboard

Challenge: Existing methods for listwise reranking exhibit intrinsic position bias . existing methods are constrained by an inherent trade-off between efficiency and flexibility .
Approach: They propose a training-free framework that mechanically decouples positional bias from ranking decisions.
Outcome: a training-free framework decouples position bias from ranking decisions . evaluations show it outperforms training-based methods and outperformed expensive methods .
The Devil is in the Details: On the Pitfalls of Event Extraction Evaluation (2023.findings-acl)

Copied to clipboard

Challenge: Event extraction (EE) is a fundamental information extraction task aimed at extracting events from plain texts.
Approach: They propose to specify data preprocessing, standardize outputs, and provide pipeline evaluation results to avoid these pitfalls.
Outcome: The results show that the evaluations are reliable and lack pipeline evaluations.
From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing definitions of streaming LLMs are fragmented and lack a systematic taxonomy . large language models are pre-trained on static and full-context corpora .
Approach: They propose a systematic taxonomy of current streaming Large Language Models and propose underlying methodologies for streaming LLMs.
Outcome: The proposed model is based on data flow and dynamic interaction to clarify existing ambiguities.
Augmenting Reasoning Capabilities of LLMs with Graph Structures in Knowledge Base Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent work uses Large Language Models (LLMs) for semantic parsing to address Knowledge Base Question Answering tasks.
Approach: They propose a framework that augments reasoning capabilities of LLMs with Graph Structures in Knowledge Base Question Answering to retrieve question-related graph structures.
Outcome: The proposed framework outperforms existing methods on GrailQA and WebQSP under the few-shot setting.
Scaling Unverifiable Rewards: A Case Study on Visual Insights (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to scale complex, open-ended tasks with unverifiable rewards are not scalable to multi-stage pipelines.
Approach: They propose a process-based refinement framework that scales inference across stages of a multi-agent pipeline, instead of refining a single output over time.
Outcome: The proposed framework scales inference across stages of a multi-agent pipeline, instead of refining a single output over time as in prior work.
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on IE tasks have focused on recognizing and analyzing cross-modal information . a multimodal large language model (MLLM) is developed to analyze IE across modalities .
Approach: They propose a multimodal large language model (MLLM) capable of grounding information from all modalities.
Outcome: The proposed framework provides a framework to analyze IE tasks over various modalities and their fine-grained groundings.
GOME: Grounding-based Metaphor Binding With Conceptual Elaboration For Figurative Language Illustration (2024.emnlp-main)

Copied to clipboard

Challenge: Existing Large Language Models (LLMs) and multimodal models are unable to illustrate figurative language based on literal objects, ignoring the underlying groundings and associations across disparate metaphorical domains.
Approach: They propose a grounding-based method for metaphor illustration that integrates metaphorical knowledge into systematic instructions for existing large language models.
Outcome: The proposed method is superior to existing LLMs, diffusion models, or their direct collaboration.
Long-form Hallucination Detection with Self-elicitation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for hallucination detection tend to decompose text into isolated statements, unable to understand contextual semantics.
Approach: They propose a framework to leverage self-generated thoughts derived from prior statements as catalysts to elicit the expression of intrinsic knowledge and understand contextual semantics.
Outcome: The proposed framework enables self-elicitation to elicit expressions of knowledge and understand semantics.
CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has begun to address routing instability in VQA models by grouping similar concepts or routing based on examples.
Approach: They propose a Concept-Guided Routing framework which incorporates semantics of the answer options to guide expert selection in the training phase.
Outcome: The proposed framework delivers strong performance across multiple VQA tasks, demonstrating the effectiveness of the proposed framework.
Self-Reinforcing Controllable Synthesis of Rare Relational Data via Bayesian Calibration (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to synthesis of relational/structured tabular data lack effective feedback mechanism to optimize quality of generated data.
Approach: They propose a relational data generator with dynamic guidance framework that uses chain-of-thought steps to generate tabular data for enhancing downstream imbalanced classification performance.
Outcome: The proposed framework outperforms existing approaches in both data fidelity and downstream imbalanced classification performance on real and synthetic datasets.
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to verify agent behaviors in complex environments rely on rule-based verifiers or LLM-as-a-Judge models.
Approach: They propose a benchmark to evaluate Agent-as-a-Judge across three domains . the benchmark covers search, data systems, and graphical user interfaces - with 155 tasks and 516 trajectories .
Outcome: The proposed benchmark outperforms existing benchmarks in search, data systems, and GUI domains while revealing open challenges in agent-based verification.
ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to climate research are limited to simple Q A tasks . a lack of data and computational expertise has created bottlenecks .
Approach: They propose a general-purpose autonomous framework to perform end-to-end climate research tasks across diverse climate sub-fields.
Outcome: The proposed framework outperforms state-of-the-art benchmarks in rigorousness and practicality.
ACEBench: A Comprehensive Evaluation of LLM Tool Usage (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for evaluating LLMs’ tool usage face several limitations: limited evaluation scenarios, lacking assessments in real multi-turn dialogue contexts; narrow evaluation dimensions, with insufficient detailed assessments of how LLM use tools; and reliance on LLM or real API executions for evaluation, which introduces significant overhead.
Approach: ACEBench is a benchmark for evaluating tool usage in Large Language Models . it categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent.
Outcome: ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent.
Chain-of-Thought Reasoning in Tabular Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to extend chain-of-thought reasoning into large language models are not viable in the scenario of privatization deployment or limited resources.
Approach: They propose a framework that extends chain-of-thought reasoning into tabular language models . framework coordinates two TaLMs responsible for CoT generation and answer inference .
Outcome: The proposed framework outperforms the state-of-the-art ChatGPT on the TABMWP dataset by 9.55% (82.60%92.15% in accuracy) with less parameters (0.8B).
Towards Persona-Based Empathetic Conversational Models (2020.emnlp-main)

Copied to clipboard

Challenge: Empathetic conversational models have been shown to improve user satisfaction and task outcomes in numerous domains.
Approach: They propose a task towards persona-based empathetic conversations and propose e-learning model CoBERT that can be used to train persona on emmpathetic conversations.
Outcome: The proposed model improves empathetic responding more when trained on e-mpathetic conversations than non-empathy ones.
MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language Navigation (2025.acl-long)

Copied to clipboard

Challenge: Vision-language navigation (VLN) is a key task in Embodied AI . traditional approaches rely on historical observations as spatio-temporal contexts for decision making .
Approach: They propose a vision-language navigation model that leverages an annotation system to replace historical frames.
Outcome: The proposed model can be used as a new memory representation method in vision-language navigation . it can be applied to simulated and real-world environments, and it is validated by experiments .
Faster MoE LLM Inference for Extremely Large Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing inference optimizations for coarse-grained Mixture-of-Experts models implicitly assume a fixed activation budget, which is poorly understood.
Approach: They propose a training-free policy that adapts token-level activation using router confidence and entropy while remaining within the model’s original budget.
Outcome: The proposed skipping policy can provide substantial throughput gains, but optimal static schedules vary significantly across models and routing mechanisms.
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to training large language models suffer from unstable value estimation, whereas outcome supervision struggles with credit assignment due to sparse, trajectory-level rewards.
Approach: They propose a framework that integrates process supervision into group relative policy optimization.
Outcome: The proposed framework outperforms standard GRPO on knowledge-intensive benchmarks by 5.0% and 6.3% on Qwen3-1.7B.
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have witnessed remarkable advancements in recent years, prompting the exploration of tool learning.
Approach: They propose a virtual API server and stable evaluation system to assess the stability of large-scale real-time APIs.
Outcome: The proposed benchmarks demonstrate the stability of the proposed system and its caching system.
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents (2026.acl-long)

Copied to clipboard

Challenge: Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, but systematic evaluation of GUI–shortcut hybrid agents remains underexplored.
Approach: They propose a benchmark that evaluates GUI-shortcut hybrid agents with a specific focus on the mobile domain.
Outcome: MAS-Bench evaluates agent's ability to generate shortcuts by discovering and creating reusable, low-cost workflows.
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) can reveal toxic or offensive content inadvertently or intentionally.
Approach: They propose to control the diversity of both sides according to the number of samples for fine-tuning, which can directly reflect their impact.
Outcome: The proposed approach improves the performance of large language models after fine-tuning.
Zero-Shot Spoken Language Understanding via Large Language Models: A Preliminary Study (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown promising results in zero-shot settings, which motivates us to explore prompt-based methods.
Approach: They propose a two-stage framework which transforms the SLU task into a question-answering problem by directly prompting LLMs.
Outcome: The proposed framework can be built by directly prompting LLMs to understand user needs without training data.
AutoAlign: Get Your LLM Aligned with Minimal Annotations (2025.acl-demo)

Copied to clipboard

Challenge: Automated Alignment (ALM) is a set of algorithms designed to align Large Language Models (LLMs) with human intentions and values while minimizing manual intervention.
Approach: They propose an open-source toolkit that integrates mainstream automated algorithms through a consistent interface and an accessible workflow supporting one-click execution for prompt synthesis and automatic alignment signal construction.
Outcome: The proposed framework enables easy reproduction of existing results through extensive benchmarks and facilitates the development of novel approaches via modular components.
Explain the Synth: Interpretable Evaluation of LLM Data Synthesis (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used to generate tabular data.
Approach: They propose a framework that uses a rule-based model as a shared explanatory language to examine the explanation of real versus synthetic data.
Outcome: The proposed framework compares the explanatory structure induced by real versus synthetic data.
RASD: Retrieval-Augmented Speculative Decoding (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for generating draft tokens rely on lightweight draft models or additional model structures to generate tokens and retrieve context from databases.
Approach: They propose to use a pruning method to enhance model-based speculative decoding by combining the best-fit model with the best retrieval tree.
Outcome: The proposed method achieves state-of-the-art inference acceleration across tasks such as DocQA, Summary, Code, and In-Domain QA.
AgentReview: Exploring Peer Review Dynamics with LLM Agents (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods of peer review analysis do not address multivariate nature of the process, account for latent variables, and are constrained by privacy concerns due to the sensitive nature of data.
Approach: They propose a large language model based peer review simulation framework which effectively disentangles the impacts of multiple latent factors and addresses the privacy issue.
Outcome: The proposed framework disentangles the impacts of multiple latent factors and addresses privacy concerns.
Look Beyond Feeling: Unveiling Latent Needs from Implicit Expressions for Proactive Emotional Support (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are gaining popularity as scalable tools for mental health support . however, nearly half of individuals do not receive timely support due to limited selfawareness or reluctance to seek help.
Approach: They propose a proactive emotional support framework that leverages principles of active listening to uncover implicit user needs.
Outcome: The proposed model elicits implicit emotional needs and delivers empathetic support compared to baselines .
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark (2025.acl-long)

Copied to clipboard

Challenge: Evaluation benchmarks based on predefined domains and human-labeled data face limitations in addressing evaluation needs for emerging domains.
Approach: They propose an automated information retrieval benchmark based on predefined domains and human-labeled data . AIR-Bench is automated and Heterogeneous with three key features .
Outcome: The proposed benchmarks are based on predefined domains and human-labeled data.
Linking Adaptive Structure Induction and Neuron Filtering: A Spectral Perspective for Aspect-based Sentiment Analysis (2024.lrec-main)

Copied to clipboard

Challenge: incorporating structure information can improve the performance of aspect-based sentiment analysis.
Approach: They propose a method to conduct neuron-level manipulations on word representations in the frequency domain.
Outcome: The proposed method can achieve or come close to state-of-the-art in ABSA.
IM-TQA: A Chinese Table Question Answering Dataset with Implicit and Multi-type Table Structures (2023.acl-long)

Copied to clipboard

Challenge: Existing benchmarks only evaluate model performance on tables with explicit table structures, which means headers are explicitly annotated and treated as model input during inference.
Approach: They propose a new Table Question Answering (TQA) dataset with implicit and multi-type table structures that requires the model to understand tables without directly available header annotations.
Outcome: The proposed framework outperforms baselines on a dataset with implicit and multi-type table structures and can handle multi-table tables including previously neglected complex tables.
Text Style Transfer Back-Translation (2023.acl-long)

Copied to clipboard

Challenge: Current methods require large amount of bilingual training data, which is challenging and sometimes impossible task.
Approach: They propose a method to modify the style of inputs by modifying the source side of BT data.
Outcome: The proposed method significantly improves translation quality against popular BT benchmarks on high-resource and low-resourced language pairs.
A Neural Question Answering Model Based on Semi-Structured Tables (C18-1)

Copied to clipboard

Challenge: Existing question answering systems rely on raw text and structured knowledge graphs.
Approach: They build an end-to-end system to answer multiple choice questions with semi-structured tables as its knowledge.
Outcome: The proposed system improves on the state-of-the-art question answering system with tabMCQ dataset.
Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents (2024.lrec-main)

Copied to clipboard

Challenge: Existing document AI approaches fail to consider key-value relations in visually-rich documents . a few-shot approach is proposed to extract key- value relation triplets in VRDs .
Approach: They propose a few-shot relational learning approach targeting the extraction of key-value relation triplets in Visually-Rich Documents.
Outcome: The proposed method outperforms existing methods in visually-rich documents.
MoralDial: A Framework to Train and Evaluate Moral Dialogue Systems via Moral Discussions (2023.acl-long)

Copied to clipboard

Challenge: A moral dialogue system aligned with users’ values could enhance conversation engagement and user connections.
Approach: They propose a framework to train and evaluate moral dialogue systems based on communication mechanisms of morality and a method to construct moral discussions between simulated users and the dialogue system.
Outcome: The proposed framework can train and evaluate moral dialogue systems based on simulated users and their values .
Entity-centered Cross-document Relation Extraction (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for relation extraction only use text snippets surrounding target entities in multiple documents.
Approach: They propose a relation-extraction model that uses cross-path entity relation attention to detect the semantic relations between entities in a given text.
Outcome: The proposed method outperforms the state-of-the-art methods in the dataset CodRED by 10%.
Collective Human Opinions in Semantic Textual Similarity (2023.tacl-1)

Copied to clipboard

Challenge: Existing benchmarks for semantic textual similarity (STS) use averaged human ratings as gold standard.
Approach: They propose to use a Chinese sentence-to-sentence dataset to study collective human opinions in semantic textual similarity (STS) neither a scalar nor a single Gaussian fits a set of observed judgments adequately, they argue .
Outcome: The proposed dataset does not capture disagreements on individual instances, but rather the confidence over the aggregate dataset.
RAPID: Efficient Retrieval-Augmented Long Text Generation with Writing Planning and Information Discovery (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for knowledge-intensive long texts struggle with issues like hallucinations, topic incoherence, and significant latency.
Approach: They propose a retrieval-augmented long text generation framework with writing P**lanning and I**nformation to address these challenges.
Outcome: The proposed framework outperforms state-of-the-art methods on a freshWiki-2024 dataset.
MERIT: Multi-Agent Collaboration for Unsupervised Time Series Representation Learning (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to time series representation learning are time-consuming and expert-dependent, which are difficult to generalize across different tasks.
Approach: They propose to use large language model agent to guide unsupervised time series representation learning and a framework to integrate three LLM agents to collaboratively generate positive views for time series data.
Outcome: The proposed framework integrates large language model (LLM) agent to guide unsupervised time series representation learning and compares it with state-of-the-art baselines on multiple time series datasets.
LETI: Learning to Generate from Textual Interactions (2024.findings-naacl)

Copied to clipboard

Challenge: Existing techniques fine-tune on input-output pairs or with numerical rewards that gauge the output quality are not effective.
Approach: They propose to fine-tune pre-trained language models with binary labels and a Python interpreter to get textual feedback from the inputs.
Outcome: The proposed model outperforms the base model on unseen problems and achieves comparable or better performance on humanEval.
SimRPD: Optimizing Recruitment Proactive Dialogue Agents through Simulator-Based Data Evaluation and Selection (2026.acl-industry)

Copied to clipboard

Challenge: High-quality data in training proactive dialogue agents is scarce, despite fine-tuning and reinforcement learning . a recent study has shown that the effectiveness of supervised fine-touring is limited by the lack of high-quality, domain-specific training data.
Approach: They propose a framework for training recruitment proactive dialogue agents using a high-fidelity user simulator and a multi-dimensional evaluation framework based on Chain-of-Intention.
Outcome: The proposed framework outperforms existing simulator-based data selection strategies in a real-world recruitment scenario.
On the Influence of Masking Policies in Intermediate Pre-training (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that inserting an intermediate pre-training stage improves performance of masked language models.
Approach: They propose methods to automate the discovery of optimal masking policies via direct supervision or meta-learning.
Outcome: The proposed method outperforms the heuristic of masking named entities on TriviaQA and can be generalizable beyond that task.
Intra-Event and Inter-Event Dependency-Aware Graph Network for Event Argument Extraction (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models do not build dependency information among event argument roles . Existing methods do not learn the interactions between different roles based on event structure .
Approach: They propose an intra-event and inter-e event dependency-aware graph network to model dependencies between roles . they use event structure as the fundamental unit to construct role dependencies within events .
Outcome: The proposed model improves on the ACE05, RAMS, and WikiEvents datasets.
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio-Language Models suffer from hallucinations, e.g., generating text not grounded in the audio input.
Approach: They propose a framework to address hallucination problems in large audio-language models . they use a preference dataset to test the model's accuracy .
Outcome: The proposed model outperforms the latest SOTA methods in terms of performance and generalization.
Multi-Defendant Legal Judgment Prediction via Hierarchical Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for predicting judgment results for multiple defendants are ineffective.
Approach: They propose a method to predict the judgment results for each defendant in multi-defendant cases . they formalize the multi-diffendant judgment process as hierarchical reasoning chains .
Outcome: The proposed method can predict the judgment results for multiple defendants in multi-defendant cases.
Never Lost in the Middle: Mastering Long-Context Question Answering with Position-Agnostic Decompositional Training (2024.acl-long)

Copied to clipboard

Challenge: Large language models suffer from severe hallucinations, compromising performance in knowledge-oriented QA, dialogue, and writing.
Approach: They propose to enhance the information searching and reflection ability of large language models by training them in position-agnostic multi-step QA tasks to improve their model's accuracy.
Outcome: The proposed model improves in multi-doc QA and other benchmarks by 13.7% absolute gain in shuffled settings, by 21.5% in passage retrieval task.
Prompt Tuning for Unified Multimodal Pretrained Models (2023.findings-acl)

Copied to clipboard

Challenge: Prompt tuning has demonstrated success in natural language pretraining and even vision pretraining.
Approach: They propose to apply prompt tuning to a unified sequence-to-sequence pretrained model by adding a sequence of learnable embeddings to each layer and finetuning the pretrained models on downstream tasks.
Outcome: The proposed method outperforms other parameter-efficient tuning methods on multimodal models and is robust against adversarial attacks.
Token-level Proximal Policy Optimization for Query Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have improved search engines and recommendation systems through their text understanding capabilities.
Approach: They propose a token-level proximal policy optimization approach to empower LLMs to perform better in query generation through fine-tuning.
Outcome: The proposed approach outperforms existing LLMs on an open-source and industrial dataset.
DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are susceptible to generating harmful content when prompted with carefully crafted inputs, a vulnerability known as LLM jailbreaking.
Approach: They propose an end-to-end generative approach for jailbreak rewriting inspired by diffusion models that uses a sequence-tosequence (seq2sequ) diffusion model as a generator, conditioning on the original prompt and guiding the denoising process with a novel attack loss.
Outcome: Experiments on Advbench and Harmbench show that the proposed method outperforms autoregressive jailbreak models across evaluation metrics including ASR, fluency, diversity and diversity.
End-to-End Learnable Psychiatric Scale Guided Risky Post Screening for Depression Detection on Social Media (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to detect depression from social media posting history are limited by frozen screening models and lack of learning.
Approach: They propose to use a frozen screening model to train a risky post detection model with psychiatric scales to enable a learnable end-to-end learning process.
Outcome: The proposed model outperforms several strong baseline methods and qualitative analysis confirms that it better captures users’ mental states than others.
Plug-and-Play Data Module for Code RL: Adaptive Ambiguity Replay (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to reinforcement learning (RL) rely on static, in-epoch metrics that overlook training dynamics, often introducing low-utility or outdated data.
Approach: They propose a plug-and-play module that prioritizes cross-epoch ambiguous samples to neutralize the noise from stale experiences.
Outcome: Extensive experiments on nine LLMs show that Adaptive Ambiguity Replay outperforms state-of-the-art baselines on real-world code editing tasks.
To Pretrain or Not to Pretrain: Examining the Benefits of Pretrainng on Resource Rich Tasks (2020.acl-main)

Copied to clipboard

Challenge: Existing studies on pretraining NLP models with variants of Masked Language Model (MLM) objectives have shown that the number of training samples used in the downstream task is limited.
Approach: They propose to use MLM objectives to pretrain NLP models with variants of Masked Language Model (MLM) objectives to improve accuracy on downstream tasks.
Outcome: The proposed model can reach a diminishing return point as the supervised data size increases significantly.
Self-Critique Guided Iterative Reasoning for Multi-hop Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable reasoning capabilities, but they still face challenges in knowledge-intensive multi-hop reasoning.
Approach: They propose a method that uses self-critique feedback to guide iterative reasoning by enabling iteration and self-evaluation of its intermediate reasoning steps.
Outcome: The proposed method surpasses the previous SOTA by 8.6% on three multi-hop reasoning datasets.
MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods neglect stylistic modeling and rely on static thresholds, which greatly limits the detection performance.
Approach: They propose a framework that enables stylistics-aware uncertainty quantification through conditional threshold estimation.
Outcome: The proposed framework achieves an average improvement 11.34% in detection performance compared to baselines.
TLUE: A Tibetan Language Understanding Evaluation Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Low-resource languages, like Tibetan, remain underrepresented in large language models' evaluations.
Approach: They propose a Tibetan Language Understanding Evaluation Benchmark to assess LLMs' proficiency in Tibetan . they use a multi-task understanding benchmark and a safety benchmark to evaluate models .
Outcome: The proposed benchmark shows that most large language models perform below the random baseline, especially in Tibetan language processing.
Retrieved In-Context Principles from Previous Mistakes (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in in-context learning (ICL) have limited customization and inadequate error coverage.
Approach: They propose a method to retrieve in-context principles from mistakes to improve model performance.
Outcome: The proposed framework enhances model performance when applied to various prompting strategies.
DVD: Dynamic Contrastive Decoding for Knowledge Amplification in Multi-Document Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) generate information with hallucinations due to uneven retrieval quality and irrelevant contents.
Approach: They propose a decoding strategy which dynamically amplifies knowledge from selected documents during the generation phase.
Outcome: The proposed method outperforms other decoding strategies on ALCE-ASQA, NQ, TQA and PopQA benchmarks.
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training–Inference Mismatch (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has shown that reinforcement learning with simple rule-based reward functions (RLVR) can induce emergent reasoning behaviors and yield gains in challenging domains such as math problem solving.
Approach: They propose a rollout-alignment-quantization-aware RL which aligns training-side forward with the quantized rollout to minimize mismatch.
Outcome: The proposed approach outperforms quantized-rollout training by +5.5 on Qwen3-30B-A3B MoE for math problems while maintaining low-bit throughput.
Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning (2023.emnlp-main)

Copied to clipboard

Challenge: In-context learning (ICL) is a promising capability for large language models (LLMs) but its underlying mechanism remains unexplored.
Approach: They propose a demonstration compression technique to expedite inference and an analysis framework for diagnosing ICL errors in GPT2-XL.
Outcome: The proposed method improves ICL performance and expedites inference.
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for group-relative policy optimization rely on scalar correctness rewards that are often non-injective with respect to semantic content.
Approach: They propose a framework that calibrates the reward signal using the semantic density of sampled groups.
Outcome: The proposed framework outperforms strong baselines on five math benchmarks with 7,000 samples and 55 cost.
PolCLIP: A Unified Image-Text Word Sense Disambiguation Model via Generating Multimodal Complementary Representations (2024.acl-long)

Copied to clipboard

Challenge: Existing models for word sense disambiguation lack images or senses in textual and visual datasets.
Approach: They propose a unified image-text WSD model that uses image-sense complementarity to generate visual representations for word senses and a disambiguation-oriented image-sensor dataset to provide implicit textual representations.
Outcome: The proposed model achieves 2.53% F1-score increase over state-of-the-art models on Textual-WSD and 2.22% HR@1 improvement on Visual-WSS.
Effective Long-Context Scaling of Foundation Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are rapidly deployed and continue to evolve through scaling.
Approach: They propose a method to train strong long-context LLMs that are capable of utilizing massive context windows of up to 32,000 tokens.
Outcome: The proposed model can surpass gpt-3.5-turbo-16k's overall performance on long-context benchmarks with a cost-effective instruction tuning procedure that is free of expensive annotations.
Enhanced Multi-Channel Graph Convolutional Network for Aspect Sentiment Triplet Extraction (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to extract aspect triplets ignore the relationships between words . Enhanced Multi-Channel Graph Convolutional Network model can be used to learn relation-aware node representations.
Approach: They propose an Enhanced Multi-Channel Graph Convolutional Network model to fully utilize the relations between words for ASTE task.
Outcome: The proposed model outperforms state-of-the-art methods significantly on a benchmark dataset.
Tiny Scales, Great Challenges: The Limits of Multimodal LLMs in Scale Recognition (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on a single type of quantity or a specific format, lacking a comprehensive evaluation of scale recognition capabilities.
Approach: They propose a visual scale recognition benchmark built using images from COCO, Open Images, and Flickr to evaluate scale recognition capabilities of multimodal large language models.
Outcome: The proposed model achieves 42.60% accuracy, lower than the 97.40% of humans.
Better Zero-Shot Reasoning with Role-Play Prompting (2024.naacl-long)

Copied to clipboard

Challenge: Recent years have witnessed a paradigm shift in natural language processing, driven by large language models such as GPT-3, PaLM, and Llama.
Approach: They propose a strategy for role-play prompting and assess its performance under the zero-shot setting.
Outcome: The proposed method outperforms the standard zero-shot prompting approach across 12 reasoning benchmarks.
Glancing Transformer for Non-Autoregressive Neural Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: Existing non-autoregressive neural machine translation methods are either inferior to Transformer or require multiple decoding passes, leading to reduced speedup.
Approach: They propose a Glancing Language Model (GLM) for single-pass parallel generation models and Glancing Transformer (GLAT) with only single- pass decoding, GLAT is able to generate high-quality translation with 8-15 speedup.
Outcome: The proposed model outperforms all previous non-autoregressive methods on multiple language directions and is nearly comparable to Transformer.
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have remarkable reasoning capabilities in complex tasks such as mathematics and coding.
Approach: They propose an entropy-modulation method that adaptively reweighs tokens based on theoretically-estimated entropic variations.
Outcome: The proposed method outperforms state-of-the-art methods in six mathematical reasoning and three coding benchmarks.
Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Event extraction is a task in natural language processing that involves identifying and extracting event information from unstructured text.
Approach: They propose a paradigm that combines schema paraphrasing with schema retrieval-augmented generation.
Outcome: The proposed paradigm retrieves paraphrased schemas and accurately generates targeted structures.
UNIMO-2: End-to-End Unified Vision-Language Grounded Learning (2022.findings-acl)

Copied to clipboard

Challenge: Existing methods for vision-language pre-training can only learn from aligned image-caption data and rely heavily on expensive regional features.
Approach: They propose an end-to-end unified-modal pre-training framework for joint learning . they propose to conduct grounded learning on both images and texts via a sharing grounded space .
Outcome: The proposed model improves visual and visual semantic alignment on images and texts.
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization (2025.findings-emnlp)

Copied to clipboard

Challenge: HKVE selectively accepts gradient optimization results based on the distribution of attention scores across different layers, ensuring that every optimization step positively contributes to the attack.
Approach: They propose a framework that selectively accepts gradient optimization results based on the distribution of attention scores across different layers and selectively takes them into account when calculating the attack success rate.
Outcome: The proposed framework outperforms existing methods by achieving success rates of 75.08% on MiniGPT4, 85.84% on LLaVA and 81.00% on Qwen-VL.
LEPO: Latent Reasoning Policy Optimization for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing latent reasoning methods that use chain of thought (CoT) are limited to selecting one discrete token at each reasoning step, which potentially induces information loss.
Approach: They propose a framework that injects controllable stochasticity into latent reasoning via Gumbel-Softmax, restoring LLMs' exploratory capacity and enhancing their compatibility with Reinforcement Learning (RL).
Outcome: The proposed framework preserves richer information for more comprehensive reasoning and is compatible with Reinforcement Learning (RL).
Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context Learning (2025.acl-long)

Copied to clipboard

Challenge: Recent learning-based demonstration selection methods have proven beneficial to in-context learning (ICL) by choosing more useful exemplars.
Approach: They propose two methods to capture task-agnostic similarities between input and output of LLMs.
Outcome: The proposed methods integrate task-agnostic similarities of different levels between input and output of exemplars and test cases to eliminate costly data collection.
RAST: Domain-Robust Dialogue Rewriting as Sequence Tagging (2021.emnlp-main)

Copied to clipboard

Challenge: Existing models for dialogue rewriting suffer from the robustness issue, i.e., performances drop dramatically when testing on a different dataset.
Approach: They propose a sequence-tagging-based approach that reduces the search space while preserving the core of the task.
Outcome: The proposed model significantly reduces the search space while still covering the core of the task.
Uncertainty Calibration for Tool-Using Language Agents (2024.findings-emnlp)

Copied to clipboard

Challenge: Language agents are increasingly used to perform tasks and interact with a variety of external tools to achieve specific, goal-oriented objectives.
Approach: They propose a tool calibration tool called ProbeCal which recalibrates the internal probabilities of tool-using language agents to better reflect the actual effectiveness of tool.
Outcome: The proposed model significantly improves off-the-shelf language models in tool-using applications.
latent-GLAT: Glancing at Latent Variables for Parallel Text Generation (2022.acl-long)

Copied to clipboard

Challenge: Recent advances in text generation have limited applications due to multimodality problem.
Approach: They propose a method which uses latent variables to capture word categorical information and invoke an advanced curriculum learning technique to overcome multi-modality problem.
Outcome: The proposed method outperforms strong baselines without an autoregressive model, which further broadens the application scenarios of the parallel decoding paradigm.
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error (2024.acl-long)

Copied to clipboard

Challenge: Existing work on tool-augmented LLMs focuses on the broad coverage of tools and the flexibility of adding new tools.
Approach: They propose a biologically inspired method for tool-augmented LLMs that orchestrates three key mechanisms for successful tool use behaviors in the biological system: trial and error, imagination, and memory.
Outcome: The proposed method improves tool learning for LLMs under both in-context learning and fine-tuning settings, bringing a boost of 46.7% to Mistral-Instruct-7B and outperforms GPT-4.
SafetyMem: Adaptive Jailbreak Defense via Dual-Component Safety Memory (2026.acl-long)

Copied to clipboard

Challenge: Existing defenses for Large Language Models suffer from a 'memory gap' parameter-modifying methods are computationally expensive and inference-time filters cannot retain or reuse defense knowledge across interactions.
Approach: They propose a framework that secures Large Language Models through a dual-component safety memory system.
Outcome: The proposed framework significantly reduces attack success rates while preserving interpretability and efficiency.
BalanceSFT: Improving LLM Function Calling with Balanced Training Signals and Data Hardness (2026.findings-acl)

Copied to clipboard

Challenge: Currently, Supervised Fine-Tuning (SFT) is the prevailing method for equipping Large Language Models (LLMs) with function calling capabilities, but its effectiveness is often compromised by two challenges: 1) lengthy Chain-of-Thought (CoT) reasoning tokens dominate training signals over concise function calls in the learning objective; 2) scarcity of hard training examples.
Approach: They propose a framework that uses a self-adjusted signal balancing loss and a hard data re-sampling strategy to selectively generate new, high-quality complex data guided by model errors.
Outcome: The proposed framework surpasses state-of-the-art models like GPT-5 in function calling performance.
DocTrack: A Visually-Rich Document Dataset Really Aligned with Human Eye Movement for Machine Reading (2023.findings-emnlp)

Copied to clipboard

Challenge: Document AI models that can read visually rich documents have a long way to go before they can read them as accurately, continuously, and flexibly as humans do.
Approach: They propose a visually-rich document dataset that aligns with human eye-movement information using eye-tracking technology.
Outcome: The proposed dataset can help in designing better document AI models and human reading robots in the future.
Tree-CoT-RT: An Explainable Multi-Path Tree-Guided Chain-of-Thought and Reinforcement Learning Framework for Aspect Sentiment Quad Prediction (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods lack explainability and generalization, making it difficult to justify inference decisions and detect implicit sentiment across domains and varied expression patterns.
Approach: They propose an explainable multi-path tree-guided chain-of-thought framework specifically designed for ASQP.
Outcome: Experiments on benchmark datasets show that Tree-CoT-RT outperforms baselines.
Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger (2025.acl-long)

Copied to clipboard

Challenge: Existing research expands the tool arrays of large language models (LLMs), but the necessity of using these tools is often overlooked, leading to indiscriminate tool invocation.
Approach: They propose a meta-cognition proxy proxy for LLMs self-assessment of their capabilities, reflecting the model’s awareness of its own limitations.
Outcome: The proposed strategy is fine-tuned-free and costs minimal.
Learning to Generate Question by Asking Question: A Primal-Dual Approach with Uncommon Word Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing automatic question generation methods focus on encoding passage and answer to generate question.
Approach: They propose an automatic question generation approach which integrates question generation with its dual problem, question answering, into a unified primal-dual framework.
Outcome: The proposed approach outperforms existing methods on SQuAD and HotpotQA benchmarks.
OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Code LLMs lack reproducible data pipelines and training protocols for reproducible advancements in code intelligence.
Approach: They propose a top-tier code LLM that releases model weights and inference code . reproducible data pipelines, rigorous experimental ablation results and training protocols are included .
Outcome: The proposed model achieves comparable performance to leading models and serves as an "open cookbook" reproducible training data, rigorous experimental ablation results, and detailed training protocols are also included in the model.
Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to enhance reasoning capabilities of language models are expensive and often lack the ability to perform complex reasoning tasks.
Approach: They propose a token-level multi-model collaboration strategy to enhance reasoning capabilities in language models by selecting the optimal tokens from the next token distributions.
Outcome: The proposed method is superior to existing methods and will be released soon.
ScholarGraph:a Chinese Knowledge Graph of Chinese Scholars (L18-1)

Copied to clipboard

Challenge: ScholarSpace integrates chinese academic information from chin scholars and science . data integration system needs to be focused on scholars, says dr. s. k. o. j. nielson .
Approach: a data integration system is built to integrate chinese academic information from chin scholars and science. a system can give you an academic portrait about a chinoise scholar with the form of a knowledge graph.
Outcome: a data integration system called ScholarSpace can integrate chinese academic information from chin scholars and science.
KoCo-Bench: Can Large Language Models Leverage Domain Knowledge in Software Development? (2026.acl-long)

Copied to clipboard

Challenge: Existing domain-specific code benchmarks focus on assessing what knowledge LLMs possess rather than how they acquire and apply new knowledge.
Approach: They propose a benchmark to evaluate domain specialization methods in real-world software development.
Outcome: KOCO-bench is a new benchmark for evaluating domain specialization methods in real-world software development.
GRV-KBQA: A Three-Stage Framework for Knowledge Base Question Answering with Decoupled Logical Structure, Semantic Grounding and Structure-Aware Validation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for Knowledge Base Question Answering generate non-executable queries and inefficiencies in query execution.
Approach: a framework that decouples logical structure generation from semantic grounding is proposed . the framework explicitly enforces KB constraints to improve alignment between generated logical forms and KB structures.
Outcome: GRV-KBQA decouples logical structure generation from semantic grounding and incorporates structure-aware validation to enhance accuracy.
GLiM: Integrating Graph Transformer and LLM for Document-Level Biomedical Relation Extraction with Incomplete Labeling (2025.findings-acl)

Copied to clipboard

Challenge: Document-level relation extraction (DocRE) solves problems of document quality . number of entities and entity-pair relations increases, causing incomplete annotations .
Approach: a framework that reduces the problem space using a graph-enhanced Transformer-based model is proposed . GLiM leverages large language models for reasoning to reduce the problem-space .
Outcome: GLiM boosts average recall and F1 scores on biomedical datasets . compared with existing models, GLim outperforms existing models on biomedicine benchmarks compared to existing models .
AceGPT, Localizing Large Language Models in Arabic (2024.naacl-long)

Copied to clipboard

Challenge: Significant concerns emerge when addressing cultural sensitivity and local values.
Approach: They propose a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models.
Outcome: The proposed model sets the state-of-the-art standard for open Arabic LLMs across various benchmarks.
Compilable Neural Code Generation with Compiler Feedback (2022.findings-acl)

Copied to clipboard

Challenge: Existing deep-learning approaches model code generation as text generation, but few of them account for compilability of the generated programs.
Approach: They propose a three-stage pipeline utilizing compiler feedback for compilable code generation to improve compilability.
Outcome: The proposed pipeline improves compilability of generated programs by combining compiler feedback, language model fine-tuning, and compilable discrimination.
QAEncoder: Towards Aligned Representation Learning in Question Answering Systems (2025.acl-long)

Copied to clipboard

Challenge: Modern QA systems entail retrieval-augmented generation (RAG) for accurate and trustworthy responses, but the inherent gap between user queries and relevant documents hinders precise matching.
Approach: They propose a retrieval-augmented generation (RAG)-based approach to bridge this gap by attaching document fingerprints to the embedding to estimate the expectation of potential queries.
Outcome: Experiments across diverse datasets, languages, and embedding models confirm the proposed solution is simple-yet-effective with zero additional index storage, retrieval latency, training costs, or catastrophic forgetting and hallucination issues.
A Relaxed Matching Procedure for Unsupervised BLI (2020.acl-main)

Copied to clipboard

Challenge: Recent studies have shown that unsupervised bilingual lexicon induction is even on par with supervised methods.
Approach: They propose a relaxed matching procedure to find a more precise matching between two languages by aligning source and target embedding space bidirectionally.
Outcome: The proposed method significantly outperforms previous unsupervised methods on standard benchmarks.
Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing text watermarking technologies lack consistency when texts are translated into different languages.
Approach: They propose a cross-lingual watermark removal attack to bypass watermarking by first obtaining a response from an LLM in a pivot language and then translating it into the target language.
Outcome: The proposed method can remove watermarks without performance loss by obtaining a response from an LLM in a pivot language and then translating it into the target language.
Explainable Recommendation with Personalized Review Retrieval and Aspect Learning (2023.acl-long)

Copied to clipboard

Challenge: Recent years have witnessed a growing interest in the development of explainable recommendation models.
Approach: They propose a model that combines prediction and generation tasks to produce more persuasive explanations by obtaining additional information from the training sets.
Outcome: The proposed model outperforms state-of-the-art models on three datasets and shows that it is more persuasive than previous models.
Modelling Long-distance Node Relations for KBQA with Global Dynamic Graph (2020.coling-main)

Copied to clipboard

Challenge: Existing studies rely on deep graph neural networks (GNNs) to capture rich structural information, but they lack the structural information needed for QA.
Approach: They propose a framework which captures structural information from KBs and models long-distance node relations from two perspectives.
Outcome: The proposed framework models long-distance node relations from two perspectives . it is based on two widely used multi-hop KBQA datasets .
Advancing Sequential Numerical Prediction in Autoregressive Models (2025.acl-short)

Copied to clipboard

Challenge: Autoregressive models are the de facto choice for sequence generation tasks, but standard approaches treat digits as independent tokens and apply cross-entropy loss, overlooking the coherent structure of numerical sequences.
Approach: They propose a novel approach to entropy loss by extending the Earth Mover’s Distance to preserve ordinal relationships between numerical values and sequence-level to penalize the overall discrepancy between predicted and actual sequences.
Outcome: Extensive experiments show that NTIL improves numerical prediction and integrates effectively with LLMs/MLLMs.
Learning to Simulate Natural Language Feedback for Interactive Semantic Parsing (2023.acl-long)

Copied to clipboard

Challenge: Existing work on interactive semantic parsing relies on human annotations to train a model . prior work relied on human-annotated feedback data, which is prohibitively expensive and not scalable .
Approach: They propose a task of simulating NL feedback for interactive semantic parsing . they propose evaluators to assess the quality of the simulated feedback .
Outcome: The proposed simulator can generate high-quality NL feedback to boost the error correction ability of a specific parser.
SciAgent: Tool-augmented Language Models for Scientific Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: SciAgent surpasses other LLMs with the comparable size by more than 8.0% in absolute accuracy.
Approach: They propose a tool-augmented scientific reasoning setting that supplements LLMs with scalable toolsets and builds a benchmark to evaluate LLM’s abilities with tool assistance.
Outcome: The proposed setting augments LLMs with scalable toolsets and shifts the focus from pursuing an omniscient problem solver to a proficient tool-user.
An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor Isomorphism (2022.acl-long)

Copied to clipboard

Challenge: Existing methods focus on graph representation learning, but decoding is a key part of the process.
Approach: They propose an EA Decoding Algorithm via Third-order Tensor Isomorphism (DATTI) they combine two sets of isomorphic equations to enhance the decoding process .
Outcome: The proposed algorithm can deliver significant performance improvements even on the most advanced methods while the extra required time is less than 3 seconds.
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings (2024.emnlp-main)

Copied to clipboard

Challenge: Attaching suffixes to harmful instructions can hack the defense of Large language models (LLMs) However, due to the unreadable of adversarial suffix, it can be relatively easily penetrated by common defense methods such as perplexity filters.
Approach: They propose an algorithm to embed adversarial suffixes into coherent and understandable text to attack Large language models (LLMs) using a Advbench dataset.
Outcome: The proposed approach reduces the computation time of adversarial suffixes and achieves a much better attack success rate than existing techniques.
MAF: Multimodal Alignment Framework for Weakly-Supervised Phrase Grounding (2020.emnlp-main)

Copied to clipboard

Challenge: Existing work on phrase localization uses caption-image datasets as weak supervision . existing work on supervised phrase localisation uses a large-scale annotated dataset .
Approach: They develop a multimodal alignment framework to leverage more widely available caption-image datasets to model phrase relevance.
Outcome: The proposed model improves on the widely-adopted Flickr30k dataset . it also improves the previous best unsupervised result by 5.56% .
Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support Conversation (2026.acl-long)

Copied to clipboard

Challenge: Existing models for ESC ignore cognitive distortions in help-seekers' expressions . current models provide basic emotional comfort, rather than helping help- seekers address psychological distress at a deeper cognitive level.
Approach: They propose a Large Language Model framework to enhance LLMs' ability to diagnose and intervene cognitive distortions in help-seekers.
Outcome: The proposed framework outperforms 15 state-of-the-art baselines in terms of distortion diagnosis accuracy, intervention strategy effectiveness, and safety risk control.
ERNIE-Code: Beyond English-Centric Cross-lingual Pretraining for Programming Languages (2023.findings-acl)

Copied to clipboard

Challenge: ERNIE-Code is a unified pre-trained language model for 116 NLs and 6 PLs.
Approach: They propose a unified pre-trained language model for 116 NLs and 6 PLs . they employ span-corruption language modeling that learns patterns from monolingual NL or PL .
Outcome: The proposed model outperforms previous multilingual models for NL or NL across end tasks.
MAVEN-FACT: A Large-scale Event Factuality Detection Dataset (2024.findings-emnlp)

Copied to clipboard

Challenge: Event factuality detection is under-explored due to the lack of high-quality large-scale data . efd is a subfield of event understanding, which aims to determine the factuity of textual events.
Approach: They propose a large-scale EFD dataset with factuality annotations of 112,276 events . they find that adopting event arguments and relations helps in event factuity detection .
Outcome: The proposed dataset includes factuality annotations of 112,276 events . it is the largest EFD dataset and is challenging for fine-tuned models and large language models .
Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Existing studies on self-supervised pretraining for machine translation have focused on the jointly pretrained decoder .
Approach: They propose a method to improve neural machine translation by jointly pretrained decoder . they propose two strategies to remedy the domain and objective discrepancies .
Outcome: The proposed approach improves translation performance and model robustness on three language pairs.
Benchmarking Large Language Models on Communicative Medical Coaching: A Dataset and a Novel System (2024.findings-acl)

Copied to clipboard

Challenge: Existing applications of natural language processing (NLP) focus on patient-centered services, but the potential of NLP to benefit inexperienced doctors remains unexplored.
Approach: They propose a human-AI cooperative framework to assist medical learners in practicing communication skills during patient consultations.
Outcome: The proposed framework enables medical learners to practice communication skills during patient consultations while a coach agent provides immediate, structured feedback.
Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation (2026.acl-long)

Copied to clipboard

Challenge: Current systems often fall short of this goal in settings where translation hinges on culturally grounded entities such as books, films, places, songs and idioms.
Approach: They propose a framework that anchors supervision on a verifiable, entity-level reward signal and incorporates lightweight structural gates to stabilize optimization.
Outcome: The proposed framework improves on XC-Translate and shows that it can learn a robust reasoning process rather than imitating reference translations.
FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches that extend the mask language modeling to other modalities require careful multi-task tuning, complex reconstruction target designs, or additional pre-training data.
Approach: They propose a centralized multimodal graph contrastive learning strategy to unify self-supervised pre-training for all modalities in one loss.
Outcome: The proposed model achieves state-of-the-art performance on FUNSD, CORD, SROIE and Payment benchmarks with a more compact model size.
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents (2025.findings-acl)

Copied to clipboard

Challenge: Conversational agents are increasingly woven into individuals’ personal lives, yet users underestimate the privacy risks associated with them.
Approach: They propose a framework that allows users to reformulate out-of-context information in user prompts by identifying and reformulating out- of-content information in the context.
Outcome: The proposed framework can achieve strong gains in contextual privacy while preserving the user’s intended interaction goals.
LOG: A Local-to-Global Optimization Approach for Retrieval-based Explainable Multi-Hop Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches to multi-hop question answering emphasize single-step and multi-step iterative decomposition or retrieval, which are susceptible to failure in long-chain reasoning due to the progressive accumulation of erroneous information.
Approach: They propose a Local-tO-Global optimized retrieval method to discover more beneficial information and improve tuplet objective loss.
Outcome: The proposed method outperforms state-of-the-art models and significantly improves multi-hop reasoning.
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks (2026.acl-long)

Copied to clipboard

Challenge: Existing large language models for software engineering rely on coarse-grained pass rates obscuring specific cognitive bottlenecks.
Approach: They propose a repository-level benchmark that dissects coding capabilities through atomized tasks.
Outcome: The proposed framework achieves a 78.55% validity yield, surpassing the 31.7% retention rate of SWE-bench-Verified.
VulAgent: Hypothesis-Validation Driven Multi-Agent Architecture for Vulnerability Detection (2026.findings-acl)

Copied to clipboard

Challenge: Recent reports indicate that software vulnerabilities caused by insecure coding practices remain a major security threat.
Approach: They propose a multi-agent vulnerability detection framework based on hypothesis validation . they use multi-view analyzers to localize and localize security-sensitive operations .
Outcome: The proposed framework reduces false positives and increases accuracy by 6.6 percentage points on PrimeVul and SVEN.
HCRE: LLM-based Hierarchical Classification for Cross-Document Relation Extraction with a Prediction-then-Verification Strategy (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to cross-document relation extraction (RE) focus on identifying relations between head and tail entities from single sentence or document.
Approach: They propose a hierarchical relation tree-based LLM-based hierarchic classification model for cross-document relation extraction (HCRE) based on predefined relations, the model can perform hierarchically classification level by level.
Outcome: The proposed model outperforms existing baselines and validates its effectiveness.
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Mixture-of-Experts (MoE) scales capacity via conditional computation, but lacks knowledge lookup primitive.
Approach: They propose a conditional memory instantiated via Deep Sparse Embedding (DSE) they propose 'u-shaped scaling law' that identifies optimal balance between MoE experts and DSE memory .
Outcome: The proposed model outperforms an iso-parameter and isoFLOPs MoE baseline across knowledge and reasoning benchmarks and is infrastructure-efficient.
MRT: Multi-modal Short- and Long-range Temporal Convolutional Network for Time-sync Comment Video Behavior Prediction (2024.lrec-main)

Copied to clipboard

Challenge: Using time-sync comments, it is difficult to understand user behavior due to complexity of interactions between users, videos, and comments.
Approach: They propose a novel time-sync comment behavior prediction model that takes historical behavior into account and optimizes it on the basis of user preferences.
Outcome: The proposed model improves the performance of time-sync comments on visual frames and textual comments on two cats playing simultaneously.
UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation (2023.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have impressive capabilities but need for task-specific prompt engineering can hinder their generalization.
Approach: They propose a lightweight and versatile retriever that automatically retrieves prompts for a given zero-shot task input.
Outcome: The proposed model is universally applicable across tasks and models . it mitigates hallucination problem in chatGPT, and it improves even the strongest LLMs.
SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Low-resource language tokens are often routed to different experts than those activated by high-resourced inputs, which hinders their efficacy in multilingual contexts.
Approach: They propose a framework to transfer specialized capabilities from high-resource languages as anchors to low-resourced languages by using a symmetric Jensen-Shannon constraint.
Outcome: The proposed framework outperforms standard instruction tuning on 5 low-resource languages and 3 benchmarks.
Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human Evaluations (2026.acl-long)

Copied to clipboard

Challenge: Existing static strategies for mitigating hallucinations do not explicitly model the information gain from interacting with the external environment.
Approach: They propose a calibration-driven interactive learning strategy that selects clarification queries by optimizing calibration error.
Outcome: The proposed method provides theoretical guarantees and empirical gains for reliability.
LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization (2025.emnlp-main)

Copied to clipboard

Challenge: Recent research focuses on integrating reasoning capabilities into the realm of retrieval-augmented generation (RAG) via outcome-supervised reinforcement learning (RL).
Approach: They propose a process-level reward module to mitigate the unawareness of intermediate reasoning steps in outcome-level supervision without additional annotation.
Outcome: The proposed framework can boost LLMs’ reasoning ability by integrating external knowledge sources through retrieval-augmented generation (RAG) The proposed model can mitigate the unawareness of intermediate reasoning steps in outcome-level supervision without additional annotation.
Interpret and Control Dense Retrieval with Sparse Latent Features (2025.naacl-short)

Copied to clipboard

Challenge: Dense embeddings deliver strong retrieval performance but lack interpretability and controllability.
Approach: They propose a novel approach using sparse autoencoders to interpret and control dense embeddings via latent sparsity.
Outcome: The proposed approach retains the same retrieval accuracy as the original dense vectors, affirming their faithfulness.
An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a powerful tool for test-time scaling, but they are often used under time constraints.
Approach: They propose to use LLMs to make models think before answering questions . they also use self-correction and best-of-N decoding to encourage deeper thinking .
Outcome: The proposed models are able to achieve higher inference accuracy with extra inference computation under time constraints.
Large Language Models Might Not Care What You Are Saying: Prompt Format Beats Descriptions (2025.findings-emnlp)

Copied to clipboard

Challenge: In-context learning has improved performance of large language models, but descriptive instructions are still under-explored.
Approach: They propose an ensemble prompt framework to describe selection criteria of multiple in-context examples. preliminary experiments on machine translation confirm that this framework boosts ICL performance.
Outcome: The proposed framework improves on commonsense, math, logical reasoning and hallucination tasks with three LLMs.
Hopscotch: Discovering and Skipping Redundancies in Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Hopscotch is a method that skips attention blocks with least contributions to a task . it does not modify model weights or require access to pretraining or instruction-tuning data.
Approach: Hopscotch proposes a method that skips attention blocks with least contributions to a task . it introduces lightweight scaling parameters to attention and MLP blocks .
Outcome: The proposed method reduces the drop in performance by 2% even after skipping four attention blocks.
Self-supervised Preference Optimization: Enhance Your Language Model with Preference Degree Awareness (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have focused on replacing the reward model in Reinforcement Learning with Human Feedback (RLHF) methods for Large Language Models (LLMs).
Approach: They propose a self-supervised preference optimization framework that replaces the reward model with a preference loss and alignment loss to improve LLMs' ability to understand human preferences.
Outcome: The proposed framework can be integrated with existing preference optimization methods and significantly boost their performance.
IoTMigrator: LLM-driven Embedded IoT Code Migration across Different OSes for Cloud-device Integration (2025.findings-emnlp)

Copied to clipboard

Challenge: Neither outline-based code generation nor common code translation techniques can adequately address this challenge, despite their prevalence in existing systems.
Approach: They have developed an algorithm that employs a multi-agent pipeline to handle embedded code migration under the TSL paradigm.
Outcome: The proposed algorithm outperforms the baseline by 50.5% for pass rate and 13.0% for completeness across all tasks in RIOT and Zephyr.
MTRouter: Cost-Aware Multi-Turn LLM Routing with History–Model Joint Embeddings (2026.acl-long)

Copied to clipboard

Challenge: Multi-turn, long-horizon tasks require dozens of sequential model calls per episode.
Approach: They propose a cost-aware multi-turn LLM routing tool which encodes interaction history and candidate models into joint history–model embeddings and learns an outcome estimator from logged trajectories to predict turn-level model utility.
Outcome: The proposed model reduces cost and performance by 58.7% on ScienceWorld and on Humanity’s Last Exam (HLE) and even reduces costs for held-out tasks.
Improving Text Generation with Student-Forcing Optimal Transport (2020.emnlp-main)

Copied to clipboard

Challenge: Maximum likelihood estimation (MLE) is used to train models, but during testing, the model is conditioned on previously generated tokens, resulting in exposure bias.
Approach: They propose to use optimal transport to match the sequences generated in MLE and test modes to reduce exposure bias.
Outcome: The proposed method is validated on machine translation, text summarization, and text generation tasks.
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems (2025.acl-long)

Copied to clipboard

Challenge: Existing reward models focus on human preferences, neglecting verifiable correctness signals.
Approach: They propose a reward system that combines human preference rewards with verifiable correctness signals to provide reliable rewards.
Outcome: The proposed reward agent significantly outperforms vanilla reward models on benchmarks and inference-time best-of-n searches on real-world tasks.
ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language Understanding (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to model coarse-grained linguistic information do not integrate coarse-gram information into pre-training.
Approach: They propose an explicitly n-gram masking method to enhance integration of coarse-grained linguistic information into pre-training.
Outcome: The proposed method outperforms existing models on English and Chinese text corpora and fine-tunes on 19 downstream tasks.
Multi-Granularity Self-Attention for Neural Machine Translation (D19-1)

Copied to clipboard

Challenge: Existing neural machine translation models use a deep multi-head self-attention network with no explicit phrase information.
Approach: They propose a neural network that combines multi-head self-attention and phrase modeling to train attention heads to attend to phrases in either n-gram or syntactic formalisms.
Outcome: The proposed approach improves on English-to-German and NIST Chinese-to English translation tasks.
CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved remarkable performance on various NLP tasks, yet their potential in more challenging task like finance, has not been fully explored.
Approach: They propose a benchmark to assess the financial knowledge of large language models (LLMs) in China.
Outcome: The proposed benchmark is the most comprehensive evaluation benchmark to date for LLMs in finance.
IDPG: An Instance-Dependent Prompt Generation Method (2022.naacl-main)

Copied to clipboard

Challenge: Existing prompt tuning methods use a fixed prompt in each input instance during the model training stage.
Approach: They propose a conditional prompt generation method to generate prompts for each input instance.
Outcome: The proposed method outperforms other prompt tuning methods while tuning fewer parameters.
HopRAG: Multi-Hop Reasoning for Logic-Aware Retrieval-Augmented Generation (2025.findings-acl)

Copied to clipboard

Challenge: Traditional retrieval systems focus on lexical or semantic similarity rather than logical relevance.
Approach: They propose a new RAG framework that augments retrieval with logical reasoning . hopRAG uses a retrieve-reason-prune mechanism to explore multi-hop neighbors .
Outcome: The proposed framework outperforms conventional retrieval systems and state-of-the-art benchmarks on multi-hop QA tasks.
Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study (2025.findings-acl)

Copied to clipboard

Challenge: Inference-time scaling of chain-of-thought (CoT) has been demonstrated as a promising approach for addressing multi-modal reasoning tasks.
Approach: They propose to integrate visual and textual modalities within the reasoning process . they adopt a consistency-enhanced verifier to ensure effective guidance for both methods across different thought paradigms.
Outcome: The proposed method outperforms text-only reasoning on 10 tasks spanning diverse domains and requires higher token consumption for processing richer visual inputs.
LongTableBench: Benchmarking Long-Context Table Reasoning across Real-World Formats and Domains (2025.findings-emnlp)

Copied to clipboard

Challenge: Evaluating 52 LLMs reveals that only the strongest models maintain robust performance under increasing context lengths and format diversity.
Approach: They propose a benchmark for evaluating long-context reasoning over semi-structured tables across diverse formats, tasks, and domains.
Outcome: The proposed model outperforms compression-based approaches on tasks requiring semantic integration.
Bayes-enhanced Lifelong Attention Networks for Sentiment Classification (2020.coling-main)

Copied to clipboard

Challenge: Existing deep learning paradigms focus on learning a model from training data of a single task and the learned model is also tested on the same task.
Approach: They propose a Bayes-enhanced lifelong attention network to learn attention knowledge from a sequence of sentiment classification tasks and build lifelong ones.
Outcome: The proposed model is able to learn attention knowledge from a set of sentiment classification tasks and build lifelong attentions.
InterIDEAS: Philosophical Intertextuality via LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a new dataset aims to bridge philosophy, literary studies, and natural language processing (NLP) by integrating theories of intertextuality with bibliometric techniques.
Approach: They propose a dataset that bridges philosophy, literary studies, and natural language processing (NLP) it combines theories of intertextuality from literary studies with bibliometric techniques and recent LLMs .
Outcome: a new dataset bridges philosophy, literary studies, and natural language processing (NLP) to analyze intertextuality . the proposed method helps scholars understand the intellectual, social, and historical relations embedded in texts . it also contributes to the development of language models, authors say .
On Unifying Misinformation Detection (2021.naacl-main)

Copied to clipboard

Challenge: On any given day, 2.5 quintillion bytes of information are created on the Internet, a figure that is only expected to increase in the coming years.
Approach: They propose a general-purpose misinformation model that jointly models multiple domains of misinformation with a single, unified setup.
Outcome: The proposed model is useful for few-shot learning of unseen misinformation tasks/datasets and generalizability to unseense events.
NavA3: Understanding Any Instruction, Navigating Anywhere, Finding Anything (2026.acl-long)

Copied to clipboard

Challenge: Existing embodied navigation methods struggle with such tasks due to their limitations in comprehending high-level human instructions and localizing objects with an open vocabulary.
Approach: They propose a hierarchical framework for long-horizon navigation that integrates human instructions with 3D scene views.
Outcome: The proposed model achieves SOTA results and can complete long-horizon navigation tasks across different robot embodiments in real-world environments.
Rethinking Translation Memory Augmented Neural Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to enhance neural machine translation (NMT) by using a TM have been reported to be effective.
Approach: They propose a translation memory augmented neural machine translation model that is good at fitting data but more sensitive to fluctuations in training data.
Outcome: The proposed model achieves consistent gains over conventional and existing models under two variance-preferable scenarios as well as the high resource scenario.
Exploiting Unlabeled Data for Target-Oriented Opinion Words Extraction (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to extract opinion words from sentences are limited due to the expensive annotation process.
Approach: They propose to exploit massive unlabeled data to reduce distribution shift risk . they propose to use two filters specifically for TOWE to filter noisy data . results indicate superiority of MGCR over current state-of-the-art methods .
Outcome: The proposed method reduces the risk of distribution shifts by increasing the exposure of the model to varying distribution shift.
What Factors Influence LLMs’ Judgments? A Case Study on Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies indicate that Large Language Models perform at a level comparable to humans with advantages of speed and cost-effectiveness in different fields.
Approach: They propose to introduce four unexplored factors and a new dimension of question difficulty to provide a more comprehensive understanding of LLMs’ judgments across varying question intricacies.
Outcome: The proposed dimensions of question difficulty and answer quantity provide valuable insights into optimizing LLMs’ performance as judges.
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities.
Approach: They propose a visual-textual embedding framework that integrates textual and visual features for robust document representation.
Outcome: The proposed visual-textual embedding framework surpasses existing methods while preserving semantic fidelity.
EnsLM: Ensemble Language Model for Data Diversity by Semantic Clustering (2021.acl-long)

Copied to clipboard

Challenge: Existing studies have shown that data diversity affects the performance of LMs if we train a single LM over the entire dataset.
Approach: They propose an autoencoding topic model with a mixture prior to perform clustering for the data.
Outcome: The proposed model can learn knowledge from different samples while extracting cluster-specific features.
DataArc-SynData-Toolkit: A Unified Closed-Loop Framework for Multi-Path, Multimodal, and Multilingual Data Synthesis (2026.acl-demo)

Copied to clipboard

Challenge: Existing synthetic data tools are limited by convoluted workflows, fragmented data standards, and limited scalability across modalities.
Approach: They develop an open-source framework that aims to reduce the technical barrier to synthetic data generation and subsequent model training.
Outcome: The proposed framework achieves an optimal balance between generation efficiency and data quality.
LMDX: Language Model-based Document Information Extraction and Localization (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models have revolutionized Natural Language Processing but their application in extracting information from visually rich documents has not been successful.
Approach: They propose a language model-based document information extraction and localization methodology to reframe the document information extract task for a LLM.
Outcome: The proposed method enables extraction of singular, repeated, and hierarchical entities with and without training data.
Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Continual Learning (CL) for Large Language Models faces a fundamental Stability-Plasticity Dilemma . Rank-Blindness enforces a single rank constraint across diverse tasks, leading to catastrophic forgetting of earlier tasks and underfitting on complex new ones.
Approach: They propose a rank-spectrum-based rehearsal-free framework that explicitly disentangles knowledge into two orthogonal subspaces.
Outcome: The proposed framework achieves a superior stability-plasticity balance compared to single-rank baselines.
SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to training GUI agents on dynamic tasks are based on SFT or Behavior Cloning.
Approach: They propose a framework that integrates global trajectory insights directly into offline learning . they reconstruct diverse rollout candidates from static data and detect first failure point .
Outcome: The proposed framework improves long-horizon task completion rates and robustness compared to baselines.
OpenRLHF: A Ray-based Easy-to-use, Scalable and High-performance RLHF Framework (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing RLHF frameworks face inference bottlenecks and complexity barriers restricting their accessibility for newcomers.
Approach: They propose an open-source RLHF framework that can be used to train large language models.
Outcome: The proposed framework achieves superior training efficiency with speedups ranging from 1.22 to 1.68 across different model sizes compared to state-of-the-art frameworks, while requiring significantly fewer lines of code for implementation.
LONG2RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall (2024.findings-emnlp)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) is a promising approach to address limitations of fixed knowledge in large language models.
Approach: They propose a benchmark and a metric to assess LLMs' ability to generate long-form responses that exploit retrieved information.
Outcome: The proposed benchmarks lack a comprehensive evaluation method to assess LLMs' ability to generate long-form responses that effectively exploits retrieved information.
Activation-Guided Local Editing for Jailbreaking Attacks (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for jailbreaking Large Language Models (LLMs) are limited and produce incoherent or unreadable inputs.
Approach: They propose a two-stage framework that performs a one-shot, scenario-based generation of context and rephrases the original malicious query to obscure its harmful intent.
Outcome: The proposed framework achieves state-of-the-art Attack Success Rate, with gains of up to 37.74% over the strongest baseline, and excellent transferability to black-box and large-scale models.
The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches focus on information processing and strategy selection, overlooking the significance of persuasive communication in social deduction games.
Approach: They propose a reinforcement learning framework that trains agents to optimize influential utterances for persuasive impact by formalizing turn-based dialogue as a Stackelberg competition .
Outcome: The proposed framework outperforms baselines across four social deduction benchmarks and shows that it is effective in persuasive communication.
Generative Zero-Shot Prompt Learning for Cross-Domain Slot Filling with Inverse Prompting (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to slot filling only learn surface mapping of slot types between D S and D T and get poor generalization capability or robustness.
Approach: They propose a generative zero-shot prompt learning framework for cross-domain slot filling which improves generalization and robustness than previous work.
Outcome: The proposed framework improves generalization and robustness on unseen slots and an efficient prompt tuning strategy boosts performance.
ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for visually rich document understanding lack layout-centered knowledge . experimental results show that ERNIE-Layout improves layout awareness .
Approach: They propose a document pre-training solution with layout knowledge enhancement in the whole workflow to learn better representations that combine the features from text, layout, and image.
Outcome: The proposed model outperforms existing models on key downstream tasks.
Semi-Supervised Bilingual Lexicon Induction with Two-way Interaction (2020.emnlp-main)

Copied to clipboard

Challenge: Existing semisupervised methods do not fully utilize the knowledge hidden in annotated and nonannotated data, which hinders further improvement of their performance.
Approach: They propose a semi-supervised BLI framework to encourage interaction between supervised signal and unsupervised alignment.
Outcome: The proposed framework can incorporate any supervised and unsupervised BLI methods based on optimal transport and bi-directional lexicon update.
A Web-scale system for scientific knowledge exploration (P18-4)

Copied to clipboard

Challenge: a system that organizes scientific knowledge into a hierarchical concept structure is needed to enable efficient exploration of Web-scale knowledge.
Approach: They propose a system that organizes scientific knowledge into a hierarchical concept structure . system allows researchers to identify hundreds of thousands of scientific concepts . it also allows researchers tagging scientific publications into millions of concepts based on text and graph structure based model .
Outcome: The proposed system builds the most comprehensive cross-domain scientific concept ontology published to date, with more than 200 thousand concepts and over one million relationships.
When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling (2026.findings-acl)

Copied to clipboard

Challenge: Existing research implicitly assumes that longer thinking leads to better results . a recent study suggests that test-time compute scaling is more effective than model scaling .
Approach: They challenge the assumption that longer thinking yields better results . they show that models exhibit overthinking and marginal returns diminish at higher budgets .
Outcome: The proposed framework reduces computation significantly while maintaining comparable accuracy.
LAiW: A Chinese Legal Large Language Models Benchmark (2025.coling-main)

Copied to clipboard

Challenge: Xie et al., 2023) show that large language models (LLMs) can generate legal text, but lack the legal syllogism . legal experts are cautious about their practical application due to the opaque nature of the LLMs.
Approach: They propose a Chinese legal LLM benchmark structured around the legal syllogism . they evaluate LLMs across three levels of capability, each reflecting a more complex stage of legal .
Outcome: The proposed benchmark identifies that LLMs lack the legal syllogism, which hinders trust and understanding from legal experts.
A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners (2024.emnlp-main)

Copied to clipboard

Challenge: a new hypothesis-testing framework is developed to assess whether large language models possess genuine reasoning abilities or primarily depend on token bias.
Approach: They propose a framework to assess whether large language models have genuine reasoning abilities or primarily depend on token bias.
Outcome: The proposed framework outlines a list of hypotheses where token biases are readily identifiable . the results suggest that most LLMs still struggle with logical reasoning .
Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show superior cross-model transferability . Prompt injection attacks are among the most critical threats .
Approach: They propose an activations-guided prompt injection attack framework to address the impracticality of existing white-box/gray-box methods and the poor transferability of black-box approaches.
Outcome: The proposed framework achieves 49.6% success rate and 34.6% improvement over human-crafted prompts on five mainstream LLMs.
Speak No Evil, Just Prompt: Low-resource Multilingual Toxic Speech Detection with Audio Language Model (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for toxic speech detection rely on high-resource languages and lack acoustic cues.
Approach: They propose a prompt-based adaptation framework that performs end-to-end toxicity detection without ASR.
Outcome: The proposed framework achieves a micro-averaged ROC-AUC of 98.07% on polySpeechTox . it is based on a frozen audio language model and can perform end-to-end toxicity detection without ASR .
CompKBQA: Component-wise Task Decomposition for Knowledge Base Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge base question answering methods struggle with complex queries.
Approach: They propose a framework that optimizes the process of fine-tuning a LLM for generating logical forms by enabling it to learn relevant sub-tasks like skeleton generation, topic entity generation, and relevant relations generation.
Outcome: The proposed framework achieves state-of-the-art on two benchmark KBQA datasets, WebQSP and CWQ.
UKP-SQuARE v3: A Platform for Multi-Agent QA Research (2023.acl-demo)

Copied to clipboard

Challenge: Current approaches to QA models are multi-dataset models, but combining expert agents can yield large performance gains over multi-agent models.
Approach: They extend an online platform for QA research to support three families of multi-agent systems: agent selection, early-fusion of agents, and late-fusion.
Outcome: The proposed model can be compared with multi-dataset models and achieve high inference speed and performance.
ENPAR:Enhancing Entity and Entity Pair Representations for Joint Entity Relation Extraction (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for joint entity relation extraction use multitask learning frameworks, but annotations for additional tasks are hard to obtain.
Approach: They propose a pre-training method to improve the joint extraction performance with just extra entity annotations.
Outcome: The proposed method outperforms existing methods on ACE05, SciERC, and NYT and outperformed BERT on other tasks.
Imitation Learning for Non-Autoregressive Neural Machine Translation (P19-1)

Copied to clipboard

Challenge: Existing non-autoregressive translation models lack parallel decoding, which is a bottleneck for NMT decoding.
Approach: They propose a framework for non-autoregressive machine translation that emulates the autoregressive model by sampling sentence length in parallel.
Outcome: The proposed model achieves 31.85 BLEU on WMT16 RoEn and 30.68 BLUE on IWSLT16 EnDe on the IWSLD16, WMT14 and WMT15 datasets.
IntelliCockpitBench: A Comprehensive Benchmark to Evaluate VLMs for Intelligent Cockpit (2025.findings-acl)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a key task in vehicular systems.
Approach: They propose a benchmark that encompasses diverse automotive scenarios . they use images from front, side, and rear cameras, various road types, weather conditions, and interior views .
Outcome: The proposed benchmark includes images from front, side, and rear cameras, various road types, weather conditions, and interior views.
CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for instruction tuning rely on expensive human-annotated seed data or powerful external teacher models.
Approach: They propose a framework that achieves fully seed-free instruction tuning by employing a dual self-training loop where two models are bootstrapped solely from raw, unlabeled text.
Outcome: The proposed framework outperforms seed-driven back-translation baselines and achieves comparable performance to strongly supervised methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations