Papers with LLMs

300 papers
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias (2024.naacl-short)

Copied to clipboard

Challenge: Position bias is a tendency of a model unfairly prioritizing information from certain parts of the input text over others, leading to undesirable behavior.
Approach: They propose to measure position bias in large language models for zero-shot summarization tasks by measuring position bias.
Outcome: The proposed model performance and position biases lead to new insights and discussion on zero-shot summarization tasks.
Towards LLM-powered Attentive Listener: A Pragmatic Approach through Quantity Self-Repair (2025.acl-short)

Copied to clipboard

Challenge: Quantity Maxims dictates that human speakers aim for optimal quantity of information during conversation.
Approach: They propose to use heuristic path-finding to enable decoder-only LLMs to travel among multiple "Q-alternatives" and search for optimal quantity in coordination with a conversation goal.
Outcome: The proposed techniques are based on heuristic path-finding and can be used to construct human-like, user-centered conversation agents.
Towards Automated Error Discovery: A Study in Conversational AI (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work shows that LLMs require information about the nature of an error or hints about its occurrence for accurate detection.
Approach: They propose an encoder-based approach to detect and define errors in conversational AI.
Outcome: The proposed framework outperforms baselines across multiple error-annotated dialogue datasets and shows strong generalization to unknown intent detection.
PropGenie: A Multi-Agent Conversational Framework for Real Estate Assistance (2026.eacl-demo)

Copied to clipboard

Challenge: PropGenie is a multi-agent framework based on large language models (LLMs) it provides comprehensive real estate assistance in real-world scenarios .
Approach: They propose a multi-agent framework based on large language models to deliver comprehensive real estate assistance in real-world scenarios.
Outcome: The proposed framework outperforms a general-purpose LLM and a domain-specific chatbot in real-world scenarios.
CHATREPORT: Democratizing Sustainability Disclosure Analysis through LLM-based Tools (2023.emnlp-demo)

Copied to clipboard

Challenge: a lack of transparency in sustainability reporting is a key challenge due to the sheer volume and complexity of sustainability reports . only a few entities worldwide have the resources to analyze these reports at scale . a novel LLM-based system to automate the analysis of corporate sustainability reports is needed .
Approach: They propose a novel LLM-based system to automate the analysis of corporate sustainability reports.
Outcome: The proposed system automates the analysis of corporate sustainability reports.
How LLMs React to Industrial Spatio-Temporal Data? Assessing Hallucination with a Novel Traffic Incident Benchmark Dataset (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have a high potential to digitize and enhance the health & public services industry.
Approach: They propose to use a cross-lingual benchmark dataset to assess the robustness of state-of-the-art LLMs in the spatio vs temporal domain for traffic incident classification.
Outcome: The proposed model performs well in the spatio-temporal domain and in the non-English context.
Big AI is Accelerating the Metacrisis: What Can We Do? (2026.acl-short)

Copied to clipboard

Challenge: LLM engineering is at the core of the problem of ecological, meaning, and language crises . big AI is fueling global crises and creating wealth and power for a handful of individuals and corporations while causing existential harm to life on earth.
Approach: et al., 2025, p162ff) argue that big AI is escalating global crises and creating a metacrisis.
Outcome: the field of natural language processing is at the core of the problem . it is being leveraged to create unprecedented wealth and power for a handful of individuals and corporations while causing existential harm to life on earth.
Guardrails and Security for LLMs: Safe, Secure and Controllable Steering of LLM Applications (2025.acl-tutorials)

Copied to clipboard

Challenge: Pretrained generative models provide novel ways for users to interact with computers.
Approach: This tutorial provides an overview of key guardrail mechanisms developed for LLMs along with evaluation methodologies and a detailed security assessment protocol.
Outcome: This tutorial provides an overview of key guardrail mechanisms developed for LLMs, along with evaluation methodologies and a detailed security assessment protocol.
NarrativePlay: Interactive Narrative Understanding (2024.eacl-demo)

Copied to clipboard

Challenge: Existing systems for interactive agents focus on specific capabilities in predetermined scenarios.
Approach: They propose a novel system that allows users to role-play a fictional character and interact with other characters in narratives in an immersive environment.
Outcome: The proposed system generates human-like responses guided by personality traits extracted from narratives.
Building Knowledge-Guided Lexica to Model Cultural Variation (2024.naacl-long)

Copied to clipboard

Challenge: Cultural variation exists between nations, but also within regions . Historically, it has been difficult to computationally model cultural variation due to a lack of training data and scalability constraints.
Approach: They propose a method to measure cultural variation using a knowledge-guided lexical model using geolocated tweets.
Outcome: The proposed method could help us better understand the way people communicate and build more culturally-aware NLP systems.
Tree-of-Evolution: Tree-Structured Instruction Evolution for Code Generation in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Data synthesis is a key research area in large language models (LLMs).
Approach: They propose a framework that models code instruction synthesis process with a tree structure and optimization-driven evolution to alleviate constraints of unidirectional synthesis and randomness-driven generation.
Outcome: The proposed framework outperforms open-weight code LLMs on five widely-used benchmarks.
ReasonGraph: Visualization of Reasoning Methods and Extended Inference Paths (2025.acl-demo)

Copied to clipboard

Challenge: Large Language Models (LLMs) reasoning processes are complex and lack of organized visualization tools creates barriers to understanding, evaluation, and improvement.
Approach: They propose a web-based platform for visualizing and analyzing LLM reasoning processes.
Outcome: The proposed platform shows high parsing reliability, efficient processing, and excellent usability across various downstream applications.
Learning to Verify Summary Facts with Fine-Grained LLM Feedback (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have significantly enhanced the text summarization performance, but hallucination issues still occur in summaries.
Approach: They propose a large-scale dataset containing fine-grained factual feedback on summaries that can be fine tuned by using Large Language Models (LLMs) they employ 10 distinct LLMs for diverse summary generation and Llama-3-70B-Instruct for feedback.
Outcome: The proposed model outperforms models trained on smaller human-annotated datasets while maintaining high performance.
Leveraging Product Catalog Patterns for Multilingual E-commerce Product Attribute Prediction (2025.emnlp-industry)

Copied to clipboard

Challenge: E-commerce stores increasingly use Large Language Models to improve catalog data quality . a critical challenge is accurately predicting missing structured attribute values .
Approach: They propose a retrieval-augmented system that leverages existing product catalog entries to guide LLM predictions for missing attributes.
Outcome: The proposed system improves catalog data quality by 34% and accuracy by 0.8% . the proposed model can predict missing attributes in multilingual product catalogs .
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)

Copied to clipboard

Challenge: Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models .
Approach: They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets.
Outcome: The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets.
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency (2024.emnlp-industry)

Copied to clipboard

Challenge: Large language models (LLMs) are widely used in commercial applications . low latency is crucial due to system latency, query concurrency, and computational resources constraints.
Approach: They propose a system that can be resource-efficiently served by addressing bottlenecks beyond LLM inference . they propose 4.3 speed up over vLLM and 1.5 higher throughput .
Outcome: The proposed system outperforms state-of-the-arts with 1.5 higher throughput . it achieves 4.3 speed up with 64 concurrent requests on Mixtral 8x7B .
EduPulse: A Practical LLM-Enhanced Opinion Mining System for Vietnamese Student Feedback in Educational Platforms (2026.eacl-industry)

Copied to clipboard

Challenge: EduPulse is a system designed specifically to analyze student feedback in Vietnamese.
Approach: They propose a system that analyzes student feedback in Vietnamese to improve opinion mining.
Outcome: The proposed system performs four opinion analysis tasks in Vietnamese . it is scalable and maintainable, and it is cost-effective, the authors show .
Enhancing Future Link Prediction in Quantum Computing Semantic Networks through LLM-Initiated Node Features (2025.coling-industry)

Copied to clipboard

Challenge: Quantum computing is rapidly evolving in both physics and computer science due to its potential to solve complex quantum physics problems and accelerate computational processes.
Approach: They propose to initialize node features using LLMs to enhance node representations for link prediction tasks in graph neural networks.
Outcome: The proposed method compared to traditional node embedding techniques on a quantum computing semantic network and demonstrated efficacy compared with other methods.
A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets (2023.findings-acl)

Copied to clipboard

Challenge: Currently, the evaluation of large language models (LLMs) such as ChatGPT in academic datasets is difficult due to the difficulty of evaluating the generative outputs produced by this model against the ground truth.
Approach: They evaluate ChatGPT across 140 tasks and analyze 255K responses it generates in academic datasets.
Outcome: The proposed model performs well on 140 tasks and generates 255K responses in these datasets.
PEARL: Prompting Large Language Models to Plan and Execute Actions Over Long Documents (2024.eacl-long)

Copied to clipboard

Challenge: Using chain-of-thought prompting, large language models perform better on complex reasoning tasks.
Approach: They propose a prompting framework that decomposes a question into a sequence of actions and executes them over the document to obtain the answer.
Outcome: The proposed framework outperforms zero-shot and chain-of-thought prompting on a QuALITY dataset . it proposes a plan based on actions mined from a training set and executes it step by step .
Resonance RoPE: Improving Context Length Generalization of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated their potential across a wide spectrum of natural language processing tasks.
Approach: They propose a novel approach to narrow the generalization gap in TSTL scenarios by refining the interpolation of RoPE features for OOD positions.
Outcome: The proposed approach improves performance without additional online computational costs on train-short-test-long scenarios.
FlashBack: Efficient Retrieval-Augmented Language Modeling for Fast Inference (2025.findings-acl)

Copied to clipboard

Challenge: Retrieval-Augmented Language Modeling (RALM) is a popular approach for large language models.
Approach: They propose a modular RALM that integrates large language models with documents from an external corpus to improve inference efficiency.
Outcome: The proposed method improves inference efficiency with appending context pattern while maintaining decent performance after fine-tuning by Low-Rank Adaption.
Corpus-Steered Query Expansion with Large Language Models (2024.eacl-short)

Copied to clipboard

Challenge: Recent studies show query expansions generate hypothetical documents that answer queries as expansions.
Approach: They propose a corpus-steered query expansion to promote incorporation of knowledge embedded within the corpus.
Outcome: et al. analyzed corpus-based Query Expansion (CSQE) using LLMs to generate hypothetical documents that answer the query.
Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are trained on vast corpora that contain substantial knowledge but their outputs often contain confidently stated inaccuracies.
Approach: They propose to encode truthfulness as a distinct linear feature, termed the "truth direction", which can classify truthfulness reliably.
Outcome: The proposed model can generalize to logical transformations, question-answering tasks, in-context learning, and external knowledge sources.
From Generating Answers to Building Explanations: Integrating Multi-Round RAG and Causal Modeling for Scientific QA (2025.naacl-industry)

Copied to clipboard

Challenge: Application of Large Language Models to complex causal question answering can be stymied by their opacity and propensity for hallucination.
Approach: They propose a causal QA approach that combines iterative RAG with a formal model of causation.
Outcome: The proposed approach is implemented into a Collaborative Research Assistant (Cora) and evaluated in the life sciences domain.
Sample Design Engineering: An Empirical Study on Designing Better Fine-Tuning Samples for Information Extraction with LLMs (2024.emnlp-industry)

Copied to clipboard

Challenge: Prompt Engineering (PE) is renowned for improving IE performance through prompt modifications, but the realm of sample design for downstream fine-tuning remains unexplored.
Approach: They propose a methodical approach to enhancing LLMs’ post-tuning performance by refining input, output, and reasoning designs.
Outcome: The proposed approach outperforms heuristic design strategies on three complex IE tasks with four additional LLMs.
EvoAgentX: An Automated Framework for Evolving Agentic Workflows (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing MAS frameworks often require manual workflow configuration and lack native support for dynamic evolution and performance optimization.
Approach: They propose an open-source platform that automates generation, execution, and evolutionary optimization of multi-agent workflows.
Outcome: The proposed platform automates generation, execution, and evolutionary optimization of multi-agent workflows.
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator (2025.coling-industry)

Copied to clipboard

Challenge: Established metrics such as ROUGE and BERTScore have a relatively low correlation with human judgments and fail to capture nuanced errors.
Approach: They propose a framework that uses a three-step assessment of individual error types, multi-agent discussion for decision refinement, and feedback-based self-training to refine error definition understanding and alignment with human judgment.
Outcome: The proposed framework achieves high correlation with human judgment and a consistent rating and adaptability to custom error guidelines.
AIDA-SEAT: Towards Reliable AI Doctor Assistant via State-Evaluation-Action Tree Enhanced LLMs in Online Hospital (2026.acl-industry)

Copied to clipboard

Challenge: Existing systems rely on large language models or retrieval-augmented generation (RAG) but these methods lack the explicit logical pathways essential for multi-step reasoning.
Approach: They propose an AIDA-SEAT framework to provide reliable clinical decision-making support by transforming and modifying medical documents and doctors' state-evaluation-action trees.
Outcome: The proposed framework achieves 1.01% higher than current state-of-the-art (SOTA) baselines across five departments, including common RAG-based methods.
IndicJR: A Judge-Free Benchmark of Jailbreak Robustness in South Asian Languages (2026.eacl-industry)

Copied to clipboard

Challenge: Indic Jailbreak Robustness (IJR) is a judge-free benchmark for adversarial safety across 12 languages.
Approach: They propose a judge-free benchmark for adversarial safety across 12 languages . they find contracts inflate refusals but do not stop jailbreaks .
Outcome: The proposed benchmarks cover 45,216 prompts in JSON and Free tracks.
Improving Few-Shot Cross-Domain Named Entity Recognition by Instruction Tuning a Word-Embedding based Retrieval Augmented Large Language Model (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing approaches to named entity recognition are domain specific and require a domain specific architecture.
Approach: They propose a retrieval augmented large language model for Named Entity Recognition . the model uses word-embedding over sentence-level embedding to fine tune .
Outcome: The proposed model outperforms existing models on the CrossNER dataset.
Are U a Joke Master? Pun Generation via Multi-Stage Curriculum Learning towards a Humor LLM (2024.findings-acl)

Copied to clipboard

Challenge: Existing research has demonstrated that the ability of large language models (LLMs) to generate humorous sentences is limited to producing 25 unique jokes.
Approach: They propose a multi-stage curriculum preference learning framework to optimize both pun structure preferences and humor preferences by a Chinese Pun dataset.
Outcome: The proposed method significantly outperforms baseline models on Chinese and English benchmark datasets.
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs (2024.findings-naacl)

Copied to clipboard

Challenge: Visual language models (VLMs) are achieving increasingly strong performance on multimodal tasks.
Approach: They propose to transfer reasoning capabilities from large-language models to VLMs by constructing a 20x larger dataset and a larger dataset to improve general reasoning capabilities.
Outcome: The proposed model outperforms larger models without an upstream OCR system while keeping inference time constant.
Self-Criticism: Aligning Large Language Models with their Understanding of Helpfulness, Honesty, and Harmlessness (2023.emnlp-industry)

Copied to clipboard

Challenge: Recent studies have shown that large language models are useful, honest, harmless (HHH) however, RLHF requires high hardware resources and human efforts.
Approach: They propose a framework that allows LLMs to align themselves with HHH . they use IF and reinforcement learning from human feedback to fine-tune their models .
Outcome: The proposed framework achieves similar performance to RLHF and human-generated models with a minimal alignment tax.
TimeRes: A Turkish Benchmark For Evaluating Temporal Understanding of Large Language Models (2026.eacl-srw)

Copied to clipboard

Challenge: Existing benchmarks focus on English and underexplore how linguistic structure contributes to temporal meaning.
Approach: They propose a Turkish benchmark to evaluate temporal understanding of Large Language Models (LLMs) their benchmark examines Reichenbach’s temporal points and reported speech through date arithmetic .
Outcome: The proposed model fails to resolve reported speech and fails to generalize across word order variations.
Neural Topic Modeling with Large Language Models in the Loop (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated promising capabilities in topic discovery, but their direct application to topic modeling suffers from issues such as incomplete topic coverage, misalignment of topics, and inefficiency.
Approach: They propose a novel LLM-in-the-loop framework that integrates Large Language Models with Neural Topic Models (NTMs) global topics and document representations are learned through the NTM, while an LLM refines these topics using an Optimal Transport (OT)-based alignment objective.
Outcome: The proposed framework improves topic interpretability while preserving the efficiency of existing NTMs.
From Sentences to Proof Trees: Leveraging Language Models for Structured Reasoning (2026.eacl-srw)

Copied to clipboard

Challenge: Multi-hop reasoning requires a chain of facts to reflect the reasoning behind the answer.
Approach: They propose an inference-guided prompting approach that performs well in natural language questions . they propose a neuro-symbolic approach to reasoning using large language models .
Outcome: The proposed model outperforms all prompting strategies and fine-tunes LLMs trained specifically for proof generation.
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection (2025.findings-naacl)

Copied to clipboard

Challenge: Recent research suggests using Large Language Models (LLMs) to automate the annotation process, reducing these costs while maintaining data quality.
Approach: They propose to use Large Language Models to automate annotation process and train classifiers on large datasets.
Outcome: The proposed model outperforms all of the annotator LLMs on two media bias benchmark datasets (BABE and BASIL) while maintaining data quality.
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in LLMs have demonstrated superior performance in a variety of reasoning tasks (Liu et al., 2023b; Chan e t al, 2024; Qin eetal., 2023) However, to truly achieve conscious processing, the integration of System II reasoning ability is essential.
Approach: They propose a three-step process for reasoning with distributional changes, termed as a metaphysical resoning, and propose 'MARS' task to assess LLMs' reasoning abilities.
Outcome: The proposed task is based on a three-step discriminative process and is compared with a standard model with 20 LLMs of varying sizes and methods.
NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debugging (2025.emnlp-main)

Copied to clipboard

Challenge: Early debugging efforts focused on code-level analysis, which often fails when addressing complex programming errors.
Approach: They propose a framework that employs natural language as an intermediate representation to improve code debugging by debuggating at a natural language level.
Outcome: The proposed framework outperforms traditional debugging methods and enables a broader modification space through direct refinement guided by execution feedback.
Agentic Economic Modeling (2026.acl-industry)

Copied to clipboard

Challenge: AEM is a framework that aligns synthetic LLM choices with small-sample human evidence for reliable econometric inference.
Approach: They introduce a framework that aligns synthetic LLM choices with small-sample human evidence for reliable econometric inference.
Outcome: The proposed framework improves RCT efficiency and establishes a foundation method for LLM-based counterfactual generation.
To Forget or Not? Towards Practical Knowledge Unlearning for Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing unlearning paradigms are mired in vague forgetting boundaries, erasing knowledge indiscriminately.
Approach: They propose a benchmark to evaluate if unlearning erases essential knowledge . they propose 'knowUnDo' which uses copyrighted content and privacy domains .
Outcome: The proposed method is superior to existing methods in both precise knowledge unlearning and general knowledge retaining of LLMs.
Divergent Thoughts toward One Goal: LLM-based Multi-Agent Collaboration System for Electronic Design Automation (2025.naacl-long)

Copied to clipboard

Challenge: Electronic design automation (EDA) is indispensable for the design of integrated circuits.
Approach: They propose a multi-agent collaboration system where multiple agents harbor divergent thoughts converge towards a common goal.
Outcome: The proposed system shows superior performance compared to single-agent systems.
Efficiently Aligned Cross-Lingual Transfer Learning for Conversational Tasks using Prompt-Tuning (2024.findings-eacl)

Copied to clipboard

Challenge: Cross-lingual transfer of language models trained on high-resource languages such as English has been limited due to the high cost of obtaining non-English conversational data.
Approach: They introduce a parallel and large-scale multilingual conversation dataset that is used for cross-lingual alignment pretraining by translating the English-only Schema-Guided Dialogue dataset into 105 other languages.
Outcome: The proposed model performs well on slot-filling and intent classification tasks, and is able to perform well in other languages.
ITERATE: Image-Text Enhancement, Retrieval, and Alignment for Transmodal Evolution with LLMs (2025.coling-main)

Copied to clipboard

Challenge: a new framework for visual annotation of text-based questions is needed to improve performance . obtaining corresponding images through manual annotation often entails high costs .
Approach: They propose a framework that uses visual modality to enhance the performance of text-based questions.
Outcome: The proposed framework improves the alignment between text and images by using search engines or web scraping techniques.
Static Models, Dynamic World: A Unified Perspective on Temporal Perception in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are trained on static corpora but deployed in a dynamic world . a foundational tension remains between time and the ability to understand it .
Approach: They formalize temporal queries in an information-theoretic framework based on parametric reachability of temporal premises and answers.
Outcome: The proposed framework formalizes temporal queries in an information-theoretic framework based on parametric reachability of temporal premises and answers . the framework induces four temporal information regimes corresponding to internal reasoning, answer recency, premise anchoring, and genuine world indeterminacy .
Increasing Coverage and Precision of Textual Information in Multilingual Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to generate knowledge graphs are unable to handle non-English textual information.
Approach: They propose a task of automatic Knowledge Graph Completion to bridge the gap between English and non-English textual information.
Outcome: The proposed method bridges the gap between the quantity and quality of textual information between English and non-English languages.
Patentformer: A Novel Method to Automate the Generation of Patent Applications (2024.emnlp-industry)

Copied to clipboard

Challenge: Patentformer is a novel method for generating patent specification by fine-tuning the generative models with diverse sources of information, e.g., patent claims, drawing text, and brief descriptions of the drawings.
Approach: They propose a method for generating patent specification by fine-tuning generative models with diverse sources of information, e.g., patent claims, drawing text, and brief descriptions of the drawings.
Outcome: The proposed method generates patent specification in legal writing style and human-like quality may be better than the actual specification.
UNIWIZ: A Unified Large Language Model Orchestrated Wizard for Safe Knowledge Grounded Conversations (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have made significant progress in integrating safety and knowledge alignment, but excessive focus on safety alignment can lead to unintended hallucinations.
Approach: They propose a "safety-priming" method to generate synthetic safety data and overcome safety bottlenecks.
Outcome: The proposed framework generates synthetic safety data and overcomes safety bottlenecks.
When Facts Change: Temporal Knowledge Conflict Resolution in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are increasingly used in retrieval-augmented generation systems to reconcile knowledge conflicts between parametric memory and contextual inputs.
Approach: They propose to use mutability to resolve temporal misalignment in large language models to compare stable and recently updated facts from Wikidata to determine if mutable models can serve as a mediating signal in this process.
Outcome: The proposed model can produce reasoning for facts that actually changed but rarely for stable ones, whereas smaller models rarely detect conflict, while larger models detect it but fail to act on mutability judgments.
RoleCDE: Benchmarking and Mitigating Role–Alignment Trade-offs in Role-Playing Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for role-playing agents only evaluate surface-level fidelity and provide limited insight into decision making under role–alignment value conflicts.
Approach: They propose a benchmark to evaluate RPAs under role–alignment value conflicts . they use 8k diverse role profiles and 240k dilemma instances to evaluate role-aware decision making .
Outcome: The proposed benchmark covers 8k diverse role profiles and scenarios and nearly 240k dilemma instances across three difficulty levels and eight role categories.
CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are limited by their narrow language pairs and tasks, failing to adequately assess their code-mixing abilities.
Approach: They propose a benchmark to assess large language models' (LLMs) code-mixing abilities that covers eight tasks and 18 languages from seven language families.
Outcome: The proposed method combines word substitution with GPT-4 prompting to generate large-scale synthetic code-mixed texts.
C2LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown significant promise, yet their evaluation raises concerns regarding data contamination due to the lack of access to proprietary training data.
Approach: They propose a bilingual benchmark that offers a holistic evaluation and systematic contamination prevention.
Outcome: The proposed evaluations of 15 open-source and proprietary models show that they are reliable and free of data contamination.
When the Model Said ‘No Comment’, We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified (2026.eacl-long)

Copied to clipboard

Challenge: Existing work uses SFT and MoE to align Large Language Models, but these work face challenges in multi-objective settings.
Approach: They propose a framework that uses prompt-injected fine-tuning to extract axis-specific task features . it deploys a MoCaE module that calibrates expert routing using fractal and natural geometry .
Outcome: The proposed framework achieves significant gains on Alpaca, BeaverTails, TruthfulQA and TruthfulQ with +171.5% win rate and +110.1% truthfulness-informativeness.
Fine-Tuned LLMs are “Time Capsules” for Tracking Societal Bias Through Books (2025.naacl-long)

Copied to clipboard

Challenge: We develop a corpus comprising 593 fictional books across seven decades (1950-2019) to track bias evolution.
Approach: They develop a method to trace and quantify bias evolution using fine-tuned LLMs on fictional books across seven decades to track bias evolution.
Outcome: The proposed method traces and quantifies bias evolution in a corpus of 593 fictional books across seven decades.
Mixture of Diverse Size Experts (2024.emnlp-industry)

Copied to clipboard

Challenge: Recent large language models (LLMs) have shown superior performance in a variety of tasks due to the sub-linearly increasing computational costs.
Approach: They propose a new MoE architecture with designed layers where experts have different sizes to mitigate this defect.
Outcome: The proposed architecture surpasses existing MoEs by adaptively assigning the parameter budget to experts while maintaining the same total parameter size and number of experts.
Course-Correction: Safety Alignment Using Synthetic Preferences (2024.emnlp-industry)

Copied to clipboard

Challenge: Recent studies show that large language models generate harmful content, but the potential for generating harmful content is an escalating concern.
Approach: They propose to fine-tune LLMs with preference learning to emphasize the preference for timely course-correction by using an automated pipeline.
Outcome: The proposed model improves course-correction skills without affecting general performance and resists jailbreak attacks.
Investigating Agency of LLMs in Human-AI Collaboration Tasks (2024.eacl-long)

Copied to clipboard

Challenge: We examine how LLMs can be measured and managed for Agency . a model that manifests high Intentionality, Motivation, Self-Efficacy, and Self-Regulation is more likely to be perceived as strongly agentive.
Approach: They collect a dataset of 83 human-human collaborative interior design conversations containing 908 conversational snippets annotated for Agency features.
Outcome: The proposed models show that they manifest high Intentionality, Motivation, Self-Efficacy, and Self-Regulation, and are more likely to be perceived as agentive.
Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for code retrieval struggle to balance scalability and annotation quality.
Approach: They propose a method that integrates functions called within the repository and information on third-party APIs to enhance the annotation context.
Outcome: The proposed method improves the annotation context by incorporating functions called within the repository and information on third-party API functionalities.
CLASS: A Design Framework for Building Intelligent Tutoring Systems Based on Learning Science principles (2023.findings-emnlp)

Copied to clipboard

Challenge: CLASS empowers ITS with two key capabilities: first, it equips it with essential problem-solving strategies, and second, it facilitates natural language interactions, fostering engaging student-tutor conversations.
Approach: They propose a design framework called Conversational Learning with Analytical Step-by-Step Strategies (CLASS) that empowers ITS with two key capabilities: first, a carefully curated dataset and second, facilitating natural language interactions.
Outcome: The proposed framework empowers ITS with two key capabilities: first, it equips it with essential problem-solving strategies, and second, it facilitates natural language interactions, fostering engaging student-tutor conversations.
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in activation quantization methods cause outliers in tokens, causing extra overhead and speedup . a method to quantize per-tensor activation is currently challenging due to the outlier activation outlier.
Approach: They propose a method to find a set of key-value cache which mitigates outliers in subsequent tokens when inserted as a prefix.
Outcome: The proposed method surpasses the established baseline of per-tensor activation quantization and can be seamlessly integrated with the recent activation quantitative method.
NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models have shown impressive performance on various benchmarks, yet their ability to engage in deliberate reasoning remains questionable.
Approach: They propose to penalize quick, intuitive "System 1" thinking by combining linguistic isolation with resistance to intuitive shortcuts to assess model's reasoning abilities.
Outcome: The proposed model penalizes quick, intuitive “System 1” thinking, isolating fundamental reasoning skills.
Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models are increasingly used in verbal creative tasks.
Approach: They propose a divergent association task that focuses on novelty, ignoring appropriateness, a core component of creativity.
Outcome: The proposed model scores are lower than baselines with no creative abilities, undermining its validity for model evaluation.
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for hallucination management fail to integrate both detection and mitigation without external knowledge sources.
Approach: They propose a black-box framework that leverages fine-grained cross-model consistency to detect and mitigate hallucinations in LLM outputs without external knowledge sources.
Outcome: The proposed framework improves hallucination detection scores by 6-39% on a FELM dataset . it achieves 9 percentage points improvement in answer accuracy on the GPQA-diamond dataset compared to existing approaches .
Safe-Unsafe Concept Separation Emerges from a Single Direction in Language Models Activation Space (2026.eacl-long)

Copied to clipboard

Challenge: Existing approaches to ensuring the safety of Large Language Models (LLMs) rely on invasive fine- tuning or external generation-based checks, which can be opaque and resource-inefficient.
Approach: They propose a mechanistic method that identifies the layer where safe and unsafe concepts are maximally separable within a pretrained representation space.
Outcome: The proposed method can be used across multiple domains, diverse tasks, and 16 non-English languages on encoder and decoder architectures.
Head-wise Shareable Attention for Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) suffer from huge number of parameters, which restricts their deployment on edge devices.
Approach: They propose two methods that share parameters across attention heads to reduce memory usage and reduce performance drop by using coarse-grained weight sharing rules.
Outcome: The proposed methods reuse pre-trained weights without retraining and then share, denoted as PostShare.
Rethinking Data Selection at Scale: Random Selection is Almost All You Need (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing data selection techniques are designed for small data pools, a study finds . filtering data by token length is an efficient method for improving results .
Approach: They use self-scoring methods that do not rely on external help to perform fine-tuning . they also find that filtering data by token length offers a stable and efficient method .
Outcome: The proposed methods outperform random selection on large datasets on large data pools.
\mathsf{Con Instruction}: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities (2025.acl-long)

Copied to clipboard

Challenge: Existing attacks communicate instruction through text, accompanied by a toxic image or audio . a novel gray-box attack method generates adversarial images or audio to convey harmful instructions to MLLMs .
Approach: They propose a gray-box attack method that generates adversarial images or audio to convey specific harmful instructions to MLLMs by following non-textual instruction.
Outcome: The proposed method achieves highest success rates on visual and audio-language models . larger models are more susceptible toCon Instruction, compared to their underlying models - the results will be released .
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time (2025.emnlp-main)

Copied to clipboard

Challenge: Existing LLMs struggle to reliably detect subtle reasoning errors in ASAS tasks.
Approach: They propose a dual-model framework with a dedicated Critic model trained for effective reflection that generates precise verbal feedback.
Outcome: The proposed framework outperforms existing ASAS benchmarks and provides valuable insights into the performance of the proposed framework.
Revisiting Automated Prompting: Are We Actually Doing Better? (2023.acl-short)

Copied to clipboard

Challenge: Recent work demonstrates that Large Language Models are great few-shot learners, and prompting significantly increases their performance on a range of downstream tasks.
Approach: They revisit techniques for automated prompting on six different downstream tasks and a larger range of K-shot learning settings.
Outcome: The proposed approach outperforms manual prompting on six different downstream tasks and a larger range of K-shot learning settings.
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective (2025.acl-long)

Copied to clipboard

Challenge: if rewards are imperfect, they can adversely affect the alignment of large language models (LLMs).
Approach: They propose a bias-agnostic method to address the issue of reward unfairness from a resource allocation perspective without specifically designing for each type of bias . they apply methods Fairness Regularization and Fairness Coefficient to achieve fairness in rewards.
Outcome: The proposed method achieves fairness in rewards while minimizing biases . it can be applied to verification and reinforcement learning scenarios .
CODERL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models excel at code generation by learning from vast code corpora, but a fundamental semantic gap remains between training on textual patterns and the goal of functional correctness . reinforcement learning with verifiable rewards (RLVR) approaches are inefficient for establishing a well-aligned connection between the textual representation of code and its execution semantics.
Approach: They propose a novel approach that integrates execution semantics alignment into the RLVR training pipeline for code generation.
Outcome: The proposed model outperforms baseline training and RLVR and shows strong applicability across RL and LLMs.
Direct Evaluation of Chain-of-Thought in Multi-hop Reasoning with Knowledge Graphs (2024.findings-acl)

Copied to clipboard

Challenge: Prior research on evaluating large language models focused on answer accuracy, neglecting the correctness of the generated CoT.
Approach: They propose a discriminative and generative CoT evaluation paradigm to assess LLMs’ knowledge of reasoning and the accuracy of the generated CoT.
Outcome: The proposed evaluation paradigm assesses LLMs’ knowledge of reasoning and the accuracy of the generated CoT.
AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages (2026.eacl-long)

Copied to clipboard

Challenge: Text embeddings are an essential building component of several NLP tasks.
Approach: They propose a regional expansion of MTEB covering 59 languages, 14 tasks, and 38 datasets, including six newly added datasets.
Outcome: The proposed model outperforms baselines and mE5 in hate speech detection, intent detection, and emotion classification tasks.
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description (2025.findings-naacl)

Copied to clipboard

Challenge: Existing 3D facial emotion modeling models are constrained by limited emotion classes and insufficient datasets.
Approach: They propose a 3D facial emotion modeling dataset that spans a wide spectrum of human emotions . they use large language models to generate a diverse array of textual descriptions .
Outcome: Emo3D is an extensive dataset that spans human emotions with images and 3D blendshapes.
DS2-Instruct: Domain-Specific Data Synthesis for Large Language Models Instruction Tuning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing data synthesis methods focus on general-purpose tasks and fail to capture domain-specific terminology and reasoning patterns.
Approach: They propose a framework that generates domain-specific instruction datasets without human supervision by pairing task-informed keywords with different cognitive levels from Bloom’s Taxonomy.
Outcome: The proposed framework generates domain-specific instruction datasets without human supervision and achieves significant improvements over existing methods.
Towards Better Evaluation for Generated Patent Claims (2025.acl-long)

Copied to clipboard

Challenge: Existing studies highlight inconsistencies between automated evaluation metrics and human expert assessments for patent claims.
Approach: They propose a multi-dimensional evaluation method specifically designed for patent claims that incorporates features annotated by patent experts.
Outcome: The proposed method achieves highest correlation with human expert evaluations across all assessment criteria across all tested metrics.
Conformity in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Conformity is a form of social influence that affects the way people respond to information.
Approach: They adapt psychological experiments to examine the extent of conformity in large language models.
Outcome: The proposed interventions mitigate conformity by reducing the naturalness of majority tones and reducing instruction-tuned models.
Task Matters: Knowledge Requirements Shape LLM Responses to Context–Memory Conflict (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that large language models favor parametric knowledge under conflict, but this setting assumes that tasks should always rely on the provided passage.
Approach: They propose a model-agnostic diagnostic framework that holds underlying knowledge constant while injecting controlled conflicts across tasks with varying knowledge requirements.
Outcome: Evaluating representative open-source LLMs, the proposed framework holds underlying knowledge constant while injecting controlled conflicts across tasks with varying knowledge requirements.
GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods of acceleration require fine-tuning of considerably large models, such as Llama-7B, posing a challenge for average users.
Approach: They propose a Guidance-based Knowledge Transfer framework that leverages a larger LLM as a 'teacher' and a smaller 'student' model to finalize responses.
Outcome: The proposed framework achieves a maximum accuracy improvement of 14.18%, along with a 10.72 times speed-up on GSM8K and an accuracy improvement 14.00% along with 7.73 times speed up in CSQA.
Improving Zero-shot Reader by Reducing Distractions from Irrelevant Documents in Open-Domain Question Answering (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) enable zero-shot approaches in open domain question answering (ODQA), yet with limited advancements as the reader is compared to the retriever.
Approach: They propose to use a distraction-aware answer selection framework to mitigate the impact of irrelevant documents in the retrieved set and the overconfidence of the generated answers to enhance the performance of zero-shot readers.
Outcome: The proposed approach handles distraction across diverse scenarios, enhancing the performance of zero-shot readers.
AdaEdit: Advancing Continuous Knowledge Editing For Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing knowledge editing methods that can efficiently update knowledge in LLMs are limited due to budget constraints.
Approach: They propose a method that can enhance the performance of edited LLMs in large-size continuous editing regimes.
Outcome: Extensive empirical evaluations on multiple LLMs show that the proposed method outperforms existing methods without compromising the general abilities of these models.
Word Surprisal Correlates with Sentential Contradiction in LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Existing models are primarily optimized for task-specific performance, lacking well-defined objectives or linguistic grounding.
Approach: They propose a token-to-word decoding algorithm that extends theoretically grounded probability estimation to open-vocabulary settings.
Outcome: The proposed algorithm can localize sentence-level inconsistency at the word level, establishing a quantitative link between lexical uncertainty and sentential semantics.
EcoSafeRAG: Efficient Security through Context Analysis in Retrieval-Augmented Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing defense methods rely on internal knowledge of the model, which conflicts with the design concept of Retrieval-Augmented Generation (RAG).
Approach: EcoSafeRAG uses sentence-level processing and bait-guided context diversity detection to identify malicious content .
Outcome: EcoSafeRAG uses sentence-level processing and bait-guided context diversity detection to identify malicious content.
CONSCENDI: A Contrastive and Scenario-Guided Distillation Approach to Guardrail Models for Virtual Assistants (2024.naacl-long)

Copied to clipboard

Challenge: A major challenge in deploying LLM-based virtual conversational assistants in real world settings is ensuring they operate within what is admissible for the task.
Approach: They propose to use large language models (LLMs) to generate training data with two key LLM components: scenario-augmented generation and contrastive training examples.
Outcome: The proposed model improves over baselines in multiple dialogue domains.
PII-Bench: Evaluating Query-Aware Privacy Protection Systems (2026.acl-long)

Copied to clipboard

Challenge: Existing models do not detect PII in user prompts, despite their convenience . current models show significant limitations in determining PI I query relevance .
Approach: They propose a query-unrelated PII masking strategy and propose PIi-Bench . they propose 'quick-and-easy' PI I masking with a user query and context description .
Outcome: The proposed model performs well in basic PII detection, but shows significant limitations in query relevance.
Boundary-Aware LLM Augmentation for Low-Resource Event Argument Extraction (2026.eacl-long)

Copied to clipboard

Challenge: Event argument extraction (EAE) is a crucial task in information extraction but its performance heavily depends on expensive annotated data.
Approach: They investigate argument replacement, adjunction rewriting, their combination, and annotation generation using four LLM-based augmentation strategies.
Outcome: The proposed methods improve performance over boundary-agnostic methods and provide detailed analysis of quality from multiple perspectives.
Improving In-Context Learning with Prediction Feedback for Sentiment Analysis (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved promising results in sentiment analysis through the in-context learning paradigm.
Approach: They propose a framework that incorporates prior predictions and feedback to improve sentiment understanding by incorporating prior feedback and leveraging a feedback-driven prompt.
Outcome: The proposed framework improves on nine sentiment analysis datasets with an average improvement of 5.95% over conventional methods.
Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to detect large language models (LLMs) use binary or ternary classifications, which can only distinguish pure human/LLM text or collaborative text at best.
Approach: They propose a fine-grained method that characterizes distinct signatures of creator and editor by using Rhetorical Structure Theory to construct a logic graph for creator's foundation and extracting Elementary Discourse Unit (EDU)-level features for the editor's style.
Outcome: The proposed method outperforms 12 baselines in identifying fine-grained types with low false alarms, offering a policy-aligned solution for LLM regulation.
M-QALM: A Benchmark to Assess Clinical Reading Comprehension and Knowledge Recall in Large Language Models via Question Answering (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on adapting large language models to perform a variety of tasks in high-stakes domains such as healthcare lack understanding of the extent and contributing factors that allow them to recall relevant knowledge and combine it with presented information.
Approach: They propose to use multiple choice and abstractive question answering to investigate the extent and contributing factors that allow LLMs to recall relevant knowledge and combine it with presented information in the clinical and biomedical domain.
Outcome: The proposed models perform better on 22 datasets in three generalist and three specialist biomedical sub-domains, and show that they can generalise to unseen sub- domains.
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: End-to-end speech-to speech (S2S) dialogue systems face key challenges in incorporating external knowledge into their models.
Approach: They propose a framework that directly retrieves relevant textual knowledge from speech queries.
Outcome: The proposed framework improves the performance of end-to-end speech-tospeech dialogue systems while achieving higher retrieval efficiency.
When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives (2024.emnlp-main)

Copied to clipboard

Challenge: Using sports data, an LLM can analyze sports narratives to infer points from actions, identify related entities, attribute points accurately to players and teams, and draw conclusions.
Approach: They propose a method to synthesize NBA basketball game narratives using real NBA basketball data and propose 'SportsGen' they find that most models fail to accurately aggregate basketball scores due to frequent scoring patterns and open-source models suffer from significant score hallucinations.
Outcome: The proposed method can evaluate LLMs’ reasoning capabilities under complex scenarios with varying narrative lengths and density of information.
K-order Ranking Preference Optimization for Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing list-wise methods focus on optimizing list ranking consistency for LLMs to improve ranking abilities.
Approach: They propose to extend the Plackett-Luce model to accommodate top-K ranking by extending the DPO’s Plact-Lucer model to dynamically determine appropriate K for different samples.
Outcome: The proposed model can be extended to accommodate top-K ranking and improve training efficiency.
Enabling Autoregressive Models to Fill In Masked Tokens (2026.findings-eacl)

Copied to clipboard

Challenge: Autoregressive (AR) and masked language modeling (MLM) models are incapable of mucked infilling, which is the ability to predict mangled tokens between past and future context.
Approach: They propose a method that leverages the strengths of autoregressive and masked language modeling to achieve state-of-the-art mucked infilling performance.
Outcome: The proposed approach outperforms existing methods on masked infilling tasks.
WavLLM: Towards Robust and Adaptive Speech Large Language Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have expanded their scope to encompass multimodal functions.
Approach: They propose a robust and adaptive speech large language model with dual encoders . they validate the model on universal speech benchmarks and apply it to specialized speech-question-answer datasets based on a CoT approach .
Outcome: The proposed model achieves state-of-the-art performance across a range of speech tasks on the same model size.
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances have extended DPO to multimodal scenarios, achieving strong performance.
Approach: They propose to use a sentence-level preference optimization technique to optimize individual sentences for more precise preference optimization without additional models or parameters.
Outcome: Experiments show that Adaptive Sentence-level Preference Optimization significantly improves the alignment of multimodal models.
How Do Humans Write Code? Large Models Do It the Same Way Too (2024.emnlp-main)

Copied to clipboard

Challenge: Program-of-Thought (PoT) replaces natural language-based Chain-ofThough (CoT) but introduces more reasoning errors, such as incorrect formulas or flawed logic, compared to CoT.
Approach: They propose a method that integrates CoT and Program-of-Thought to achieve more accurate reasoning and reinforcement learning.
Outcome: The proposed method achieves an average improvement of 6.5% on the Llama-Base model and 4.3% on the Mistral-Bass model across 8 mathematical calculation datasets.
Aligning Complex Knowledge Graph Question Answering as Knowledge-Aware Constrained Code Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing frameworks that generate LF using Large Language Models (LLMs) in a few-shot setting are limited due to little exposure to the LF during pre-training.
Approach: They propose a framework that aligns the LF generation as code generation that incorporates LF-specific constraints.
Outcome: The proposed framework surpasses all few-shot baselines on KQA Pro by 21%.
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that large language models can cause harmful, human-like biases against various demographics.
Approach: They propose a causal formulation for bias measurement in generative language models based on a list of desiderata for designing robust bias benchmarks and a bias-measuring procedure to investigate occupational gender bias.
Outcome: The proposed framework is generalizable and can be extended to include other datasets.
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models’ Posteriors (2025.naacl-long)

Copied to clipboard

Challenge: In-context Learning (ICL) is the primary method for performing natural language tasks with Large Language Models.
Approach: They examine whether aggregation is a confounding factor in the modeling of subjective tasks . they find it is possible for minority annotators to better align with LLMs .
Outcome: The proposed method is based on aggregation of annotations in a dataset with appropriate priors.
From Single to Multi: How LLMs Hallucinate in Multi-Document Summarization (2025.findings-naacl)

Copied to clipboard

Challenge: a recent study investigated hallucinations in multi-document summarization tasks . but, it is unclear how challenges arising from handling multiple documents affect outputs .
Approach: They investigate how hallucinations manifest in large language models when summarizing topic-specific information from a set of documents.
Outcome: The proposed benchmarks show that the models generate more hallucinations than baselines . the results highlight the need for more effective approaches to mitigate hallucinosity in MDS .
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean (2024.lrec-main)

Copied to clipboard

Challenge: Existing benchmark datasets for Korean cultural and linguistic knowledge are derived from the English counterparts through translation, so they overlook cultural contexts.
Approach: They propose to use Korean cultural and linguistic intelligence to assess Korean model performance by providing fine-grained annotations of cultural and cultural knowledge.
Outcome: The proposed dataset includes 1,995 QA pairs and is based on 1,992 Korean exams and textbooks.
Large Language Models are good multi-lingual learners : When LLMs meet cross-lingual prompts (2025.coling-main)

Copied to clipboard

Challenge: Experimental results show that Large Language Models can generate rule-based data in long contexts without following all specified rules.
Approach: They propose a novel prompting strategy Multi-Lingual Prompt which automatically translates the error-prone rule that an LLM struggles to follow into another language, thus drawing greater attention to it.
Outcome: The proposed framework outperforms state-of-the-art prompting methods on public datasets across various tasks, with a specific case study in text-to-MIP instances.
LLatrieval: LLM-Verified Retrieval for Verifiable Generation (2024.naacl-long)

Copied to clipboard

Challenge: Large language models struggle with factual errors and often produce non-factual and fabricated content.
Approach: They propose to use large language models to generate text with supporting documents to enable the user to flexibly verify the answer.
Outcome: Experiments on ALCE show that LLatrieval significantly outperforms extensive baselines and achieves state-of-the-art results.
UTMath: A Benchmark for Math Evaluation with Unit Test (2025.findings-emnlp)

Copied to clipboard

Challenge: Prevailing benchmarks for mathematical reasoning include MATH and AIME . predicated on single-instantiation problems with fixed numbers, these models leave generalization on isomorphic problem variants untested.
Approach: They propose a mathematical reasoning benchmark that quantifies solution accuracy and solution space generality.
Outcome: The proposed model solves 1,053 problems spanning 9 mathematical domains . the best-performing model solved only 32.57% of the problems .
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds (2025.naacl-long)

Copied to clipboard

Challenge: Evaluating role-playing capabilities in large language models is challenging due to complex dynamics involved in role-playering.
Approach: They propose a simulation sandbox that generates situational fine-grained character behavior trajectories to enhance LLM performance.
Outcome: The proposed model generates situational fine-grained character behavior trajectories to enhance performance.
Progra: Progress-Aware Reinforcement Learning for Multi-Turn Function Calling (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for multi-turn function calling are limited by redundancy and lack explicit integration of progress awareness into training.
Approach: They propose a framework that explicitly integrates progress awareness into LLM training for multi-turn function calling.
Outcome: Empirical results show that Progra outperforms existing methods on two public benchmarks.
QUDeval: The Evaluation of Questions Under Discussion Discourse Parsing (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics poorly approximate parser quality, says a new study . questions under discussion is a linguistic framework that views discourse as asking questions and answering them .
Approach: They propose a framework for automatic evaluation of QUD parsing . they use a dataset of fine-grained evaluation of 2,190 QUD questions .
Outcome: The proposed framework shows that satisfying constraints of QUD is still challenging for modern LLMs.
Towards Modern Topic Models: A Survey of Taxonomies and Paradigm Shifts from Algorithm-Centric to LLM-Centered Topic Analysis (2026.findings-acl)

Copied to clipboard

Challenge: Topic modeling (TM) is a classic unsupervised learning task in the field of natural language processing.
Approach: They propose a new taxonomy that emphasizes the role of LLMs and the design of end-to-end workflows.
Outcome: The proposed taxonomy emphasizes the role of LLMs and the design of end-to-end workflows.
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in reasoning with large language models have popularized Long Chain-of-Thought (LCoT) a framework that converts sequential LCoTs into hierarchical tree structures enables deeper structural analysis of LLM reasoning.
Approach: They propose a framework that converts sequential LCoTs into hierarchical tree structures and enables deeper structural analysis of LLM reasoning.
Outcome: The proposed framework can be used to analyze LLM reasoning in a variety of tasks and models.
TRANSIENTTABLES: Evaluating LLMs’ Reasoning on Temporally Evolving Semi-structured Tables (2025.naacl-long)

Copied to clipboard

Challenge: a recent study shows that large language models are limited in their ability to reason over time due to static datasets.
Approach: They present a dataset that includes 3,971 questions derived from over 14,000 tables . they introduce a template-based question-generation pipeline that harnesses LLMs to refine questions .
Outcome: The proposed model improves on the TRANSIENTTABLES dataset . it demonstrates that the model can reason over time, even when it is not static .
The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding? (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to self-improvement rely on external supervision signals in the form of seed data and/or assistance from third-party models.
Approach: They propose a framework for generating high-quality synthetic question-answer data in a fully autonomous manner.
Outcome: The proposed framework generates high-quality synthetic question-answer data in a fully autonomous manner.
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has shown that self-citing large language models (LLMs) fail to faithfully reflect their context usage throughout the generation process.
Approach: They propose a plug-and-play approach using model internals for faithful answer attribution in RAG applications that detects context-sensitive answer tokens and pairs them with retrieved documents contributing to their prediction.
Outcome: The proposed approach achieves citation quality and efficiency comparable to self-citation while allowing for a finer-grained control of attribution parameters.
Multimodal Conversation Structure Understanding (2026.eacl-long)

Copied to clipboard

Challenge: a new set of tasks is being developed to parse the structure of conversation . female characters are 1.2 times more likely to be cast as an addressee or side-participant .
Approach: They propose a set of tasks and release an annotated dataset for multimodal conversation structure understanding.
Outcome: The proposed model outperforms the baseline model, but performance drops when character identities are anonymized.
Distilling Instruction-following Abilities of Large Language Models with Task-aware Curriculum Planning (2024.findings-emnlp)

Copied to clipboard

Challenge: Instruction tuning aims to align large language models (LLMs) with open-domain instructions and human-preferred responses.
Approach: They propose a multi-round distillation framework that uses an oracle LLM to select instructions that are difficult for a student LLM.
Outcome: The proposed framework outperforms large language models and user-tuned models on several widely recognized benchmarks and multiple student LLMs.
Exploring Large Language Models for Multi-Modal Out-of-Distribution Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: Out-of-distribution (OOD) detection is essential for reliable and trustworthy machine learning.
Approach: They propose to apply world knowledge to enhance OOD detection performance through selective generation from large language models (LLMs) they propose to extract visual objects from each image to fully capitalize on the aforementioned world knowledge.
Outcome: The proposed method outperforms the state-of-the-art on visual OOD detection on in-distribution (ID) samples.
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) focus on replicating human cognition in specific contexts, overlooking the inherently dynamic nature of cognition.
Approach: They propose a task to assess cognitive dynamics of large language models (LLMs) they introduce a benchmark and two evaluation metrics to validate the benchmark and evaluate it through participant surveys.
Outcome: The proposed task overcomes the limitations of existing methods and is available for download.
QUITO-X: A New Perspective on Context Compression from the Information Bottleneck Theory (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for compressing context by removing redundant tokens are inconsistent with the objective of retaining the most important tokens when conditioning on a given query.
Approach: They propose a method that uses information bottleneck theory to compress context . they propose to remove redundant tokens using metrics such as self-information or perplexity .
Outcome: The proposed method achieves a 25% increase in compression rate compared to the state-of-the-art .
SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing safety mechanisms for large language models (LLMs) are inadequate to fully leverage their internal cognitive processes.
Approach: They propose a framework that regulates unsafe outputs by utilizing the prober-based internal state monitor that actively detects harmful intentions.
Outcome: The proposed framework reduces harmful outputs by approximately 80% while maintaining strong utility.
Detection and Measurement of Syntactic Templates in Generated Text (2024.emnlp-main)

Copied to clipboard

Challenge: Existing diversity evaluation focuses primarily on word-level features.
Approach: They propose a method for evaluating diversity over syntactic features to characterize general repetition in large language models.
Outcome: The proposed method shows that models produce templated text in downstream tasks at a higher rate than what is found in human-reference texts.
From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment (2026.findings-acl)

Copied to clipboard

Challenge: a framework for sentence-level interpretability of rubric-based scoring is proposed . aaron e. smith: automated scoring models provide little insight into why scores are produced .
Approach: They propose a framework for sentence-level interpretability of rubric-based scoring that combines Shapley-value attributions with rationales generated by large language models.
Outcome: The proposed framework compares fine-tuned pretrained language models with large language models . it shows that fine- tuned models outperform LLMs in prediction accuracy but exhibit label compression toward mid-scale scores .
A Survey of Large Language Models in Psychotherapy: Current Landscape and Future Directions (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) can handle extensive context and multi-turn reasoning.
Approach: They propose a taxonomy dividing psychotherapy into stages of assessment, diagnosis, and treatment to examine LLM advancements and challenges.
Outcome: The proposed taxonomy reveals imbalances in current research, such as a focus on common disorders, linguistic biases, fragmented methods, and limited theoretical integration.
DA-Pred: Performance Prediction for Text Summarization under Domain-Shift and Instruct-Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) often don’t perform as expected under Domain Shift or after Instruct-tuning.
Approach: They propose a method that uses the known performance in high-resource domains and fine-tuning settings to predict performance in low-resourced domains or base models.
Outcome: The proposed method can help researchers decide if resources should be allocated for data labeling and LLM Instruct-tuning.
On Evaluating LLMs’ Capabilities as Functional Approximators: A Bayesian Evaluation Framework (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized the way we can formulate tasks in text-in-text-out format.
Approach: They propose a new evaluation framework to comprehensively assess LLMs’ function modeling abilities by adopting a Bayesian perspective of function modeling.
Outcome: The proposed evaluation framework enables LLMs to excel in utilizing prior knowledge to develop a strong understanding of the underlying function.
Sneaking Syntax into Transformer Language Models with Tree Regularization (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for incorporating syntactic inductive biases into transformers are limited . we introduce auxiliary loss function that converts bracketing decisions into differentiable orthogonality constraints on vector hidden states.
Approach: They propose to introduce syntactic inductive biases into transformer circuits through a structured regularizer.
Outcome: The proposed approach could unlock more robust and data-efficient learning in transformer language models . it integrates seamlessly with the standard LM objective, requiring no architectural changes.
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies show that Large Language Models are biased towards a Western and Anglo-centric worldview.
Approach: They propose to extend the Octopus test to measure "cultural awareness" they argue that cultural awareness is needed for AI systems to be useful across cultures .
Outcome: The proposed method argues that cultural awareness is not cultural knowledge, but meta-cultural competence . the proposed method is based on the octopus test, which shows it is impossible to learn meaning from real-world concepts without knowing intent and meaning .
Iterative Knowledge Graph Refinement and Integration for Medical Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Existing graph-based RAG methods heuristically retrieve and refine question-relevant subgraphs, potentially introducing redundant and noisy factual information that is difficult for LLMs to process.
Approach: They propose to integrate knowledge graphs (KGs) through retrieval-augmented generation methods to improve LLM reasoning by incorporating external trustworthy knowledge resources.
Outcome: The proposed framework achieves state-of-the-art against baseline competitors on three medical QA benchmark datasets.
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos (2025.findings-acl)

Copied to clipboard

Challenge: Large Reasoning Models (LRMs) have advanced beyond traditional Large Language Models, yet they pose heightened safety risks.
Approach: They propose a first jailbreak attack targeting Large Reasoning Models . they exploit a Chaos Machine component to transform attack prompts with diverse one-to-one mappings based on the reasoning chain .
Outcome: The proposed attack exploits the unique vulnerabilities of LRMs by integrating a Chaos Machine. success rates of the mousetrap attack are as high as 96%, 86% and 98% respectively.
StepKE: Stepwise Knowledge Editing for Multi-Hop Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing knowledge editing methods overlook interplay with pre-existing knowledge, leading to inconsistent edit propagation.
Approach: stepKE integrates edited and existing knowledge for coherent multi-hop reasoning . stepKE decomposes multi-step questions into sequential single-hop sub-questions .
Outcome: Experiments show that StepKE generates more accurate and consistent responses than baselines.
Identifying Semantic Induction Heads to Understand In-Context Learning (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable performance, but lack of transparency in their inference logic raises concerns about their trustworthiness.
Approach: They conduct a detailed analysis of the operations of attention heads to understand their in-context learning of LLMs.
Outcome: The proposed analysis of attention heads reveals that they increase the output logits of object tokens and recall objects . the proposed model is a novel approach to understand the in-context learning of large language models.
Detecting Conceptual Abstraction in LLMs (2024.lrec-main)

Copied to clipboard

Challenge: a novel approach to detecting noun abstraction within a large language model is proposed . a first step towards the explainability of conceptual abstraction in LLMs is shown .
Approach: They propose a method to detect noun abstraction within a large language model . they instantiate taxonomic relationships and analyze attention matrices produced by BERT .
Outcome: The proposed approach can detect hypernymy in a large language model . the results are a first step towards the explainability of conceptual abstraction in LLMs .
Text Embedding Inversion Security for Multilingual Language Models (2024.acl-long)

Copied to clipboard

Challenge: storing sensitive information as embeddings is susceptible to security breaches, as text can be reconstructed from embeddables . study explores multilingual inversion attacks using a masking defense .
Approach: They propose a simple masking defense that can be used to decode embedded text . they define the problem of black-box multilingual and crosslingual inversion attacks .
Outcome: The proposed defense is effective for both monolingual and multilingual models.
EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees (2024.emnlp-main)

Copied to clipboard

Challenge: Modern Large Language Models (LLMs) are expensive and time-consuming.
Approach: They propose a new technique of context-aware dynamic draft tree into drafting modeling.
Outcome: The proposed method achieves speedup ratios of up to **5x**, which is 1.3x that of EAGLE.
Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models (2025.findings-naacl)

Copied to clipboard

Challenge: Using large vision-language models to understand cultural contexts is a critical area of research.
Approach: They conduct a thorough evaluation of multimodal models at different scales, focusing on their alignment with cultural values.
Outcome: The proposed models show that they exhibit sensitivity to cultural values but their performance is highly context-dependent.
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination? (2024.naacl-long)

Copied to clipboard

Challenge: Existing large language models (LLMs) suffer from hallucinations and unfaithful reasoning due to keyword/entity biases.
Approach: They propose a new probing method and benchmark to quantify this phenomenon by using a keyword/entity biases-based probing technique called EUREQA.
Outcome: The proposed method achieves 62% accuracy on multi-hop and complex QA benchmarks.
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond (2025.naacl-long)

Copied to clipboard

Challenge: Afrispeech-Dialog is a benchmark dataset of 50 simulated medical and non-medical African-accented English conversations . a 10%+ performance degradation is found in ASR systems on long-form, accented speech .
Approach: They propose to use a dataset to evaluate automatic speech recognition systems on African-accented conversations.
Outcome: The proposed dataset compares state-of-the-art speech recognition systems on accented conversations with native accents and shows a 10%+ performance degradation.
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved remarkable success in natural language tasks, yet understanding their reasoning processes remains a significant challenge.
Approach: They propose a dataset that includes 24204 instances where each instance interprets the LLM’s reasoning behavior using knowledge graphs and graph attention networks (GAT).
Outcome: The proposed explanation framework reduces hallucinations and improves grounded explanation generation in large language models.
Semi-Supervised Reward Modeling via Iterative Self-Training (2024.findings-emnlp)

Copied to clipboard

Challenge: Reward models capture values and preferences of humans and are used in Reinforcement Learning with Human Feedback (RLHF) Traditionally, training large language models relies on extensive human-annotated preference data, which poses significant challenges in terms of scalability and cost.
Approach: They propose a method that enhances RM training using unlabeled data.
Outcome: The proposed approach improves reward models without incurring additional labeling costs on unlabeled datasets.
Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect (2025.emnlp-main)

Copied to clipboard

Challenge: Prior work has focused largely on binary grammatical contrasts, but how do they encode their syntactic knowledge internally?
Approach: They propose to use a multidimensional hierarchical grammar phenomenon to identify distinct, orthogonal directions in residual space to demonstrate causal control over both grammatical features.
Outcome: The proposed model can encode tense and aspect in human-like ways, but effective steering during generation is sensitive to multiple factors and requires manual tuning or automated optimization.
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs (2026.acl-long)

Copied to clipboard

Challenge: High-quality post-training data is the primary engine driving LLM capabilities . datasets are often treated as isolated artifacts, overlooking their true developmental context .
Approach: They propose a framework to reconstruct the evolutionary graph of dataset development using data lineage.
Outcome: The proposed framework characterizes domain-specific structural patterns in Math-oriented datasets and general-domain corpora.
ANAH: Analytical Annotation of Hallucinations in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: a comprehensive and fine-grained measurement of the hallucination is crucial for LLMs' wide applications.
Approach: They propose a dataset that offers ANalytical Annotation of Hallucinations in Large Language Models.
Outcome: The proposed dataset can be used to train and evaluate hallucination annotators.
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models excel at few-shot learning but their direct application in real-world scenarios is often hindered by their high computational cost.
Approach: They propose a framework that uses Large Language Models for data generation and scoring to improve encoder model performance.
Outcome: The proposed approach improves accuracy from 28.9% to 39.3% on a few-shot MCQA task .
ARM: Alignment with Residual Energy-Based Model (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) acquire a wide range of abilities and abilities, but their behavior does not align with human preferences.
Approach: They propose to minimize a forward Kullback–Leibler divergence from a target policy to a parameteric policy instead of a reverse KL as in RLHF methods.
Outcome: The proposed method can learn an aligned policy by minimizing a forward Kullback–Leibler divergence from a target policy to a parameteric policy instead of a reverse KL as in RLHF methods.
Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown impressive language capabilities, but most of them have very unbalanced performance across different languages.
Approach: They propose to use question translation data to enhance LLMs' multilingual capabilities by using mechanistic interpretability methods.
Outcome: The proposed method improves multilingual alignment even with unannotated answers in English and a wide range of languages even with instruction-tuned LLMs.
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals (2024.acl-long)

Copied to clipboard

Challenge: Existing interpretability research focused on analyzing a single mechanism . et al., 2023) focused on finding how models copy or recall factual knowledge .
Approach: They propose a competition of mechanisms that focuses on the interplay of multiple mechanisms instead of individual mechanisms . they uncover how and where the competition of mechanism happens within LLMs using logit inspection and attention modification methods.
Outcome: The proposed model is based on two interpretability methods, logit inspection and attention modification.
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: In the recent past, a popular way of evaluating natural language understanding was to consider a model’s ability to perform natural language inference (NLI) tasks.
Approach: They focus on five different NLI benchmarks across six models of different scales and examine how their accuracies develop during training.
Outcome: The softmax distributions of models align with human label distributions in cases where statements are ambiguous or vague.
VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has focused on pushing weight-only quantization to extremely low-bit due to numerical representation limitations.
Approach: They propose a vector-based quantization approach that pushes LLMs to extremely low-bit . they propose scalar-based weight quantization that reduces memory requirements and optimizes storage costs .
Outcome: The proposed method reduces model quantization perplexity by 0.01-0.34 on LLaMA-2, 0.38-0.68 on mistral-7B, 4.41-7.34, on llaMA-3 on QA tasks on average.
Measuring memorization in language models via probabilistic extraction (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are susceptible to memorizing training data, raising concerns about the potential extraction of sensitive information at generation time.
Approach: They propose a method that splits training example into prefix and suffix, prompts LLM with suffix and deems it extractable if it generates the suffix using greedy sampling.
Outcome: The proposed method is unreliable because it does not account for non-determinism in more realistic sampling schemes.
TOP-Training: Target-Oriented Pretraining for Medical Extractive Question Answering (2025.coling-main)

Copied to clipboard

Challenge: e-health records underscore the growing significance of information extraction (IE) from these datasets.
Approach: They propose a target-oriented pre-training paradigm for extractive question-answering in the medical domain . TOP-Training moves one step further than popular domain-oriented fine-tuning .
Outcome: The proposed method improves on the Medical-EQA benchmarks.
When to Speak, When to Abstain: Contrastive Decoding with Abstention (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks by leveraging pre-trained (parametric) and external (contextual) knowledge.
Approach: They propose a training-free decoding method that allows LLMs to generate responses when relevant knowledge is available and to abstain otherwise.
Outcome: The proposed method can generate responses when relevant knowledge is available and abstain otherwise.
Hallucination Diversity-Aware Active Learning for Text Summarization (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for alleviating hallucinations require costly human annotations . Existing approaches focus on a specific type of hallucinism, which limits their effectiveness .
Approach: They propose a method to detect hallucinations from errors in semantic frame, discourse and content verifiability in LLM summarization using HAllucination Diversity-Aware Sampling.
Outcome: The proposed framework reduces the need for costly human annotations to correct hallucinations in LLM outputs.
On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMs (2025.acl-long)

Copied to clipboard

Challenge: Evidence-enhanced detectors are able to detect malicious social text, but they are prone to evidence pollution.
Approach: They propose three defense strategies to mitigate evidence pollution by large language models by machine-generated text detection and a mixture of experts.
Outcome: The proposed defense strategies could mitigate evidence pollution, but they faced limitations for practical employment.
Scaling Laws for Code: Every Programming Language Matters (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on language-agnostic settings, neglecting the inherently multilingual nature of modern software development.
Approach: They propose a proportion-dependent scaling law that prioritizes high-utility languages . they propose PLs to have varying effects during pre-training that affect model performance .
Outcome: The proposed scaling law is based on 1000+ experiments across multiple languages and models.
Can Large Language Models Win the International Mathematical Games? (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated strong mathematical reasoning abilities, even in visual contexts.
Approach: They propose a benchmark of 2,183 high-quality mathematical problems in an open-ended format that enables a structured evaluation of LLMs’ mathematical and logical reasoning abilities.
Outcome: The new benchmark spans seven age groups and a skill-based taxonomy and enables a structured evaluation of LLMs’ mathematical and logical reasoning abilities.
CONTOR: Benchmarking Strategies for Completing Ontologies with Plausible Missing Rules (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations focus on distinguishing held-out ontologies from randomly corrupted ones, which often makes the task unrealistically easy.
Approach: They propose to use the common description logic syntax for encoding ontology rules to test their effectiveness on manually annotated hard negatives.
Outcome: The proposed models are compared with existing models and have been evaluated on different ontologies.
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry (2025.findings-emnlp)

Copied to clipboard

Challenge: Detecting AI-generated poetry is difficult due to distinctive characteristics of modern Chinese poetry.
Approach: They propose a benchmark for detecting AI-generated modern Chinese poetry . they use a high-quality dataset and systematic performance assessments .
Outcome: The proposed benchmark is based on a high-quality dataset of 800 poems written by six professional poets and 41,600 poems generated by four mainstream LLMs.
PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing retrieval-based methods to pre-select tools are often misaligned with tool-calling LLMs due to separate training processes.
Approach: They propose a method to fine-tune retrievers to find useful tools by using a frozen LLM.
Outcome: The proposed method fine-tunes retrievers to find useful tools using a frozen LLM . it improves tool selection accuracy and can be generalized to new queries and tools .
Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals? (2024.findings-acl)

Copied to clipboard

Challenge: Recent work on the evaluation of large language models (LLMs) has shown unprecedented performance on diverse language generation tasks.
Approach: They investigate the controllability of large language models on scientific summarization tasks by controlling stylistic and content coverage factors.
Outcome: The proposed model outperforms humans on the MuP review generation task in terms of similarity to reference summaries and human preferences.
DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions (2024.emnlp-main)

Copied to clipboard

Challenge: a new variational approach to distractors in multiple-choice questions is needed . high-quality distractors are crucial to the assessment and pedagogical value of MCQs . a variational method that learns the error behind distractors is more effective .
Approach: They propose a variational approach that learns an interpretable representation of errors behind distractors in math MCQs.
Outcome: The proposed method outperforms state-of-the-art approaches on distractors in math MCQs.
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles (2025.findings-emnlp)

Copied to clipboard

Challenge: Figure captions are crucial for helping readers understand and remember a figure’s key message.
Approach: They propose a dataset for personalized figure caption generation with multimodal figure profiles that provide inputs and profiles for each figure .
Outcome: The proposed dataset provides inputs and profiles for personalized figure caption generation with multimodal figure profiles.
SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) demonstrate remarkable performance across diverse tasks, yet their effectiveness often depends on costly commercial APIs or cloud services.
Approach: They propose a dual-mode compatible approach that fine-tunes models through shortest-response preference optimization and a confidence-aware rejection mechanism.
Outcome: The proposed approach reduces redundant outputs and response times while reducing computational costs by over 50% and cascade latency by over 80%.
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models (2025.acl-long)

Copied to clipboard

Challenge: composition of pre-training datasets for large language models remains undisclosed . current methods for evaluating data quality are limited by single-dimensional evaluation or redundancy-focused strategies.
Approach: They propose a multi-dimensional data selection method that integrates dimensions with existing quality metrics through learned optimal weightings.
Outcome: The proposed method doubles convergence speed for 1.3B model models and improves downstream task performance by 3.23%.
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive capabilities but their application in open-ended, knowledge-intensive, complex reasoning scenarios is limited.
Approach: They propose a framework that integrates risk assessment of intermediate reasoning states with dynamic retrieval-augmented generation within a Monte Carlo tree search paradigm.
Outcome: The proposed framework outperforms the state-of-the-art KAR methods by up to 23.10% and the latest RAG-equipped large reasoning models by upto 25.37%.
Summarization of Opinionated Political Documents with Varied Perspectives (2025.coling-main)

Copied to clipboard

Challenge: Political ideologies can lead people to develop misperceptions of groups with opposing opinions, such as the 2024 US presidential election, French legislative election, or the Brexit referendum.
Approach: They propose a dataset and task for independently summarizing political perspectives in a set of opinionated news articles.
Outcome: The proposed dataset and task evaluates models of varying sizes and architectures on a set of opinionated news articles.
A Survey of LLM-based Agents in Medicine: How far are we from Baymax? (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming healthcare through their ability to understand and assist with medical tasks.
Approach: They analyze system profiles, clinical planning, medical reasoning frameworks, and external capacity enhancement.
Outcome: The findings highlight the future directions in medical reasoning, physical system integration, and training simulations.
Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models for test case generation have improved branch coverage via prompt-engineered mutations, limiting their effectiveness for discovering subtle bugs and security vulnerabilities.
Approach: They propose a program structure-aware LLM framework that integrates code property graphs and code semantics to condition test case generation on execution branches.
Outcome: Experiments on real-world projects show that GLMTest improves branch accuracy from 27.4% to 50.2% on TestGenEval benchmark compared with state-of-the-art LLMs, i.e., Claude-Sonnet-4.5 and GPT-4o-mini.
Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense Knowledge (2023.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been studied for their ability to store and utilize positive knowledge.
Approach: They propose to use a constrained keywords-to-sentence generation task and a Boolean question answering task to probe large language models on negative commonsense knowledge.
Outcome: The proposed tasks show that LLMs fail to generate valid sentences grounded in negative commonsense knowledge, yet they can correctly answer yes-or-no questions.
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) can solve reasoning and mathematical problems using the Chain-of-Thought technique, but require costly and long CoT data and fine-tuning.
Approach: They propose a method that uses Sparse Autoencoders to extract interpretable features from vanilla CoT and use them to steer the LLM's internal states.
Outcome: The proposed method uses Sparse Autoencoders (SAEs) to extract interpretable features from vanilla CoT and steer the LLM's internal states during generation.
TC–RAG: Turing–Complete RAG’s Case study on Medical LLM Systems (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to RAG neglect system state variables, resulting in poor performance and erroneous knowledge accumulation.
Approach: They propose a framework that incorporates a Turing Complete System to manage state variables and manage retrieval halting.
Outcome: The proposed framework improves on seven real-world healthcare datasets and shows that it is more accurate than existing methods.
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for detecting hallucinations in machine translation are limited for low-resource languages.
Approach: They evaluate sentence-level hallucination detection approaches using Large Language Models (LLMs) they find that the choice of model is essential for performance.
Outcome: The proposed models outperform the existing models in HRLs and LRLs on average by 0.16 MCC.
Evaluating Code-Switching Translation with Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown they can match or surpass finetuned models on many natural language processing tasks.
Approach: They propose to use in-context learning and pivot translation to improve code-switching translation.
Outcome: The proposed models show strong ability for cross-lingual understanding in a code-switching setting.
SumSurvey: An Abstractive Dataset of Scientific Survey Papers for Long Document Summarization (2024.findings-acl)

Copied to clipboard

Challenge: a growing need for long document summarization datasets with 16k input is causing problems.
Approach: They propose to use a dataset to analyze salient information in long document summarizations.
Outcome: The proposed dataset outperforms existing models and LLMs in the distribution form of salient information and the distribution of salinal information is an indicator of quality.
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification (2025.acl-long)

Copied to clipboard

Challenge: Chain-of-Thought prompting is a de facto method to elicit reasoning capabilities from large language models (LLMs).
Approach: They propose a step-aware formal verification framework Safe to address hallucinations in CoT prompting . they propose 'formal step' as a benchmark for step correctness theorem proving with 30,809 formal statements.
Outcome: The proposed framework shows significant performance improvement while offering interpretable and verifiable evidence.
Rhetorical Device-Aware Sarcasm Detection with Counterfactual Data Augmentation (2025.findings-acl)

Copied to clipboard

Challenge: Sarcasm is a complex form of sentiment expression widely used in human daily life.
Approach: They propose a device-aware sarcasm dataset with counterfactually augmented data to capture its complexity.
Outcome: The proposed dataset shows that it is more balanced than zero-shot models.
Detecting Machine-Generated Long-Form Content with Latent-Space Variables (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing zero-shot methods to distinguish machine-generated long-form texts from humans are vulnerable to domain shift including different decoding strategies, variations in prompts, and attacks.
Approach: They propose a method that incorporates abstract elements as key deciding factors by training a latent-space model on sequences of events or topics derived from human-written texts.
Outcome: The proposed method improves on baselines on three domains and significantly improves over existing methods.
Watch Out Your Industrial Copilots: Stealthy Backdoor Attack Against LLM-Based PLC Code Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are being used to generate PLC code from natural language.
Approach: They propose a stealthy backdoor attack framework targeting LLM-based PLC code generation . they incorporate six malicious logic injection patterns and a pipeline to refine stealthiness .
Outcome: The proposed framework achieves 82.92% success rate while remaining stealthy . it bypasses quality validation and is difficult to detect .
MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark (2023.emnlp-main)

Copied to clipboard

Challenge: MULTITuDE benchmarks lack authentic and machine-generated text in languages other than English . defining characteristic of new generation of LLMs is increased quality of text .
Approach: They propose a benchmarking dataset for multilingual machine-generated text detection that compares detectors with authentic and machine-generated texts in 11 languages.
Outcome: The proposed dataset compares detectors with zero-shot and fine-tuned detectors in 11 languages.
SCIURus: Shared Circuits for Interpretable Uncertainty Representations in Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for uncertainty quantification in large language models provide little insight into factors responsible for an uncertainty estimate, limiting their usefulness as practical tools for improving trustworthiness and understanding uncertainty reasoning.
Approach: They adapt causal tracing and zero-ablation techniques to study the effect of different circuits on LLM generation to identify whether factuality of generated responses and uncertainty originate in separate or shared circuits.
Outcome: The proposed methods use the well-established methods of causal tracing and zero-ablation to study the effect of different circuits on LLM generation.
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Quantization is a practical solution for deploying Large Language Models in resource-constrained environments.
Approach: They propose an outlier-safe pre-training approach that prevents outlier formation . they validate a 1.4B-parameter model on 1 trillion tokens with no outliers .
Outcome: The proposed model achieves a 35.7 average score on 1 trillion tokens with 2% training overhead.
Agentic Knowledgeable Self-awareness (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved considerable performance across various agentic planning tasks.
Approach: They propose a data-centric approach that applies agents with knowledgeable self-awareness like humans to a heuristic situation judgement criterion to mark special tokens on their self-explored trajectories for collecting training data.
Outcome: The proposed paradigm outperforms baseline models on various tasks with minimal external knowledge.
Learning When to Retrieve, What to Rewrite, and How to Respond in Conversational QA (2024.findings-emnlp)

Copied to clipboard

Challenge: Understanding users’ contextual search intent when generating responses is an understudied topic for conversational question answering (QA).
Approach: They propose a method that allows LLMs to decide when to retrieve in RAG settings given a conversational context.
Outcome: The proposed method improves on three conversational QA datasets and criticizes the quality of generated responses.
CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: a new benchmark is designed to evaluate LLMs on Chinese legal knowledge and its application in reasoning . general pre-training that ingests legal texts without specialized focus compromises reliability of LLM responses . achieving trustworthy legal reasoning in LLM requires a robust synergy of accurate knowledge retrieval and strong general reasoning capabilities.
Approach: They propose a benchmark specifically engineered to evaluate LLMs on Chinese legal knowledge and its application in reasoning.
Outcome: The proposed benchmark evaluates LLMs on Chinese legal knowledge and its application in reasoning.
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMs (2026.acl-long)

Copied to clipboard

Challenge: a lot of research aims to mitigate these problems by introducing specific computational solutions.
Approach: They examine how large language models engage in the interactive process of repair in multi-turn dialogues around solvable and unsolvable math questions.
Outcome: The models respond to user-initiated repair differently from one another . the models exhibit their own characteristic form of unreliability in the context of repair .
CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to align LLMs with recommendation tasks do not fully leverage their sequential information processing capabilities.
Approach: They propose a system that allows users to expand their vocabulary by assigning a unique ID to each item within the expanded vocabulary.
Outcome: The proposed system maximizes the sequence understanding abilities of large language models, significantly enhancing their performance on recommendation tasks.
Pragmatic Norms Are All You Need – Why The Symbol Grounding Problem Does Not Apply to LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: 'Symbol grounding problem' is a philosophical problem that arises when questionable theories of meaning are presupposed.
Approach: They argue that LLMs are vulnerable to Harnad’s symbol grounding problem (SGP), as it has been claimed recently . they trace the origins of the SGP to the computational theory of mind .
Outcome: The proposed model-theoretic semantics does not give rise to the SGP, as it has been claimed in the literature.
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft (2024.findings-emnlp)

Copied to clipboard

Challenge: In the Minecraft Collaborative Building Task, two players collaborate to build a building using 3D blocks.
Approach: They propose to use large language models to model the Builder's sequence of actions in the Minecraft Collaborative Building Task.
Outcome: The proposed methods significantly improve performance over baseline methods and provide detailed analysis for future work.
HSDreport: Heart Sound Diagnosis with Echocardiography Reports (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for heart sound diagnosis are limited to a few fixed categories and do not utilize echocardiography reports, the gold standard in the diagnosis of related diseases.
Approach: They propose a benchmark that mandates the direct utilization of heart sounds obtained from auscultation to predict echocardiography reports.
Outcome: The proposed method outperforms existing methods and existing multimodal LLMs in detecting key abnormalities in heart sounds.
Decoding Decoded: Understanding Hyperparameter Effects in Open-Ended Text Generation (2025.coling-main)

Copied to clipboard

Challenge: Generative large language models generate a high-dimensional probability distribution over all tokens in their vocabulary.
Approach: They conduct extensive sensitivity analyses to determine how hyperparameter choices shape the outputs of generative large language models.
Outcome: The proposed methods influence the distribution of diversity and coherence metrics in human-written text, but the optimal configurations vary across models and tasks.
XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for embedding binary messages into LLM-generated text suffer from key limitations, such as a poor trade-off between text quality and decoding accuracy.
Approach: They propose a method for embedding binary messages into Large Language Model (LLM)-generated text that uses a limited number of tokens to decode and recover the encoded message.
Outcome: The proposed method significantly outperforms existing methods in multiple downstream tasks and will be made publicly available upon acceptance.
PECAN: LLM-Guided Dynamic Progress Control with Attention-Guided Hierarchical Weighted Graph for Long-Document QA (2025.findings-acl)

Copied to clipboard

Challenge: Long-document Question Answering (QA) challenges with large-scale text and long-distance dependencies.
Approach: They propose a method that leverages large language models to control retrieval process . they propose 'attention-based' retrieval methods that construct hierarchical graphs .
Outcome: The proposed method achieves LLM-level performance while maintaining computational complexity comparable to RAG methods.
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks (2024.acl-long)

Copied to clipboard

Challenge: Large Vision/Language Models (LVLMs) are less capable of generating accompanying image sequences.
Approach: They propose a method that integrates a Latent Diffusion Model (LDM) with an LLM to generate captions to maintain semantic coherence of the sequence.
Outcome: The proposed method is preferred by humans in 46.6% of the cases against 26.6% for the second best method.
Simulating Crisis Cognition: A Computational Framework for Hypothesis Generation in Crisis Communication (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable fidelity in simulating social dynamics, yet using them to inform high-stakes crisis policy requires rigorous causal evaluation.
Approach: They propose a framework that functions as an in-silico hypothesis generator to evaluate communication strategies by coupling real-world telemetry with 1,813 agents.
Outcome: The proposed framework provides a rigorous testbed for evaluating strategies before human-subject trials.
To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs (2026.acl-long)

Copied to clipboard

Challenge: Misinformation is on the rise, and the strong writing capabilities of LLMs lower the barrier for malicious actors to produce and disseminate false information.
Approach: They introduce a multilingual parallel dataset of 440 misinformation generation prompt templates and 6,867 entities, spanning 8 languages and 195 countries.
Outcome: The proposed model reduces misinformation generation across languages and countries . it also reduces the risk of misinformation being spread across countries based on the model's performance .
Cross-Lingual Multi-Hop Knowledge Editing (2024.findings-emnlp)

Copied to clipboard

Challenge: Prior work on knowledge editing in monolingual settings focused on a single language, but there are significant gaps in performance between the two settings.
Approach: They propose a cross-lingual multi-hop knowledge editing paradigm for measuring and analyzing the performance of various SoTA knowledge editing techniques in a multilingual setup.
Outcome: The proposed system improves on previous methods in a cross-lingual setting and in English.
Can Large Language Models Understand Argument Schemes? (2025.findings-acl)

Copied to clipboard

Challenge: Argument schemes are stereotypical forms of reasoning that occur in everyday arguments.
Approach: They propose to use large language models (LLMs) to classify argument schemes based on Walton’s taxonomy to employ formal definitions and LLM-generated descriptions to enhance task instructions.
Outcome: The proposed models perform well on annotated and automatically generated arguments, and provide insights for advancing reasoning capabilities in computational argumentation.
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown to be a great success in a wide range of applications ranging from regular NLP-based use cases to AI agents.
Approach: They examine the robustness of existing MUL techniques for their ability to enable leakage-proof forgetting in LLMs.
Outcome: The proposed methods can be used to enable leakage-proof forgetting in LLMs.
Prediction-Augmented Generation for Automatic Diagnosis Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) adopt autoregressive architecture, predicting the next word token based on the preceding context.
Approach: They propose a method that integrates task-specific predictive models as external tools to improve model generation quality and accuracy.
Outcome: The proposed method improves the generation quality and predictive accuracy of large language models in inference-driven tasks.
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models have created significant safety concerns . factuality ability is crucial in determining whether they can be deployed and applied safely and compliantly within specific regions.
Approach: They propose a benchmark to evaluate the factuality of large language models in China . they evaluate the models' ability to provide accurate and reliable information .
Outcome: The proposed benchmark evaluates the factuality abilities of existing LLMs and compares them to LLM abilities.
Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Existing methods rely on entity vector matching, but the purpose of the question is abstract and difficult to match with specific entities. Existing approaches rely only on entity-vector matching, and there is a problem with multi-hop reasoning.
Approach: They propose a framework that constructs reasoning paths from purposes back to conditions using the KG ontology.
Outcome: Experiments on the WebQSP and CWQ datasets show that ORT significantly improves the capability of large language models in knowledge graph question answering tasks (KGQA).
Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning (2024.findings-acl)

Copied to clipboard

Challenge: Paraphrasing of offensive content is a better alternative to content removal, but supervised methods often retain a large portion of the offensiveness of the original content.
Approach: They propose to use In-Context Learning (ICL) to generate usable offensive paraphrases by using large language models.
Outcome: The proposed framework is better than supervised methods on human evaluation and lower toxicity by 76%.
Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown considerable promise for annotation purposes, but questions remain about their ability to capture human label variation (HLV) label variation is genuine disagreement between annotators observed across NLP tasks.
Approach: They investigate how label and explanation variation manifests within and across LLMs with respect to the Natural Language Inference task.
Outcome: The proposed models generate label distributions similar to humans but exhibit distinct, idiosyncratic judgments and disagreement patterns.
Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for text regression lack local grounding and rely on shared representations.
Approach: They propose a distributional regression model with quantile tokens that insert dedicated quantiles into the input sequence.
Outcome: The proposed method outperforms baseline models on the inside Airbnb and StackSample datasets.
How Retrieved Context Shapes Internal Representations in RAG (2026.findings-acl)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) is a widely adopted approach for enhancing large language models with external knowledge.
Approach: They analyze how different types of retrieved documents affect the hidden states of large language models and how these internal representation shifts relate to downstream generation behavior.
Outcome: The results show that context relevancy and layer-wise processing influence internal representations, providing explanations of LLMs’ output behaviors and insights for RAG system design.
Question Answering as Programming for Solving Time-Sensitive Questions (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that Large Language Models (LLMs) have shown remarkable intelligence in question answering.
Approach: They propose to reframe the Question Answering task as Programming to overcome this limitation by leveraging LLMs' superior ability in understanding both natural language and programming language.
Outcome: The proposed approach improves on time-sensitive question answering datasets by 14.5% over baselines.
Reinforcement Tuning for Detecting Stances and Debunking Rumors Jointly with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Social media has become a fertile ground for nurturing rumors and misinformation due to its lack of systematic moderation.
Approach: They propose a framework to enhance the joint predictive capabilities of LLMs for stance detection and rumor verification tasks.
Outcome: The proposed framework outperforms state-of-the-art methods and generalizes to non-LLMs accommodated as task models.
Updating Large Language Models’ Memories with Time Constraints (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) can modify their internal memory by incorporating the latest external knowledge, but in practical applications, outdated information may be inputted into LLMs.
Approach: They propose a two-stage decoupling framework that separates the identification and computation of time constraints into a symbolic system and propose 'selective update' of internal memory based on time constraints.
Outcome: The proposed framework improves ChatGPT performance by 60% and improves state-of-the-art LLM GPT-4.
RaDA: Retrieval-augmented Web Agent Planning with LLMs (2024.findings-acl)

Copied to clipboard

Challenge: Agents powered by large language models inherit important limitations such as the restricted context length, dependency on human-engineered exemplars, and insufficient generalization.
Approach: They propose a novel planning method for Web agents that disentangles planning into two stages: for a new given task, it decomposes tasks into high-level subtasks; and then iteratively synthesizes actions based on dynamically retrieved exemplars.
Outcome: The proposed method decomposes tasks into high-level subtasks and iteratively synthesizes actions based on dynamically retrieved exemplars.
LCR-RAG: Enhancing Logical Consistency in Retrieval-Augmented Generation via Neuro-symbolic Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) is widely used to ground large language models in external knowledge and improve factual accuracy.
Approach: They propose a framework that integrates neuro-symbolic verification with reinforcement learning to optimize logical consistency.
Outcome: The proposed framework outperforms strong RAG baselines on hotpotQA, ASQA, and TriviaQA.
How Do Multilingual Language Models Remember Facts? (2025.findings-acl)

Copied to clipboard

Challenge: Prior research has focused on English monolingual models, but how these mechanisms generalize to non-English languages remains unexplored.
Approach: They analyze three multilingual LLMs to find out how they can generalize recall mechanisms . they find that subject enrichment is language-independent, object extraction is language dependent .
Outcome: The proposed model performs better in multilingual contexts than in English models . the model is more efficient in multi-lingual context, but it is more complex in multilinguistic models compared to English models.
Semantic Aware Linear Transfer by Recycling Pre-trained Language Models for Cross-lingual Transfer (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly incorporating multilingual capabilities, fueling the demand to transfer them into target language-specific models.
Approach: They propose a novel cross-lingual transfer technique that recycles embeddings from target language Pre-trained Language Models to transmit deep representational strengths to LLMs.
Outcome: The proposed technique outperforms existing methods in cross-lingual understanding setups and achieves faster convergence and lower loss during language adaptation.
HalluMeasure: Fine-grained Hallucination Measurement Using Chain-of-Thought Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: HalluMeasure is a new LLM-based hallucination detection mechanism that decomposes an LLM response into atomic claims and evaluates each claim against the provided reference context.
Approach: They propose a new LLM-based hallucination detection mechanism that decomposes an LLM response into atomic claims and evaluates each atomic claim against the provided reference context.
Outcome: The proposed model can detect 3 major categories of hallucinations and 10 more specific subtypes which help to identify reasons behind the hallucinian errors.
AGD: Adversarial Game Defense Against Jailbreak Attacks in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing defenses, including post-training alignment and prompt engineering, struggle with adaptability to out-of-distribution (OOD) attacks.
Approach: They propose an adversarial game-based defense method that dynamically adjusts LLMs’ internal representations to achieve a balanced trade-off between helpfulness and harmlessness.
Outcome: The proposed method improves LLMs’ safety over all baselines.
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data.
Approach: They review training strategies, robustness enhancements, loss functions, and agent-based approaches and outline open challenges and future directions to guide research in this evolving field.
Outcome: The proposed model improves accuracy and accuracy while integrating external dynamic information for improved factual grounding.
ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that Large Language Models (LLMs) augmented with chain-of-thought (CoT) reasoning demonstrate impressive problem-solving abilities.
Approach: They propose a weight-editing approach to reduce overly short reasoning by steering the model along a linear direction in the representation space.
Outcome: The proposed model reduces overly short reasoning and yields significant accuracy gains on multiple math benchmarks.
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have recently exhibited performance gains owing to a wide variety of prompting techniques, including Retrieval-Augmented Generation (RAG), Chain-of-Thought (CoT), and In-Context Learning (ICL).
Approach: They propose a prompt compression method that captures the global context without compromising semantic consistency while detouring the necessity of pseudo-labels for training the compressor.
Outcome: Empirical results show that the proposed method retains key contexts while reducing the prompt length by 80%.
Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates (2026.acl-long)

Copied to clipboard

Challenge: Large language models underperform in languages absent or underrepresented in training data, creating barrier to equitable access for speakers worldwide.
Approach: They propose a selective parameter update strategy that proactively preserves source knowledge by identifying critical parameters critical to maintaining source abilities.
Outcome: Experiments in five typologically diverse languages show that SSU mitigates catastrophic forgetting.
Systematic Assessment of Factual Knowledge in Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing question-answering benchmarks for large language models have limitations regarding factual knowledge coverage, as they focus on generic domains and overlap with pretraining data.
Approach: They propose a framework to assess the factual knowledge of large language models by leveraging knowledge graphs.
Outcome: The proposed framework generates questions and expected answers from the facts stored in a given knowledge graph and evaluates them with KGs in generic and specific domains.
Language Acquisition Device in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are less data-efficient than humans, and pre-pretraining on synthetic languages has been proposed to close this gap.
Approach: They propose to pre-pretrain on MP-STRUCT, a formal language whose strings encode hierarchical composition, feature-based dependencies, and long-distance displacement via MERGE, AGREE, and MOVE.
Outcome: The proposed model outperforms k-Shuffle Dyck despite not being definable in C-RASP despite being hierarchically expressive and circuit-theoretically learnable .
Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Large Reasoning Models (LRMs) often display unstable behaviors, e.g., hallucinating unsupported premises, overthinking simple tasks, and displaying higher sensitivity to prompt variations.
Approach: They propose a graph-based analytical framework that clusters long, verbose CoT outputs into semantically coherent reasoning steps, then constructs directed reasoning graphs to capture contextual and logical dependencies among these steps.
Outcome: The proposed framework enables quantitative evaluation of internal reasoning structure and quality beyond conventional metrics and provides practical insights for prompt engineering and cognitive analysis of LLMs.
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks focus on simple attribution that retrieves textual evidence as references.
Approach: They propose a benchmark to evaluate the ability of large language models to generate reliable attributions.
Outcome: The proposed benchmark evaluates the ability of LLMs to generate long-form answers with reliable and nuanced attributions.
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models are increasingly relied upon to evaluate text outputs of other LLMs . however, concerns persist over the accuracy of these assessments and the potential for misleading conclusions.
Approach: They propose a framework to assess the reliability of Large Language Models (LLMs) they propose ' FBI' framework to examine the proficiency of Evaluator LLMs in assessing four critical abilities .
Outcome: The proposed framework assesses the performance of LLMs in text generation tasks.
Soundwave: Less is More for Speech-Text Alignment in LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing end-to-end speech large language models rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth.
Approach: They propose a training strategy and a novel architecture to address representation space gap and sequence length inconsistency in speech and text.
Outcome: The proposed model outperforms other advanced speech LLMs in speech translation and AIR-Bench speech tasks with only a fraction of the training data.
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning (2024.findings-acl)

Copied to clipboard

Challenge: Recent years have witnessed a substantial increase in the demand for legal services, especially for individuals with modest means.
Approach: They propose a diagnostic legal large language model which uses adaptive lawyer-like diagnostic questions to collect additional case information and then provides high-quality feedback.
Outcome: The proposed model surpasses classical LLMs by providing outstanding performance and a remarkable user experience in the legal domain.
Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing defenses rely on impractical assumptions about trigger settings to mitigate backdoor attacks . a recent study found that small amounts of training data can systematically induce harmful behaviors in large language models.
Approach: They propose a backdoor defense framework that requires no prior knowledge of trigger settings . they use a two-stage process to aggregate backdoor representations and fine-tune recovery .
Outcome: The proposed defense reduces the average Attack Success Rate to 4.41% across multiple benchmarks . the proposed framework generalizes across different types of backdoors, confirming its robustness in practical deployment scenarios.
HAF-RM: A Hybrid Alignment Framework for Reward Model Training (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have focused on enhancing reward models through data improvements, following the conventional training framework for reward models that directly optimizes the predicted rewards.
Approach: They propose a hybrid alignment framework **HAF-RM** that incorporates additional constraint on token-level policy probabilities in addition to the reward score.
Outcome: The proposed framework can supervise the internal preference model at the token level and optimize the mapping layer of the reward model at sequence level.
Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent training-based TTS methods, such as continued reinforcement learning, have surged in popularity, while training-free TTS approaches are gradually fading from prominence.
Approach: They propose a fine-grained sequential scaling method guided by process verification that integrates training-free TTS methods with other classical parallel scaling methods at the step level.
Outcome: Experiments on five instruction-tuned large language models (LLMs) show that training-free TTS methods can extend reasoning performance boundaries.
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have remarkable capabilities but are vulnerable to adversarial “jailbreak” attacks designed to bypass safety guardrails.
Approach: They propose to empower a large language model to be its own red teamer . safety self-play allows the model to act as both the Attacker and Defender .
Outcome: The proposed approach outperforms baselines trained on static adversarial datasets and establishes a new benchmark for proactive safety alignment.
LLMSegm: Surface-level Morphological Segmentation Using Large Language Model (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to morphological segmentation split word into its morphemes . LLMSegm is applicable in low-data settings and low-resourced languages .
Approach: They propose a novel approach to surface-level morphological segmentation leveraging large language models.
Outcome: The proposed method is applicable in low-data settings and low-resource languages.
Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly employed in high-stakes decision-making tasks such as loan approvals.
Approach: They evaluate the performance and fairness of LLMs on serialized loan approval datasets from Ghana, Germany, and the United States.
Outcome: The model’s zero-shot and in-context learning (ICL) capabilities are evaluated on loan approval datasets from Ghana, Germany, and the United States.
ToolGrad: Efficient Tool-use Dataset Generation with Textual “Gradients” (2026.findings-acl)

Copied to clipboard

Challenge: Prior work synthesizes tool-use LLM datasets by first generating a user query, then complex tool-using annotations like DFS.
Approach: They propose an agentic framework that synthesizes user queries and generates valid tool-use chains . they propose a dataset with more complex tool use, lower cost, and almost 100% pass rate .
Outcome: Experiments show that tools trained on ToolGrad outperform expensive baseline datasets and proprietary LLMs.
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Evaluating debate speeches requires a deep understanding of arguments at multiple levels.
Approach: They propose a benchmark task for LLM judges based on annotated debate speeches . they analyze the judgment capabilities and behavior of frontier LLMs .
Outcome: The proposed task requires a comprehensive understanding of argumentation and its arguments.
Generation-Augmented Retrieval: Rethinking the Role of Large Language Models in Zero-Shot Relation Extraction (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Relation Extraction (RE) emphasize Zero-Shot methodologies, aiming to recognize unseen relations between entities with no annotated data.
Approach: They propose a plug-in retrieval adjuster that allows rapid fine-tuning without accessing LLMs’ parameters.
Outcome: The proposed model demonstrates comparable performance on multiple benchmarks.
Adaptive LoRA Merge with Parameter Pruning for Low-Resource Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for adapting LLMs to low-resource tasks keep LoRA parameters frozen and the low-level problem out of their scope.
Approach: They propose a LoRA merge method that updates and prunes LoRA parameters through fine-tuning with minimal target task data.
Outcome: The proposed method improves performance on a low-resource language generation task and improves on previous methods.
Tree of Problems: Improving structured problem solving with compositionality (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable performance across multipletasks through in-context learning.
Approach: They propose a Tree of Problems (ToP) that is a simpler version of Tree of Thoughts (toT) they propose 'in-context learning' is the ability of Large Language Models (LLMs) to perform a task with the help of a few demonstrations within their context.
Outcome: The proposed approach outperforms ToT and GoT and performs better on complex reasoning tasks.
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance? (2024.emnlp-main)

Copied to clipboard

Challenge: Recent large language models (LLMs) have demonstrated remarkable generalization abilities in mathematics and reasoning tasks.
Approach: They pre-trained decoder-based language models from scratch using ten programming languages and three natural language datasets.
Outcome: The proposed models outperform natural languages on logical reasoning tasks.
Assessing French Readability for Adults with Low Literacy: A Global and Local Perspective (2025.emnlp-main)

Copied to clipboard

Challenge: illiterate individuals are persons aged 15 years and above who cannot read and write with understanding a short simple statement on their everyday life.
Approach: They propose a novel approach to assess french text readability for adults with low literacy skills using a global and segment-level difficulty scale.
Outcome: The proposed approach addresses both global (full-text) and local (segment-level) difficulty scales.
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Autoregressive (AR) decoding in large language models is latency-bounded by strictly sequential token generation.
Approach: They propose a diffusion-based drafter that proposes multi-token candidates and then verifies them in parallel by the target model.
Outcome: The proposed drafter generates multi-token proposals in a single forward pass while remaining compatible with standard AR verifiers.
Penetrating Linguistic Disguises: A Slang-aware Label-Aligned Framework for Fine-Grained Toxicity Extraction in Chinese Hate Speech Detection (2026.findings-acl)

Copied to clipboard

Challenge: Flexible word boundaries and linguistic obfuscation, particularly slang, challenge precise span-level hate speech detection in Chinese.
Approach: They propose a Slang-aware Label-Aligned Framework that maps slang to explicit hate semantics and uses task-specific branches to mitigate feature interference.
Outcome: The proposed framework reduces ambiguity by mapping obscure slang to explicit hate semantics.
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Hallucination is a significant barrier to the effective application of Large Language Models (LLMs).
Approach: They propose an Attention-Guided SElf-Reflection approach for hallucination detection in Large Language Models.
Outcome: The proposed method significantly outperforms existing methods in zero-shot hallucination detection on four widely-used LLMs across three different halluciation benchmarks.
Exploring the Role of Mental Health Conversational Agents in Training Medical Students and Professionals: A Systematic Literature Review (2025.findings-acl)

Copied to clipboard

Challenge: This systematic review analyses 38 studies on AI-powered conversational agents in mental health education and training . traditional training methods provide valuable but expensive and inherently limited learning opportunities . early pioneers like Woebot and Wysa demonstrated a groundbreaking insight: machines could engage in meaningful therapeutic interactions.
Approach: They analyse 38 studies on AI-powered conversational agents in mental health education and training . findings reveal that AI-based approaches dominate the field, with training as the application area being the most prevalent .
Outcome: The systematic review of 38 studies on AI-powered conversational agents in mental health education and training (MHET) reveals that AI-based approaches dominate the field, with training as the application area being the most prevalent.
CaT-Bench: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in Plans (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies on reasoning in plans focus on classical problems, simulated environments, or restricted language such as PDDL, but real-world plans cannot be tested to test for correctness and reliability.
Approach: They propose a benchmark question that tests whether a step must necessarily occur before or after another in cooking recipe plans.
Outcome: The proposed question-driven evaluation shows that SOTA LLMs are underwhelming and biased towards predicting dependence more often, but the best F1 result is 0.73.
RiOT: Efficient Prompt Refinement with Residual Optimization Tree (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for automatic prompt optimization face two challenges: lack of diversity and semantic drift.
Approach: They propose a framework for automatic prompt optimization that iteratively refines prompts through text gradients and selects the best prompt using perplexity.
Outcome: The proposed framework outperforms existing prompt optimization methods and manual prompting on commonsense, mathematical, logical, temporal, and semantic reasoning benchmarks.
CodeJudge: Evaluating Code Generation with Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown promising performance in code generation, but how to reliably evaluate code generated by LLMs remains a challenging problem.
Approach: They propose a framework that leverages Large Language Models to evaluate the semantic correctness of generated code without the need for test cases.
Outcome: The proposed framework outperforms existing methods on four code generation datasets and five programming languages.
Explicit Bayesian Inference to Uncover the Latent Themes of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive generative capabilities, yet their inner mechanisms remain largely opaque.
Approach: They propose a variational autoencoder-based neural topic model to interpret LLMs generation process through an explicit Bayesian framework by inferring latent topic variables via variational inference.
Outcome: The proposed model outperforms state-of-the-art topic models on intrinsic measures of coherence and diversity on multiple datasets and shows significant gains on classification and summarization tasks.
Adaptive Spatial and Temporal Redundancy Optimization for Efficient Reasoning in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing research to improve CoT efficiency falls into three categories, each with distinct limitations.
Approach: They propose a training-free framework that addresses both dimensions of CoT reasoning by applying a progressive precision reduction strategy coupled with an entropy-based confidence mechanism for adaptive termination.
Outcome: Empirical results show that the proposed framework achieves 11.3 efficiency gain without compromising accuracy.
Rule-Guided Extraction: A Hierarchical Rule Optimization Framework for Document-Level Event Argument Extraction (2025.findings-emnlp)

Copied to clipboard

Challenge: Document-level event argument extraction (EAE) is a critical task in natural language processing.
Approach: They propose an LLM-driven HiErarchical Rule Optimization framework that iteratively generates and selects optimal hierarchical rules.
Outcome: The proposed framework outperforms few-shot supervised methods and outperformed state-of-the-art prompting baselines.
Encoding Spreadsheets for Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Spreadsheets are characterized by their extensive two-dimensional grids, flexible layouts, and varied formatting options, which pose significant challenges for large language models (LLMs).
Approach: They propose a structural-anchor-based compression, inverse index translation, and data-format-aware aggregation module to compress spreadsheets effectively.
Outcome: The proposed method outperforms the existing model in GPT4 and achieves a state-of-the-art 78.9% F1 score.
Prediction Hubs are Context-Informed Frequent Tokens in LLMs (2025.acl-long)

Copied to clipboard

Challenge: Hubness is a tendency for a few points to be among the nearest neighbours of a disproportionate number of other points.
Approach: They show that only large-scale representation comparisons are not characterized by hubness . they show that hubs are the result of context-modulated frequent tokens .
Outcome: The results show that the comparison between context and unembedding vectors does not result in hubness . the findings suggest that hubness is not a negative property that needs to be mitigated when LLMs are being used for next token prediction.
From Isolated Scoring to Collaborative Ranking: A Comparison-Native Framework for LLM-Based Paper Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are currently used to evaluate scientific papers by assigning an absolute score to each paper independently.
Approach: They propose a comparison-native framework for paper evaluation that integrates comparison into both data construction and model learning.
Outcome: The proposed framework achieves an average relative improvement of 21.8% over the strong baseline DeepReview-14B, while exhibiting robust generalization to five previously unseen datasets.
Dialect Normalization using Large Language Models and Morphological Rules (2025.findings-acl)

Copied to clipboard

Challenge: Natural language understanding systems struggle with low-resource languages, including many dialects of high-resourced ones.
Approach: They propose a method that combines rule-based linguistically informed transformations and large language models with targeted few-shot prompting without any parallel data.
Outcome: The proposed method is able to transform dialectal text into a standard variety while maintaining as much of the original meaning as possible.
DIESEL: A Lightweight Inference-Time Safety Enhancement for Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models generate outputs that are not aligned with human values, such as toxic content, malicious use cases, and vulnerabilities to adversarial jailbreak attacks.
Approach: They propose a lightweight inference-guidance technique that can be seamlessly integrated into any autoregressive LLM to semantically filter undesirable content during generation.
Outcome: The proposed technique can be integrated into any autoregressive LLM to semantically filter undesirable content during generation.
ReflectDiffu: Reflect between Emotion-intent Contagion and Mimicry for Empathetic Response Generation via a RL-Diffusion Framework (2025.acl-long)

Copied to clipboard

Challenge: Existing models for empathetic dialogue generation neglect the intricate interplay between emotion and intent, leading to suboptimal controllability of empathy.
Approach: They propose a framework that integrates emotion contagion and intent mimicry to enhance empathetic response generation.
Outcome: The proposed framework outperforms existing models in relevance, controllability, and informativeness.
Hallucination Detection in LLMs Using Spectral Features of Attention Maps (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable performance across tasks but remain prone to hallucinations.
Approach: They propose a method that uses attention maps to detect hallucinations . they propose to use top-k eigenvalues of the attention maps as input to probes .
Outcome: The proposed method achieves state-of-the-art hallucination detection performance among attention-based methods.
Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items? (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for estimating the cognitive complexity of reading comprehension items are expensive, time-consuming, and subject to rater variability.
Approach: They propose to use two dimensions to estimate cognitive complexity of RC items to focus on evidence Scope and transformation level to estimate the cognitive complexity.
Outcome: The proposed models can estimate the cognitive complexity of items by focusing on two dimensions—Evidence Scope and Transformation Level—that indicate the degree of cognitive burden involved in reasoning about the answer.
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test (2025.emnlp-main)

Copied to clipboard

Challenge: Multilingual and cross-cultural WAT reveal how culture modulates perceptual and interactive patterns.
Approach: They propose to embed cultural-specific semantic associations directly within large language models (LLMs) to address cultural preference.
Outcome: The proposed model significantly improves cross-cultural alignment, capturing diverse semantic associations.
MIBench: Evaluating Multimodal Large Language Models over Multiple Images (2024.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks and MLLMs focus on single-image input scenarios, leaving performance of ML models when handling multiple images underexplored.
Approach: They propose a benchmark to evaluate fine-grained abilities of multimodal large language models in multi-image scenarios.
Outcome: The proposed benchmark categorizes the multi-image abilities into three scenarios: MII, MKS and MIC.
Do We Know What LLMs Don’t Know? A Study of Consistency in Knowledge Probing (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for probing knowledge gaps in large language models are inconsistent and inconsistent.
Approach: They propose a process based on input variations and quantitative metrics to evaluate probing methods that are inconsistent on knowledge gaps.
Outcome: The proposed process exposes two dimensions of inconsistency in knowledge gap probing.
Intent-aware Schema Generation and Refinement for Literature Review Tables (2025.findings-emnlp)

Copied to clipboard

Challenge: ambiguity in reference-based evaluations and lack of editing/refinement methods have slow progress on schema generation.
Approach: They propose a method for augmenting unannotated table corpora with synthesized intents . they propose prompted workflows and fine-tuned models to improve schema generation .
Outcome: The proposed approach significantly improves baseline performance in reconstructing reference schemas.
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming (2025.acl-long)

Copied to clipboard

Challenge: Existing jailbreak techniques rely on single-round interactions, pro-Corresponding author.
Approach: They propose a multi-turn safety alignment framework to address the challenge of securing large language models in multi-round interactions.
Outcome: The proposed framework exhibits state-of-the-art attack capabilities while improving safety performance on safety benchmarks.
Steering Away from Refusal: A Black-box Jailbreak Method Based on First-Token Distribution (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to analyze black-box jailbreaks lack direct optimization signals to refine adversarial prompts.
Approach: They propose a distribution-jailbreak attack method that selects effective jailbreak templates and iteratively optimizes adversarial suffixes by maximizing the KL divergence from the standard refusal distribution.
Outcome: The proposed method achieves state-of-the-art Attack Success Rate (ASR) on all tested open-source models and delivers over 94% ASR on GPT-4.1.
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once (2026.acl-long)

Copied to clipboard

Challenge: Recent Large Reasoning Models (LRMs) lack a narrow evaluation paradigm . a single-question evaluation setup suffers from two major limitations .
Approach: They propose a stress-testing framework that exposes LRMs to multiple problems simultaneously.
Outcome: The proposed framework outperforms existing models on reasoning benchmarks and state-of-the-art models.
Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making (2025.emnlp-main)

Copied to clipboard

Challenge: Existing safety evaluations rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail.
Approach: They propose a framework for systematically evaluating the physical safety of LLMs in embodied decision making.
Outcome: The proposed framework assesses the physical safety of LLMs in embodied decision making.
Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective in enhancing LLMs’ short-context reasoning but falters in long-contemporal scenarios requiring precise grounding and multi-hop reasoning.
Approach: They propose a framework that constructs high-difficulty, multi-hop long-context QA pairs with inherent reasoning chains to overcome this bottleneck.
Outcome: The proposed framework outperforms RLVR baselines and matches frontier LLMs while using far fewer parameters.
Beyond Chunks and Graphs: Retrieval-Augmented Generation through Triplet-Driven Thinking (2026.findings-acl)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) is critical for reducing hallucinations and incorporating external knowledge into Large Language Models (LLMs).
Approach: They propose a framework that leverages an LLM to decompose questions into searchable triplets with placeholders.
Outcome: Empirical results show that T2RAG outperforms state-of-the-art multi-round and Graph RAG methods while reducing retrieval costs by up to 45%.
DiaLLMs: EHR-Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction (2025.findings-acl)

Copied to clipboard

Challenge: Existing medical LLMs focus primarily on diagnosis recommendation, limiting their clinical applicability.
Approach: They propose a medical LLM that integrates heterogeneous EHR data into clinically grounded dialogues.
Outcome: The proposed model outperforms baselines in clinical test recommendation and diagnosis prediction.
ICL CIPHERS: Quantifying ”Learning” in In-Context Learning via Substitution Ciphers (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies suggest that In-Context Learning operates in dual modes . however, disentangling these modes remains a challenging goal .
Approach: They propose a class of task reformulations based on substitution ciphers borrowed from classic cryptography.
Outcome: The proposed model can solve tasks with a BIJECTIVE mapping, but it requires 'deciphering' the latent cipher.
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks frame long-term memory evaluation as fact retrieval from past conversations, providing limited insight into agents’ ability to consolidate memory over time or handle frequent knowledge updates.
Approach: They propose a long-term memory benchmark that evaluates three memory-grounded tasks: remembering, reasoning, and recommending.
Outcome: The proposed benchmarks evaluate three tasks: remembering, reasoning, and recommending.
LaCo: Layer-wise Compensation for Pruned Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for predicting performance degradations of Large Language Models (LLMs) neglect the structural distortions caused by sparsity.
Approach: They propose a framework that reorients the recovery paradigm from global adaptation to hierarchical representation alignment by sequentially optimizing each layer to reconstruct the model's hidden states.
Outcome: The proposed framework surpasses parameter-efficient baselines in perplexity reduction and zero-shot reasoning.
SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks conflate coordination ability with role-based priors.
Approach: They propose a role-free benchmark for evaluating free-form collaboration under information silos.
Outcome: The proposed benchmark systematically probes coordination capabilities under information silos using 54 configurations and 3 frontier LLMs.
Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for testing time scales treat reasoning traces or tokens equally, ignoring substantial variations in trajectory quality and localized logical failures.
Approach: They propose a chronological reasoning scorer that models each trajectory as a time series.
Outcome: The proposed method achieves relative improvements of 34.21% over Pass@128 and 22.70% over Maj@135 on HMMT25, highlighting its effectiveness.
Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding (2025.acl-long)

Copied to clipboard

Challenge: Existing studies show that MBR decoding improves model generation performance . however, the theoretical underpinnings of these results remain uncertain .
Approach: They propose a theoretical interpretation of MBR decoding from the perspective of bias–diversity decomposition.
Outcome: The proposed method improves the quality estimation of hypotheses by decomposing bias and diversity into two main factors.
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in education, yet their default usefulness conflicts with pedagogical principles.
Approach: They propose an adversarial student agent that they fine-tune to jailbreak LLM-based tutors and propose a benchmark to evaluate tutor robustness.
Outcome: The proposed model fine-tunes to jailbreak LLM-based tutors, and shows that they perform well under adversarial student attacks.
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Prior studies assess instruction adherence in the model’s main responses, but it is also critical for large reasoning models to follow user instructions throughout their reasoning process.
Approach: They propose a systematic benchmark for assessing reasoning instruction following to assess the model's adherence to instructions.
Outcome: The proposed benchmark reduces the risk of undesirable shortcuts, hallucinations, or reward hacking within reasoning traces.
GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQL (2026.acl-long)

Copied to clipboard

Challenge: despite growing interest in NL2GQL, benchmarking progress has been constrained by the lack of resources that are simultaneously large-scale, cross-domain, and cross-dialect.
Approach: They propose a framework that integrates NL2SQL-to-NL2GQL conversion with graph-native data generation.
Outcome: The proposed framework supports execution-based evaluation on Cypher and ISO-GQL, covering hundreds of graph databases and over 20k natural language questions for each dialect.
De-Anonymization at Scale via Tournament-Style Attribution (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are rapidly gaining widespread adoption in real-world use . authors propose a method for attributing authorship among tens of thousands of candidate texts .
Approach: They propose a large-language-model-based method for attributing authorship among tens of thousands of candidate texts.
Outcome: The proposed method improves accuracy and ranking precision over previous approaches.
A Causal Lens for Evaluating Faithfulness Metrics (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability, but they may not reflect the model’s truereasoning faithfully.
Approach: They propose a testbed framework for evaluating faithfulness metrics for natural language explanations using diagnosticity and model-editing methods.
Outcome: The proposed framework evaluates faithfulness metrics for natural language explanations on four tasks including fact-checking, analogy, object counting, and multi-hop reasoning.
FISTAPruner: Layer-wise Post-training Pruning for Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing pruning methods require inefficient retraining for billion-scale LLMs or rely on heuristicically designed metrics to determine pruning masks, leading to performance degradation.
Approach: They propose a convex optimization model that induces sparsity in large language models by leveraging FISTA.
Outcome: The proposed method can remove 50% of model parameters while retaining 98.6% and 95.6% of the zero-shot performance.
Negative Matters: Multi-Granularity Hard-Negative Synthesis and Anchor-Token-Aware Pooling for Enhanced Text Embeddings (2025.acl-long)

Copied to clipboard

Challenge: Text embedding models are used for various natural language processing tasks such as sentiment analysis, text clustering, and content-based information retrieval.
Approach: They propose a synthesis framework that leverages large language models to generate diverse negative samples with varying levels of similarity with the query.
Outcome: The proposed framework achieves state-of-the-art performance surpassing existing synthesis strategies with synthetic data and when combined with public retrieval datasets.
HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims (2025.acl-long)

Copied to clipboard

Challenge: Identifying checkworthy claims is the first step, but detection methods struggle with content that is (1) multimodal, (2) from diverse domains, and (3) synthetic.
Approach: They propose a dataset for multimodal checkworthiness detection with 27K real-world and synthetic image/claim pairs.
Outcome: The proposed dataset compares lightweight text-based encoders to multimodal models but only focus on claim-like content.
LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to optimize large language models rely on manual design or focus on optimizing individual components.
Approach: They propose a LaMDAgent framework that constructs and optimizes end-to-end post-training pipelines by exploring various model improving methods, objects, and their applied orderings based on task-based feedback.
Outcome: The proposed framework achieves a 9.0-point gain in tool-use accuracy without degrading instruction-following, and reduces computational costs.
Safety Sidecar: Reflection-Driven Runtime Control for Safer Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety controls fail to provide runtime intervention or cross-architecture portability for autonomous LLM agents.
Approach: They propose a model-agnostic, plug-and-play module to provide arbitrary agent safety control and auditability.
Outcome: The proposed module improves the secure-solution rate by 2.9–11.2 percentage points . it adds only 3.2s to end-to-end latency and a negligible average cost of 5.37 10-4 per scenario .
Reflective Agreement: Combining Self-Mixture of Agents with a Sequence Tagger for Robust Event Extraction (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for event extraction are limited in their ability to recall nuanced or rare events.
Approach: They propose a hybrid approach that leverages a self-mixture of agents and a discriminative sequence tagger to resolve ambiguities and enhance overall event prediction quality.
Outcome: The proposed approach outperforms existing state-of-the-art methods across three benchmark datasets.
Improving Rule-based Reasoning in LLMs using Neurosymbolic Representations (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) face challenges in reliably solving reasoning tasks, especially when solving tasks that require strict rule following.
Approach: They propose a method that encodes hidden states into neurosymbolic vectors and decodes them into a neurosample vector space to enable problem-solving within a neural space.
Outcome: The proposed method shows an average of 88.6% lower cross-entropy loss and 15.4 times more problems correctly solved on a suite of mathematical reasoning tasks compared to chain-of-thought prompting and supervised fine-tuning (LoRA).
CORD: Bridging the Audio–Text Reasoning Gap via Weighted On-policy Cross-modal Distillation (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models (LALMs) exhibit a degradation in knowledge and reasoning capabilities . empirical results show that CORD significantly bridges the audio–text performance gap .
Approach: They propose a framework that performs online cross-modal self-distillation to bridge the acoustic-semantic gap between LALMs and text-based models.
Outcome: The proposed framework bridges the acoustic-semantic gap between LALMs and text-based models . it employs on-policy reverse KL divergence with importance-aware weighting .
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Toxic content encompasses a wide spectrum of terminologies whose definitions vary by platform.
Approach: They propose a 2-stage framework for explainable content moderation using Large Language Models (LLMs) they leverage LLMs’ own outputs to generate synthetic explanations for correct and incorrect labels . they refine explanation quality through cross-model training, allowing weaker models to align with stronger ones.
Outcome: Experiments on 3 benchmarks show that the proposed framework achieves 13% macro-F1 improvement over few-shot baselines using only 6-57% of training data.
Explicit Learning and the LLM in Machine Translation (2025.emnlp-main)

Copied to clipboard

Challenge: a growing number of researchers are examining whether large language models can learn to translate a "new" language using grammar books.
Approach: They examine an LLM's ability to learn new languages using grammar books . authors suggest alternative fine-tuning strategies to improve explicit learning .
Outcome: The proposed model can learn low-resource languages described in grammar books but lacking extensive corpora.
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks rely heavily on text-based evaluation and largely ignore paralinguistic cues such as prosody, emotion, and speaker traits.
Approach: They propose a speech-native benchmark for evaluating instruction-following S2S models with explicit assessment of both semantic understanding and paralinguistic expression.
Outcome: The proposed system enables more natural, robust, and human-aligned speech agents.
Sounding vs. Being an Expert: Disentangling Authority, Register and Cultural Impact in Sycophantic LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models exhibit sycophancy, a tendency to align with user assertions even when they conflict with factual correctness.
Approach: They propose an adversarial evaluation framework that isolates two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register).
Outcome: The proposed framework disentangles two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register).
PruneCD: Contrasting Pruned Self Model to Improve Decoding Factuality (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to decode early exit logits in large language models are lacking in factuality and accuracy.
Approach: They propose a contrastive decoding method that constructs the amateur model via layer pruning rather than early exit.
Outcome: The proposed method improves factuality with minimal inference overhead and is robust and practical.
Chimera: Compositional Jailbreak Attacks on LLMs via Judgment-Driven Search over Heterogeneous Strategies (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for evaluating large language models face two limitations: they explore homogeneous transformations in isolation and rely on brittle judgment metrics that misclassify non-refusal hallucinations as successful attacks.
Approach: They propose a framework that generates compositional jailbreak attacks via judgment-driven search over heterogeneous strategies.
Outcome: The proposed framework generates compositional jailbreak attacks over heterogeneous strategies . strongREJECT++ improves attack success rates and transferability compared to state-of-the-art .
No Need for Explanations: LLMs can implicitly learn from mistakes in-context (2025.emnlp-main)

Copied to clipboard

Challenge: Existing literature assumes that correct answers to large language models must be accompanied by comprehensive rationales to be helpful.
Approach: They propose to show incorrect answers to Large Language Models (LLMs) as a popular strategy to improve their performance in reasoning-intensive tasks.
Outcome: The proposed approach outperforms chain-of-thought prompting in math reasoning tasks.
Exploring Large Language Models for Detecting Mental Disorders (2025.emnlp-main)

Copied to clipboard

Challenge: Detecting mental disorders and patient emotions through text analysis and machine learning is of increasing interest to researchers over the past decade.
Approach: They compare the performance of traditional machine learning methods and encoder-based models on Russian-language datasets to those of large language models.
Outcome: The proposed models outperform traditional methods on small and noisy datasets, but can perform comparable to language models when trained on patients with clinically confirmed depression.
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for large language models fail to capture complex interplay between functionality and security.
Approach: They propose a benchmark for secure code generation constructed from real-world, high-risk Java repositories.
Outcome: The proposed benchmarks highlight the gap between functional and secure code generation in LLMs.
Beyond Static Synthetic Noise: Assessing the Robustness of Large Language Models to Natural Context Variation in the Real World (2026.findings-acl)

Copied to clipboard

Challenge: Current robustness evaluation methods rely on static synthetic perturbations to stress-test models.
Approach: They propose a framework for automatically evaluating QA models under naturally occurring textual perturbations by replacing context passages with revised Wikipedia edit histories.
Outcome: The proposed framework replaces context passages with revised Wikipedia edit histories to improve model performance.
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) can automatically draft reviews, but determining whether they are trustworthy requires systematic evaluation.
Approach: They propose an automatic focus-level evaluation pipeline based on two sets of facets . authors evaluated LLM reviews at surface-level or content-level .
Outcome: The proposed framework enables automatic evaluation of paper reviews based on two sets of facets . the framework compared open review paper reviews with human experts on validity, clarity, novelty .
KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality (2026.acl-long)

Copied to clipboard

Challenge: Existing Reinforcement Learning approaches rely on outcome-oriented rewards to reinforce fabricated reasoning paths when the final answer is correct.
Approach: They propose a framework that integrates factual supervision directly into reasoning . they propose to decompose chain of thought into atomic facts and verify them against ground-truth knowledge .
Outcome: The proposed framework reduces the Incorrect Rate on SimpleQA by 20.3% while maintaining strong performance on complex reasoning benchmarks.
Too Long, Do Re-weighting for Efficient LLM Reasoning Compression (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have recently achieved remarkable progress on complex reasoning tasks by leveraging extended Chain-of-Thought (CoT) techniques.
Approach: They propose a method that uses Extended Chain-of-Thought (EFT) to reduce the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning.
Outcome: The proposed method reduces the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning.
From Selection to Refinement: Iterative Optimization for Instruction Data (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to optimize instruction tuning datasets face two main challenges: unreasonable pruning of potentially valuable low-quality data and the persistence of noise or semantic drift during revision.
Approach: They propose an automated iterative framework for instruction data optimization that prunes low-quality data and refines low quality data using feedback-driven iteration.
Outcome: The proposed framework outperforms state-of-the-art methods on seven public benchmark datasets with high data efficiency.
SiLP: Enhancing Non-Dominant Language Capabilities with a Selective Bidirectional Language Projection Framework (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to improve performance of large language models rely on additional training objectives or language-specific parameters.
Approach: They propose a bidirectional language projection framework that enables efficient multilingual alignment and language shift using the intrinsic parameters.
Outcome: The proposed framework improves performance of non-dominant languages and improves internal representations.
NITI: Neural Plan Concretization for Incremental Execution, Bridging and Trigger Inference from Underspecified Human Policies (2026.findings-acl)

Copied to clipboard

Challenge: Using NITI, we examine the performance of a safety-critical automated insulin dosing task with minimal contextualization infence overhead.
Approach: They propose a framework that treats large language models as execution-time concretizers of human intent that incrementally executes abstract policies via verifier-grounded interfaces.
Outcome: The proposed framework outperforms one-shot and chain-of-thought baselines on two structurally distinct embodied domains: a world cubing championship 22 Rubik’s Cube scramble and a safety-critical automated insulin dosing task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations