Papers with feedback

300 papers
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop (2020.acl-srw)

Copied to clipboard

Challenge: ACL 2020 student research workshop has received considerable attention . 137 submissions were received, including 12 non-archival papers .
Approach: the ACL 2020 student research workshop is a forum for student researchers in computational linguistics and natural language processing.
Outcome: the student research workshop has received considerable attention reflecting the growth of the field.
Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (P19-3)

Copied to clipboard

Challenge: The best demo paper was selected by the demo chairs based on the feedback received by reviewers.
Approach: The best demo paper was selected by the demo chairs based on the feedback received by reviewers.
Outcome: The best demo paper was selected by the demo chairs based on the feedback received by reviewers.
Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (D18-1)

Copied to clipboard

Challenge: EMNLP 2018 received a record-breaking 2,231 valid submissions, a 48% increase over EMnLP 2017 . EMNA 2018 will have 14 workshops, 6 tutorials, 3 invited speakers, 351 long paper presentations, 198 short paper presentations and 10 TACL paper presentations .
Approach: EMNLP 2018 will have 14 workshops, 6 tutorials, 3 invited speakers, 351 long paper presentations, 198 short paper presentations and 10 TACL paper presentations.
Outcome: EMNLP 2018 received a record-breaking 2,231 valid submissions, a 48% increase over EMnLP 2017 . the program co-chairs put tremendous care into every decision, big and small, and handled numerous inquiries and requests .
Beyond the Gold Standard in Analytic Automated Essay Scoring (2025.acl-srw)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) is a new approach to assessing writing practice . traditional holistic scoring methods are not reliable and lack formative feedback in the classroom.
Approach: They propose to combine analytic and holistic AES to create a system that learns from individual raters instead of gold standard labels.
Outcome: The proposed system learns from individual raters instead of gold standard labels.
ALTER: Auxiliary Text Rewriting Tool for Natural Language Generation (D19-3)

Copied to clipboard

Challenge: Generative modeling of editing text with respect to control attributes has seen increasing progress over the past few years.
Approach: They propose an auxiliary text rewriting tool that facilitates the rewrite process for natural language generation tasks.
Outcome: The proposed tool facilitates the rewriting process for natural language generation tasks, such as paraphrasing, text simplification, fairness-aware text rewrite, and text style transfer.
Multi-Programming Language Sandbox for LLMs (2025.acl-demo)

Copied to clipboard

Challenge: MPLSandbox is an out-of-the-box multi-programming language sandbox designed to provide unified and comprehensive feedback from compiler and analysis tools for Large Language Models (LLMs).
Approach: They propose a multi-programming language sandbox that provides unified feedback from compilers and analysis tools for Large Language Models.
Outcome: The proposed multi-language sandbox can provide comprehensive feedback from compilers and analysis tools for large language models (LLMs).
Doc-React: Multi-page Heterogeneous Document Question-answering (2025.acl-short)

Copied to clipboard

Challenge: Existing methods for integrating information across multiple modalities are suboptimal for multi-page, multimodal documents.
Approach: They propose an adaptive iterative framework that balances information gain and uncertainty reduction at each step.
Outcome: The proposed framework captures relevant multimodal content and achieves strong performance on complex QA tasks.
Robustness Gym: Unifying the NLP Evaluation Landscape (2021.naacl-demos)

Copied to clipboard

Challenge: Existing tools cater to specialized set of evaluations and provide no clear way to leverage or share findings from prior evaluations.
Approach: They propose a toolkit that unifies 4 evaluation paradigms to provide a common platform for evaluation.
Outcome: The proposed evaluation toolkit unifies 4 evaluation paradigms and is under active development.
Toward Optimal LLM Alignments Using Two-Player Games (2025.findings-emnlp)

Copied to clipboard

Challenge: Alignment of large language models (LLM) is a process that ensures the model’s responses to user prompts align with human intentions and social values.
Approach: They propose an alignment method based on a two-agent game consisting of an adversarial agent and a defensive agent.
Outcome: The proposed method improves on a two-agent game with an adversarial agent and a defensive agent.
Generative Reviewer Agents: Scalable Simulacra of Peer Review (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing peer review mechanisms are limited by the small fraction of researchers with established networks.
Approach: They propose a system that extends a large language model and equips agents with reviewer personas derived from historical data to enable generative reviewers.
Outcome: The proposed architecture performs comparable to human reviewers in providing detailed feedback and predicting paper outcomes.
Template-guided Grammatical Error Feedback Comment Generation (2023.eacl-srw)

Copied to clipboard

Challenge: Writing corrective feedback on learner text is widespread in language education, but it can be time-consuming for teachers.
Approach: They propose to use feedback comment generation to generate explanatory notes for learners by categorizing comments and constraining outputs of noisy classes.
Outcome: The proposed scheme can be used to generate feedback comment corpora using a broader scope than existing typologies focused on error correction.
PAIR: Prompt-Aware margIn Ranking for Counselor Reflection Scoring in Motivational Interviewing (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to provide constructive feedback to counselors are limited by the time and cost involved.
Approach: They propose a system that takes as input a client prompt and a counselor response and outputs a score indicating the level of reflection in the counselor response.
Outcome: The proposed model outperforms baselines on different metrics and can be used to provide useful feedback to counseling trainees.
QueryExplorer: An Interactive Query Generation Assistant for Search and Exploration (2024.naacl-demo)

Copied to clipboard

Challenge: Formulating effective search queries can be a daunting task for users when they lack expertise in a specific domain or are not proficient in the language of the content.
Approach: QueryExplorer is an interactive query generation, reformulation, and retrieval interface with support for Hug-gingFace generation models and PyTerrier’sretrieval pipelines and datasets.
Outcome: QueryExplorer is an interactive query generation, reformulation, and retrieval interface with support for Hug-gingFace generation models and PyTerrier’sretrieval pipelines and datasets, and extensivelogging of human feedback.
Text2DB: Integration-Aware Information Extraction with Large Language Model Agents (2024.findings-acl)

Copied to clipboard

Challenge: Current methods for information extraction (IE) focus on integrating IE output with the database . a long-overlooked question is what counts as "relevant knowledge"
Approach: They propose a task that emphasizes integration of IE output and the database . they introduce a benchmark and an LLM agent framework for this task .
Outcome: The proposed task integrates IE output and the target database (or knowledge base) it meets common demands such as data infilling, row population, and column addition .
A Neural, Interactive-predictive System for Multimodal Sequence to Sequence Tasks (P19-3)

Copied to clipboard

Challenge: a neural interactive-predictive system is used to tackle multimodal sequence to sequence tasks . it generates text predictions to different sequence to sequencing tasks, including machine translation, image and video captioning.
Approach: They present a neural interactive-predictive system for tackling multimodal sequence to sequence tasks.
Outcome: The proposed system reduces human effort during the correction process by providing alternative hypotheses.
Dia-Lingle: A Gamified Interface for Dialectal Data Collection (2025.acl-demo)

Copied to clipboard

Challenge: Dialects suffer from the scarcity of textual resources and are largely spoken rather than written.
Approach: They propose a gamified interface that combines active learning with gamification to enhance the dialect corpus.
Outcome: The proposed interface demonstrates high levels of user satisfaction while requiring minimal effort.
Vis-Eval Metric Viewer: A Visualisation Tool for Inspecting and Evaluating Metric Scores of Machine Translation Output (N18-5)

Copied to clipboard

Challenge: Many metrics have been proposed for Machine Translation (MT) that compare system translations against human references.
Approach: They propose to use BLEU and METEOR to evaluate machine translations against human translations.
Outcome: VisEval Metric Viewer (VEMV) provides visualisation of multiple evaluation scores so they can be easily interpreted by a user.
Learning to Verify Summary Facts with Fine-Grained LLM Feedback (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have significantly enhanced the text summarization performance, but hallucination issues still occur in summaries.
Approach: They propose a large-scale dataset containing fine-grained factual feedback on summaries that can be fine tuned by using Large Language Models (LLMs) they employ 10 distinct LLMs for diverse summary generation and Llama-3-70B-Instruct for feedback.
Outcome: The proposed model outperforms models trained on smaller human-annotated datasets while maintaining high performance.
Linguistic Constructs Represent the Domain Model in Intelligent Language Tutoring (2023.eacl-demo)

Copied to clipboard

Challenge: a new language-learning platform, Revita, is being developed for language learners . the platform uses a system of linguistic constructs to represent domain knowledge .
Approach: They propose to use a domain model to represent the domain knowledge of Revita's online tutoring system.
Outcome: The proposed language-learning platform, Revita, is based on the domain model of linguistic constructs . the system is undergoing pilot use with hundreds of students at several universities .
Effects of Gender Stereotypes on Trust and Likability in Spoken Human-Robot Interaction (L18-1)

Copied to clipboard

Challenge: a study investigates the influence of gender stereotypes on trust and likability of humanoid robots . explicit gender and stereotypicality of a task are manipulated to influence robot behavior . future research may look into situational variables that drive stereotypification in robot interaction .
Approach: They investigated the influence of gender stereotypes on trust and likability of robots . they used explicit (name and voice) and implicit (personality) genders to manipulate stereotypical tasks . future research may look into situational variables that drive stereotypization .
Outcome: The findings suggest that gender stereotypes need to be differentiated in robot interaction . the gender and personality characteristics of robots influence trust and likability .
PEEP-Talk: A Situational Dialogue-based Chatbot for English Education (2023.acl-demo)

Copied to clipboard

Challenge: Existing chatbots lack realistic practice scenarios for English learners . existing platforms employ hand-crafted and patternmatching rules, limiting communication ability and responding appropriately to out-of-situation utterances.
Approach: They propose a real-world situational dialogue-based chatbot for English education . it generates appropriate responses in various real-life situations while providing accurate feedback .
Outcome: The proposed chatbot generates appropriate responses in various real-life situations while providing accurate feedback to learners.
Journalist-in-the-Loop: Continuous Learning as a Service for Rumour Analysis (D19-3)

Copied to clipboard

Challenge: Existing rumour analysis tools do not scale due to the large volume and velocity of user generated content.
Approach: They propose to use a web-based rumour analysis tool that can continuously learn from journalists and integrate it into existing tools and platforms.
Outcome: The proposed system can be easily integrated as a service into existing tools and platforms used by journalists using a REST API.
GREEN: Generative Radiology Report Evaluation and Error Notation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing automated evaluation metrics fail to consider factual correctness or are limited in their interpretability.
Approach: They propose a radiology report evaluation metric that leverages natural language understanding of language models to identify and explain clinically significant errors.
Outcome: The proposed method demonstrates higher correlation with expert error counts and higher alignment with expert preferences when compared to previous methods.
A dataset for identifying actionable feedback in collaborative software development (P18-2)

Copied to clipboard

Challenge: a dataset of code reviews for the Google Chromium project analyzed linguistic features of code review feedback that elicited responsive actions from coworkers.
Approach: They analyze code reviews for Google Chromium and extract linguistic features that elicit responsive responses from coworkers.
Outcome: The proposed dataset shows that using NLP can be useful in code reviews . it also shows that it can be used to improve code reviews in a collaborative environment .
DAC: Decomposed Automation Correction for Text-to-SQL (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve text-to-SQL performance are hard to detect errors in SQL directly.
Approach: They propose to use decomposed correction to improve text-to-SQL performance . they first detect errors based on decompose subtasks, then use it to correct them .
Outcome: The proposed method improves text-to-SQL performance by 1.4% compared with previous methods .
FEAT-writing: An Interactive Training System for Argumentative Writing (2025.coling-demos)

Copied to clipboard

Challenge: Argumentative writing is a critical skill for academic success, but many students struggle to develop these skills.
Approach: They developed an online system that provides students with automated feedback and exercises for argumentative writing.
Outcome: The proposed system improves argumentative writing quality among native English speakers and english-as-a-foreign-language learners.
Learning to repair: Repairing model output errors after deployment using a dynamic memory of feedback (2022.findings-naacl)

Copied to clipboard

Challenge: Our approach pairs an LM with a growing memory of cases where the user identified an output error and provided general feedback on how to correct it.
Approach: They propose to use an existing script generator to train a model to repair output errors without retraining.
Outcome: The proposed model learns to apply user feedback to repair output errors while avoiding similar past mistakes on new, unseen examples.
T3-Vis: visual analytic for Training and fine-Tuning Transformers in NLP (2021.emnlp-demo)

Copied to clipboard

Challenge: Existing visual analytics tools have been shown to support the analysis and interpretation of deep learning models due to the inherent black-box nature of the models.
Approach: They propose to use visual analytic framework to help researchers understand the model's intrinsic properties and behaviours through interactive visualization.
Outcome: The proposed framework provides valuable insights about the model’s intrinsic properties and behaviours through interactive visualization and a suite of built-in algorithms.
Multimodal large language models for inclusive collaboration learning tasks (2022.naacl-srw)

Copied to clipboard

Challenge: This project leverages advances in multimodal large language models to build an inclusive collaboration feedback loop for participants developing general collaboration skills.
Approach: They propose to integrate advances in multimodal large language models into downstream tasks such as the learning analytics feedback loop.
Outcome: The proposed model will be used to detect, model, and feedback participants developing general collaboration skills.
Zero-Shot Sequence Labeling: Transferring Knowledge from Sentences to Tokens (N18-1)

Copied to clipboard

Challenge: Recent work has used attention weights to visualize the focus of neural models in input data.
Approach: They propose to use attention-based visualization techniques to infer token-level labels from a network trained only on sentence-level labeling.
Outcome: The proposed approach outperforms gradient-based methods on four datasets and is expected to outperfect supervised methods.
Continually Improving Extractive QA via Human Feedback (2023.emnlp-main)

Copied to clipboard

Challenge: a study of extractive question answering systems using human feedback shows promising potential for continual learning.
Approach: They study extractive question answering system by using user feedback to improve it . they design and deploy an iterative approach where users ask questions and provide feedback .
Outcome: The proposed model improves over time across different data regimes and domains . human user feedback is more affordable and abundant than annotations provided by trained experts .
What Do Users Care About? Detecting Actionable Insights from User Feedback (2022.naacl-industry)

Copied to clipboard

Challenge: a large amount of data can be used to extract actionable insights from user feedback . however, the data is unstructured and voluminous, and is underutilized for most users .
Approach: They propose an unsupervised method for finding actionable insights from user feedback . they cluster data into groups containing coherent insights, followed by theme detection .
Outcome: The proposed approach outperforms baselines on two real-world user feedback datasets and one academic dataset.
Contextualizing Argument Quality Assessment with Relevant Knowledge (2024.naacl-short)

Copied to clipboard

Challenge: Existing methods for assessing argument quality in isolation analyze their quality in the absence of context, which affects their accuracy and generalizability.
Approach: They propose a method for scoring argument quality based on contextualization via relevant knowledge that leverages large language models to provide feedback, infer hidden assumptions, supply a similar-quality argument, or give a counter-argument.
Outcome: The proposed method outperforms existing methods across multiple metrics in both in-domain and zero-shot setups.
Stimulating Creativity with FunLines: A Case Study of Humor Generation in Headlines (2020.acl-demos)

Copied to clipboard

Challenge: FunLines is an online game that allows players to generate and rate funny news headlines . it is difficult to generate data that depends on human creativity, and measuring creativity often requires more effort.
Approach: They propose a game where players edit news headlines to make them funny and rate the funniness of headlines edited by others.
Outcome: The proposed game outperforms other crowdsourcing approaches in generating humor datasets.
Dialogue Response Ranking Training with Large-Scale Human Feedback Data (2020.emnlp-main)

Copied to clipboard

Challenge: Existing open-domain dialog models can minimize the perplexity of target human responses . however, some human responses are more engaging than others, spawning more followup interactions .
Approach: They train open-domain dialog models to minimize perplexity of target human responses . they use social media feedback data to train models to predict engaging dialog turns .
Outcome: The proposed model outperforms existing models on 133M human feedback pairs . it also outperformed the conventional dialog perplexity baseline model .
LCDS: A Logic-Controlled Discharge Summary Generation System Supporting Source Attribution and Expert Review (2025.acl-demo)

Copied to clipboard

Challenge: Large language models (LLMs) are capable of generating inaccurate discharge summary content or fabricating information without valid sources.
Approach: They propose a tool for empowering LLMs with Logic-Controlled Discharge Summary generation.
Outcome: The proposed tool identifies the writing logic of discharge summaries and integrates it with EMRs to generate silver discharge summararies.
Self-Regulated Interactive Sequence-to-Sequence Learning (P19-1)

Copied to clipboard

Challenge: Existing approaches to learning from different types of feedback have not been explored.
Approach: They propose a machine learning algorithm that uses self-regulation to balance cost and effect of different types of feedback.
Outcome: The proposed model is robust under domain shift and is a promising alternative to active learning.
ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models (2026.acl-demo)

Copied to clipboard

Challenge: Existing code sandboxes fail to provide accurate verification and efficiency under high-concurrency workloads.
Approach: They propose a high-fidelity code verification system that provides sandbox feedback for RL training and evaluation.
Outcome: The proposed system outperforms heuristic-matching baselines on LiveCodeBench and training stability on high-concurrency workloads.
Reciprocal Learning of Knowledge Retriever and Response Ranker for Knowledge-Grounded Conversations (2022.coling-1)

Copied to clipboard

Challenge: Recent work on grounding dialogue agents with knowledge documents has sparked increased attention . hand-labeling data to that end is time-consuming and many datasets lack knowledge annotations .
Approach: They propose a reciprocal learning approach to optimize a knowledge retriever and a response ranker for knowledge-grounded response retrieval without ground-truth knowledge labels.
Outcome: The proposed model outperforms previous state-of-the-art methods on two public benchmarks.
Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Utility-based retrieval has emerged as a promising topic for downstream tasks . however, capturing passage utility accurately remains unexplored due to insufficient understanding .
Approach: They propose a framework for training utility-based retrievers in Retrieval-Augmented Language Models . it incorporates multi-task generalization and inter-passage interaction to improve performance .
Outcome: The proposed framework improves performance on ten datasets across different tasks.
ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQL (2026.acl-long)

Copied to clipboard

Challenge: Current Text-to-SQL reasoning models lack integrated execution feedback during generation.
Approach: They propose a text-to-SQL framework that interacts with the SQL execution engine during decoding and dynamically adjusts reasoning based on execution feedback.
Outcome: The proposed framework achieves 89.1% accuracy on Spider and 65.3% on BIRD at the 7B scale.
SELFGOAL: Your Language Agents Already Know How to Achieve High-level Goals (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to improve the performance of language agents without training are not available.
Approach: They propose an automatic approach to break down high-level goals into tree structure of more practical subgoals during interaction with environments while identifying the most useful subgoal.
Outcome: The proposed approach significantly improves the performance of language agents across various tasks, including competitive, cooperative, and deferred feedback environments.
RevieWeaver: Weaving Together Review Insights by Leveraging LLMs and Semantic Similarity (2025.naacl-industry)

Copied to clipboard

Challenge: RevieWeaver extracts key product features and provides concise review summaries . a condensed list of key features, pros, and cons, along with a brief summary of customer opinions can help mitigate this issue.
Approach: They propose a framework that extracts key product features and provides concise review summaries.
Outcome: The proposed framework scales efficiently to 30 million reviews and ensures reproducibility and controllability.
LEAF: Language Learners’ English Essays and Feedback Corpus (2024.naacl-short)

Copied to clipboard

Challenge: Current automated essay scoring models lack the granularity desired by learners and instructors seeking more detailed insights.
Approach: They present a corpus of English essays and their corresponding feedback from the “essayforum” website.
Outcome: The LEAF corpus provides valuable feedback for students and teachers . it provides insights on argumentative aspects and organizational coherence .
Efficient and Accurate Prompt Optimization: the Benefit of Memory in Exemplar-Guided Reflection (2025.acl-long)

Copied to clipboard

Challenge: Recent work utilizes feedbacks generated from erroneous cases to guide prompt optimization . previous methods rely on computational resources and powerful GPUs .
Approach: They propose an automatic prompt engineering method that leverages feedbacks from erroneous cases to guide prompt optimization.
Outcome: The proposed method surpasses state-of-the-art methods with less steps and lower computational resources.
PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers in Overleaf (2026.acl-demo)

Copied to clipboard

Challenge: Emerging AI-powered writing assistants focus on grammar fixes or simulating peer review with final scores, yet they fall short of providing concrete, actionable suggestions that help students improve their papers during drafting.
Approach: They propose a human-centered writing assistant system that delivers actionable suggestions as Overleaf-native inline comments while leaving the actual writing entirely to human authors.
Outcome: The proposed system outperforms a baseline with the skill library and provides actionable suggestions while leaving the actual writing to human authors.
Personalized Abstractive Summarization by Tri-agent Generation Pipeline (2024.findings-eacl)

Copied to clipboard

Challenge: Existing research shows that large language models do not consistently satisfy users' preferences or expectations.
Approach: They propose a tri-agent generation pipeline that includes a generator, an instructor, and an editor to enhance output personalization.
Outcome: The proposed pipeline generates outputs that better meet user expectations on two abstractive summarization datasets.
BackMATH: Towards Backward Reasoning for Solving Math Problems Step by Step (2025.coling-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive results in reasoning, but when faced with more complex mathematical problems, performance drops significantly.
Approach: They propose a backward reasoning dataset that includes 14K backward thinking problems and 100K reasoning steps.
Outcome: The proposed model achieves an accuracy of 68.1% on the GSM8K dataset and 21.9% on the MATH dataset, exceeding the SOTA by 1.6% and 2.1% respectively.
Counterfactual Augmentation for Multimodal Learning Under Presentation Bias (2023.findings-emnlp)

Copied to clipboard

Challenge: In real-world machine learning systems, labels are often derived from user behaviors that the system wishes to encourage.
Approach: They propose a method for correcting presentation bias using generated counterfactual labels by augmentation of the labels by the user.
Outcome: The proposed method improves performance in an oracle setting compared to uncorrected models and existing bias-correction methods.
Which Spurious Correlations Impact Reasoning in NLI Models? A Visual Interactive Diagnosis through Data-Constrained Counterfactuals (2023.acl-demo)

Copied to clipboard

Challenge: a spurious correlation exists when a feature correlates with the target label while there is no causal relationship between the feature and the label.
Approach: They propose a dashboard that allows users to generate diverse and challenging examples by drawing inspiration from GPT-3 suggestions.
Outcome: The proposed dashboard enables users to generate diverse and challenging examples by drawing inspiration from GPT-3 suggestions and make refinements based on the feedback.
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing systems that provide personalised, curriculum-aligned feedback are time-intensive and time-consuming.
Approach: They propose a modular, LLM-based system that generates personalised, curriculum-aligned feedback in science education.
Outcome: The proposed system generates personalised, curriculum-aligned feedback in science education.
An Individualized News Affective Response Dataset (2024.acl-srw)

Copied to clipboard

Challenge: a new dataset captures subjective affective responses to news headlines . current methods to assess emotion detection ignore subjective differences in groups and individuals .
Approach: They propose a large-scale dataset capturing subjective affective responses to news headlines . the dataset includes Facebook post screenshots from popular UK media outlets .
Outcome: The proposed dataset captures subjective affective responses to headlines from popular media outlets.
IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction (2024.acl-long)

Copied to clipboard

Challenge: Various communication frameworks assist individuals in conducting difficult conversations by providing a set of skills to apply.
Approach: They propose a language-based simulation system that provides just-in-time feedback to support the practice and learning of interpersonal effectiveness skills.
Outcome: The proposed training system improves self-efficacy and reduces negative emotions by 27% compared to the GPT-4 training system.
CodeGuard: Improving LLM Guardrails in CS Education (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly embedded in Computer Science classrooms to automate code generation, feedback, and assessment.
Approach: They propose a guardrail framework for educational AI systems that can handle unsafe and irrelevant prompts.
Outcome: The proposed framework reduces potentially harmful or policy-violating code completions by 30-65% without degrading performance on legitimate educational tasks.
Learning from Chunk-based Feedback in Neural Machine Translation (P18-2)

Copied to clipboard

Challenge: a common problem with explicit ratings of translations is that users are not qualified enough to provide reliable feedback for the whole sentence.
Approach: They propose a way to learn from partial feedback in neural machine translation . they ask users to highlight a correct chunk of a translation based on partial feedback .
Outcome: The proposed method outperforms sentence-based feedback by 2.61% BLEU absolute.
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to synthesize test cases using Large Language Models (LLMs) rely on the model’s intrinsic generation capabilities without external feedback, resulting in insufficiently diverse cases.
Approach: They propose a feedback-driven iterative framework that leverages Large Language Models to generate initial test cases, execute them against known correct and incorrect solutions, and utilizes the failed results as feedback to guide the LLM in refining the test cases toward high fidelity and discriminability.
Outcome: The proposed method outperforms the existing codecontests and codecontests+ models by 4.30% and 8.78%.
Faithfulness Beyond Plausibility: Auditing Human Explanations in Educational Assessment (2026.acl-srw)

Copied to clipboard

Challenge: a gap exists between explanation components and how scores are constructed, and whether they reflect how scores were constructed . authors: explanation components are structurally inconsistent and may not be used as post-hoc justifications .
Approach: They propose to use human tutor grading traces to test whether explanations are reliable . they find that removing rubric-level information leads to substantial changes in reconstructed scores .
Outcome: The proposed diagnostic measures how explanation components contribute to score interpretation . removing rubric-level information leads to substantial changes in reconstructed scores .
AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching (2023.acl-industry)

Copied to clipboard

Challenge: In recent years, the utilization of Artificial Intelligence (AI) in the contact center industry is on the rise.
Approach: They present a transformer-based pairwise sentence classification model that analyzes call transcripts to determine which calls are most relevant for coaching purposes.
Outcome: The proposed model can determine which calls are most relevant for coaching purposes based on quality assurance queries/questions asked by managers or supervisors .
Give Me More Feedback: Annotating Argument Persuasiveness and Related Attributes in Student Essays (P18-1)

Copied to clipboard

Challenge: Existing work on automated essay scoring has focused on holistic scoring, which summarizes the quality of an essay with a single score.
Approach: They present a corpus of essays simultaneously annotated with argument components, argument persuasiveness scores, and attributes of argument components that impact an argument’s persuasiveness.
Outcome: The proposed corpus could trigger the development of novel computational models that provide useful feedback to students on why their arguments are (un)persuasive .
A Practical Incremental Learning Framework For Sparse Entity Extraction (C18-1)

Copied to clipboard

Challenge: Existing approaches to extract entities from textual data are expensive and unattractive due to the high cost of training.
Approach: They propose a framework that integrates Entity Set Expansion and Active Learning to reduce the cost of data annotation.
Outcome: The proposed framework reduces the cost of sparse entity annotation by 85% and 45% while maintaining high accuracy.
PromptSculptor: Multi-Agent Based Text-to-Image Prompt Optimization (2025.emnlp-demos)

Copied to clipboard

Challenge: PromptSculptor automates the iterative prompt optimization process for Text-to-Image models . previous work focused on generating detailed, high-quality prompts based on user feedback .
Approach: They propose a framework that decomposes a task into four specialized agents . they use Chain-of-Thought reasoning to transform a short, vague user prompt into a comprehensive, refined prompt.
Outcome: The proposed framework significantly improves output quality and reduces iterations needed for user satisfaction.
Semi-Supervised Dialogue Policy Learning via Stochastic Reward Estimation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for dialogue policy optimization do not provide sufficient supervision signals at the end of dialogues.
Approach: They propose to learn from state-action pairs of an optimal policy to provide turn-by-turn rewards.
Outcome: The proposed approach outperforms competitive policy learning baselines on a benchmark multi-domain dataset.
Accurate Training of Web-based Question Answering Systems with Feedback from Ranked Users (2023.acl-industry)

Copied to clipboard

Challenge: Recent work shows that large-scale annotated datasets are essential for training state-of-the-art Question Answering (QA) models.
Approach: They use large-scale annotated datasets to train question answering models . they use feedback data collected from deployed QA systems to provide cheaper supervision .
Outcome: The proposed model improves on the large scale annotated datasets from QA systems . the proposed model can be easily supervised on large-scale unlabeled web data .
DUAL-REFLECT: Enhancing Large Language Models for Reflective Translation through Dual Learning Feedback Mechanisms (2024.acl-short)

Copied to clipboard

Challenge: Existing self-reflection methods lack effective feedback information, limiting the translation performance of large language models (LLMs).
Approach: They propose a framework that leverages the dual learning of translation tasks to provide effective feedback, thereby enhancing the models’ self-reflective abilities and improving translation performance.
Outcome: The proposed framework improves the models’ self-reflective abilities and improves translation accuracy and eliminating ambiguities across translation tasks.
Boot and Switch: Alternating Distillation for Zero-Shot Dense Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to enhance dense retrieval models are unwieldy, such as requiring explicit supervision, complex model architectures, or massive external models.
Approach: They propose an unsupervised method to enhance passage retrieval in zero-shot settings by iterating a loop that a dense retriever learns from supervision signals provided by a reranker.
Outcome: The proposed method outperforms leading supervised and unsupervised retrievers on the BEIR benchmark while showing strong adaptation abilities to tasks and domains that were unseen during training.
Metamo: Empowering Large Language Models with Psychological Distortion Detection for Cognition-aware Coaching (2025.emnlp-demos)

Copied to clipboard

Challenge: Metamo is a browser-based dialogue system that transforms an off-the-shelf large language model into an empathetic coach for everyday workplace concerns.
Approach: They propose a browser-based dialogue system that first identifies the cognitive distortion behind an emotion, then recognizes the user’s emotion, and finally produces a question-centered reply that invites reflection.
Outcome: Empirical tests on public corpora showed that the proposed system improved emotionrecognition quality and response diversity without sacrificing latency.
The lack of theory is painful: Modeling Harshness in Peer Review Comments (2022.aacl-main)

Copied to clipboard

Challenge: a new study shows that peer-review has a power imbalance, making it fraught for authors . authors argue that a little more effort to remain critical but be constructive would help foster a positive outcome .
Approach: They propose to use a dataset to show peer-review comments' harshness scores . they argue that this moderation could help authors to be more constructive .
Outcome: The proposed dataset shows that it can be used to make peer reviews less hurtful and more welcoming.
MARS: Multilingual Aspect-centric Review Summarisation (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for summarizing customer feedback are not able to extract actionable reviews into a specific target language.
Approach: They propose a framework involving extract-then-summarise to summariser customer feedback into a specific language.
Outcome: The proposed framework improves abstractive baselines and efficiency to real-time systems.
SeRTS: Self-Rewarding Tree Search for Biomedical Retrieval-Augmented Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing retrieval-augmented approaches to large language models face performance limitations due to the lack of publicly available training data.
Approach: They propose a plug-and-play LLM-based retrieval method called Self-Rewarding Tree Search based on Monte Carlo Tree Search and a self-rewarding paradigm to address these limitations.
Outcome: The proposed method improves the performance of the BM25 retriever and surpasses the baseline of self-reflection in both efficiency and scalability.
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem (2025.coling-main)

Copied to clipboard

Challenge: Natural language explanations (NLEs) are vital for elucidating the reasoning behind large language model (LLM) decisions.
Approach: They propose a role-modeling approach that employs two LLMs as generator and critic to generate and refine NLEs.
Outcome: The proposed model outperforms self-refine and can perform with less powerful LLMs.
ExpanRL: Hierarchical Reinforcement Learning for Course Concept Expansion in MOOCs (2020.aacl-main)

Copied to clipboard

Challenge: Existing methods for concept expansion in MOOCs are inefficient because of the diversity of MOOC courses and rapid updates.
Approach: They propose an end-to-end hierarchical reinforcement learning (HRL) model for concept expansion in MOOCs that employs a two-level mechanism of seed selection and concept expansion.
Outcome: The proposed model improves on nine real MOOC datasets and maintains competitive performance under different settings.
Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction (2025.acl-srw)

Copied to clipboard

Challenge: Existing evaluation metrics for Grammatical error correction lack explainability . lack of explainability hinders researchers from analyzing strengths and weaknesses of models .
Approach: They propose to assign sentence-level scores to individual edits to improve GEC performance . they use Shapley values, from cooperative game theory, to compute contribution of each edit .
Outcome: The proposed method shows that the evaluation metrics are consistent across edits and human evaluations.
MixRevDetect: Towards Detecting AI-Generated Content in Hybrid Peer Reviews. (2025.naacl-short)

Copied to clipboard

Challenge: Existing methods for detecting fully AI-generated peer reviews fail to detect finer-grained AI-generated points within mixed-authorship reviews.
Approach: They propose a method to identify AI-generated points in peer reviews using large language models . their approach achieved an F1 score of 88.86%, significantly outperforming existing methods .
Outcome: The proposed method outperforms existing methods in identifying AI-generated points in peer reviews.
Dialogue Act Annotation in a Multimodal Corpus of First Encounter Dialogues (2020.lrec-1)

Copied to clipboard

Challenge: a method used to annotate dialogue acts in a multimodal corpus is described . the annotations allow for analysis of how multimodal signals contribute to the structure and content of the dialogues.
Approach: They propose to annotate dialogue acts in a multimodal corpus of first encounter dialogues . they focus on which dialogue acts often follow each other across speakers and which overlap gestural behaviour .
Outcome: The method used to annotate dialogue acts in a multimodal corpus is described.
Unveiling the Achilles’ Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have highlighted various neural metrics that align well with human evaluations.
Approach: They propose a black-box adversarial framework that generates strong disagreements between human and victim evaluators.
Outcome: The proposed framework can significantly improve the performance of human and victim evaluators.
Real-time Factuality Assessment from Adversarial Feedback (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluations for assessing the factuality of news from conventional sources, such as claims on fact-checking websites, result in high accuracies over time for LLM-based detectors.
Approach: They propose a pipeline that leverages natural language feedback from a RAG-based detector to iteratively modify real-time news into deceptive variants that challenge LLMs.
Outcome: The proposed pipeline reduces the binary classification ROC-AUC by 17.5 percent for a strong RAG-based GPT-4o detector.
Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for grounding large language models suffer from inefficient querying . Existing approaches that rely on physical verification or self-reflection suffer from excessive querying.
Approach: They propose a framework that introduces Reinforced Advantage feedback for efficient self-refinement of plans.
Outcome: The proposed framework surpasses baselines in success rate and significantly decreases interaction steps of agents and query rounds of LLMs.
MOOSE-Copilot: A Web-Based Interactive Assistant for Unified Exploratory and Fine-Grained Scientific Hypothesis Discovery (2026.acl-demo)

Copied to clipboard

Challenge: Existing approaches to scientific hypothesis discovery operate autonomously with little to no human guidance.
Approach: They propose a framework that empowers scientists to steer the generative process via explicit signals.
Outcome: The proposed framework outperforms existing models in terms of quality and accuracy of the input signals.
Towards a Language for Natural Language Treebank Transductions (C18-1)

Copied to clipboard

Challenge: a transduction language suitable for natural language treebank transformations is described . linguistics and computer scientists have always liked to use trees to theorize about the structure of natural language sentences.
Approach: They propose a transduction language suitable for treebank transformations . they aim to get feedback from the NLP community to eventually converge to a standard .
Outcome: The proposed language is suitable for treebank transformations and motivates its implementation.
Explanation-Based Human Debugging of NLP Models: A Survey (2021.tacl-1)

Copied to clipboard

Challenge: In this paper, we review literature that exploits explanations to enable humans to fix bugs in NLP models.
Approach: They review literature that exploits explanations to enable humans to fix bugs in NLP models.
Outcome: The proposed approach is described in detail in this paper and is based on three dimensions of the problem explanation-based human debugging (EBHD).
Reinforcement Learning for Adversarial Query Generation to Enhance Relevance in Cold-Start Product Search (2025.acl-industry)

Copied to clipboard

Challenge: Existing methods do not incorporate feedback from the query relevance model, limiting their ability to generate queries that enhance product retrieval.
Approach: They propose an adversarial reinforcement learning framework that exposes weaknesses in query classification models by creating synthetic queries that augment the classifier's training set.
Outcome: The proposed framework improves query generation performance on public datasets and on proprietary datasets.
Feedback Attribution for Counterfactual Bandit Learning in Multi-Domain Spoken Language Understanding (2021.emnlp-main)

Copied to clipboard

Challenge: a large amount of labeled data is needed for fine-tuning.
Approach: They propose attribution methods inspired by multi-agent reinforcement learning for a feedback attribution problem in spoken language understanding.
Outcome: The proposed methods can train competitive models from user feedback.
Binary Classifier Optimization for Large Language Model Alignment (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for aligning large language models rely on preference-based approaches that require both positive and negative feedback as a pair.
Approach: They propose a binary classifier optimization technique that trains a classifier using only binary feedback and a reward shift technique which minimizes the DPO loss.
Outcome: The proposed method performs on a paired preference dataset and on 'likert-5 scale annotation dataset' it consistently demonstrates effective and robust alignment across four base LLMs and three different datasets, showcasing the strength of the proposed technique.
I-SEE: An Instruction-tuned, SOP-Enhanced Quality Evaluator for Product Content (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing approaches to content evaluation treat information uniformly without prioritizing based on customer relevance.
Approach: They propose a framework that combines domain expertise with a single instruction to improve content.
Outcome: a new framework outperforms existing models in detecting inconsistencies across 20 product categories and 150 product specific features.
What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks? (2021.acl-long)

Copied to clipboard

Challenge: Despite the importance of datasets for natural language understanding, there has been little attention on crowdsourcing methods for collecting datasets.
Approach: They compare the effectiveness of crowdsourcing methods for boosting NLU example difficulty with training crowdworkers instead of expert judgments.
Outcome: The proposed method is ineffective for boosting NLU example difficulty, but it is not effective for training crowdworkers and qualifying workers based on expert judgments.
Towards Conversational Recommendation over Multi-Type Dialogs (2020.acl-main)

Copied to clipboard

Challenge: In recent years, there has been a significant increase in the work of conversational recommendation due to the rise of voice-based bots.
Approach: They use a Chinese dialog dataset DuRecDial to study conversational recommendation in the context of multi-type dialogs where bots can proactively lead a conversation from a non-recommendation dialog to a recommendation dialog.
Outcome: The proposed dataset allows to investigate different parts of the overall problem, e.g., how to naturally lead a dialog, how interact with users for recommendation.
Annotating Customer-Oriented Behaviour in Call Centre Sales Dialogues (2024.lrec-main)

Copied to clipboard

Challenge: Customer-oriented behaviour (COB) is often hindered by a lack of clarity in its definition and lack of robust analytical, categorization, and computational approaches.
Approach: They propose a conceptual and empirical framework for customer-oriented behaviour in call centre interactions . they aim to identify facets of COB that positively impact on Customer Satisfaction .
Outcome: The proposed framework improves our understanding of the dynamics shaping sales strategies in call centres and holds promise for practical applications in optimising customer-agent interactions.
Reinforcement Learning with Token-level Feedback for Controllable Text Generation (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for controllable text generation are guided by coarse-grained feedback, which may lead to suboptimal performance owing to semantic twists or progressions within sentences.
Approach: They propose a reinforcement learning algorithm which formulates TOken-LEvel rewards for controllable text generation and employs a "first-quantize-then-noise" paradigm to enhance the robustness of the RL algorithm.
Outcome: The proposed algorithm can achieve superior performance on single-attribute and multi-attract control tasks.
Similarity-Based Content Scoring - A more Classroom-Suitable Alternative to Instance-Based Scoring? (2023.findings-acl)

Copied to clipboard

Challenge: Recent work suggests that similarity-based content scoring methods can yield comparable results to instance-based supervised learning.
Approach: They propose to use similarity-based scoring to achieve similar results . they compare different instance-based and similarity based methods on multiple data sets .
Outcome: The proposed approach has a lower need for annotated training data and better zero-shot performance, but the results are not consistent with previous studies.
ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages (2024.acl-long)

Copied to clipboard

Challenge: Existing research focuses on enhancing LLMs capabilities through tool utilization.
Approach: They propose a framework to investigate safety issues in large language models in tool learning . they propose malicious queries and jailbreak attacks in the input stage .
Outcome: The proposed framework investigates six safety scenarios for LLMs in tool learning . the data will be released upon acceptance of the proposed framework .
PUPPET: Neural-Symbolic Standardized Patients for Mental Health (2026.acl-long)

Copied to clipboard

Challenge: Existing LLM-based training approaches lack faithful responses to clinical errors and explainable feedback.
Approach: They propose a neural-symbolic virtual standardized patient governed by an OBSERVE-THINK-BEHAVE architecture that embeds LLM reasoning into a symbolic system where experts implant causal associations between intervention logic and patient mental states.
Outcome: The proposed model outperforms baselines in faithfulness and pedagogical value.
Advancing Process Verification for Large Language Models via Tree-Based Preference Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for generating step-by-step rationales fail to fully utilize the relative merits of intermediate steps, limiting the effectiveness of feedback provided.
Approach: They propose a tree-based preference learning verifier that constructs reasoning trees via a best-first search algorithm and collects step-level paired data for preference training.
Outcome: The proposed approach outperforms existing benchmarks on arithmetic and commonsense reasoning tasks.
From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications (2024.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for large language models yield numerical scores that ignore user experience.
Approach: They propose a metric that suggests revision edits that mimic the human writing process . their results show that the metric offers more insightful feedback and distinguishes between texts .
Outcome: The proposed metric can provide a self-explained text evaluation result in a human-understandable manner beyond the context-independent score.
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study shows that asking for direct user feedback can be disruptive . we examine whether incorporating the contents of user feedback improves model performance .
Approach: They analyze user feedback in the user-LLM conversation logs and harvest learning signals from it.
Outcome: The proposed approach can lead to model degradation on two user-LM interaction datasets.
Iterative Refinement of Project-Level Code Context for Precise Code Generation with Compiler Feedback (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) generate code for given contexts, such as incomplete code, class, data structure, or project-specific information.
Approach: They propose a compiler feedback-based code generation approach that leverages static analysis to identify mismatches between the generated code and the project's context.
Outcome: The proposed model outperforms retrieval-based code generation baselines and significantly outperfies the existing large language models.
Audience-Centric Natural Language Generation via Style Infusion (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to text style transfer (TST) with large volumes of parallel or non-parallel data are limiting for two reasons: it is difficult to collect large volumes and some stylistic objectives are hard to define without audience feedback.
Approach: They propose a task of style infusion - infusing stylistic preferences of audiences into pretrained language generation models by leveraging pairwise human judgments to bootstrap a style analysis model and augment a seed set of judgments.
Outcome: The proposed method generates compelling stylized examples with generic text prompts while balancing fluency and style adoption.
What’s Wrong? Refining Meeting Summaries with LLM Feedback (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for meeting summarization are limited and lack the robustness and context-based accuracy needed to maintain relevance.
Approach: They propose a multi-LLM correction approach for meeting summarization using a two-phase process that mimics the human review process: mistake identification and summary refinement.
Outcome: The proposed approach improves the quality of a given meeting summarization measured by relevance, informativeness, conciseness, and coherence.
Multi-Agent Comedy Club: Investigating Community Discussion Effects on LLM Humor Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on the use of multi-turn interaction and feedback for LLM writing focus on prompts and localized feedback.
Approach: They build a controlled multi-agent sandbox that instantiates a small standup comedy community and allows it to manipu-late whether public reception is generated, logged, and fed back into later rounds.
Outcome: The proposed model improves craft/clarity and social response with occasional increases in aggressive humor.
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values (2023.emnlp-main)

Copied to clipboard

Challenge: Incorporating human feedback into Large Language Models is a welcome development, but it introduces new biases and challenges.
Approach: They propose to survey 95 articles that use human feedback to steer, guide or tailor the behaviours of large language models.
Outcome: The proposed approaches are based on 95 articles primarily from the ACL and arXiv repositories and highlight five unresolved conceptual and practical challenges.
Estimating Summary Quality with Pairwise Preferences (N18-1)

Copied to clipboard

Challenge: Existing evaluation systems rely on gold standard summaries but they are expensive and require the availability of experts to achieve high quality.
Approach: They propose an alternative evaluation approach based on pairwise preferences of sentences to provide useful feedback in the form of pairwise preference.
Outcome: The proposed evaluation framework performs better than the three most popular versions of ROUGE with less expensive human input.
DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent (2025.findings-emnlp)

Copied to clipboard

Challenge: a new method for detecting advanced backdoors is proposed to bypass safety audits.
Approach: They propose a backdoor implantation strategy that introduces dynamic encryption to bypass safety audits.
Outcome: The proposed method achieves an attack success rate approaching 100% while maintaining a detection rate of 0%.
Presentations by the Humans and For the Humans: Harnessing LLMs for Generating Persona-Aware Slides from Documents (2024.eacl-long)

Copied to clipboard

Challenge: Existing efforts to automate document-to-slide generation have failed to adapt to the persona of target audience or duration of presentation.
Approach: They propose a concept of end-user specification-aware document to slides conversion that incorporates end- user specifications into the conversion process.
Outcome: The proposed model can create persona-aware presentations tailored to the persona of target audience and cognitive abilities of target audiences.
Alignment-Enhanced Decoding: Defending Jailbreaks via Token-Level Adaptive Refining of Probability Distributions (2024.emnlp-main)

Copied to clipboard

Challenge: Existing defenses against jailbreaks focus on perturbing or inspecting inputs, but ignore competing objectives, the underlying cause of alignment failures.
Approach: They propose a novel defense that employs adaptive decoding to address the root causes of jailbreak issues.
Outcome: The proposed defense improves safety alignment while maintaining helpfulness.
Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in diffusion models have shown impressive performance in many domains, but their ability to follow instructions is still unsatisfactory.
Approach: They propose an algorithm that aligns images to text through iterative image sampling and prompt relabeling with feedback.
Outcome: The proposed algorithm improves on the spatial relation VISOR benchmark by 15.22% compared to previous methods.
LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models struggle to detect lazy thinking in a zero-shot setting, but instruction-based fine-tuning significantly boosts performance by 10-20 performance points.
Approach: They propose to use LazyReview to train junior reviewers in the community to detect lazy thinking in peer-review sentences annotated with fine-grained lazy thinking categories.
Outcome: The proposed dataset shows that LLMs struggle to detect lazy thinking instances in a zero-shot setting, while instruction-based fine-tuning significantly boosts performance by 10-20 performance points.
Reliability and Learnability of Human Bandit Feedback for Sequence-to-Sequence Reinforcement Learning (P18-1)

Copied to clipboard

Challenge: Recent work has shown that reinforcement learning (RL) can be scaled to games with large state-action spaces, achieving human-level performance or even superhuman performance.
Approach: They propose to use bandit feedback to improve sequence-to-sequence learning by simulating reward signals by evaluation metrics such as BLEU, F1-score, or ROUGE.
Outcome: The proposed methods improve performance even from small amounts of human feedback, pointing to a great potential for applications at larger scale.
Personalized Multimodal Feedback Generation in Education (2020.coling-main)

Copied to clipboard

Challenge: In this paper, we propose a novel Personalized Multimodal Feedback Generation Network (PMFGN) that generates personalized feedback for teachers to evaluate assignments involving multimodal inputs.
Approach: They propose a Personalized Multimodal Feedback Generation Network (PMFGN) that generates personalized feedback for teachers to evaluate assignments involving multimodal inputs such as images, audios, and texts.
Outcome: The proposed model outperforms baseline models on real-world K-12 education data and detailed ablation experiments to deepen understanding of the proposed framework.
What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User Perception (2024.naacl-long)

Copied to clipboard

Challenge: Question answering models can often be black boxes, as their reasoning process is mostly opaque.
Approach: They analyze the effect of rationales generated by QA models on user feedback and how well they enable users to understand and trust model answers.
Outcome: The proposed model can be used to improve model responses by removing feedback from end users and enhancing model outputs by using natural language feedback.
System-Level Natural Language Feedback (2024.eacl-long)

Copied to clipboard

Challenge: Existing studies on NL feedback focus on instance-level approaches to refine specific examples, but we present a framework for system-level use of NL.
Approach: They propose a framework for system-level use of natural language feedback . they use feedback to formalize system-design decisions in a human-in-the-loop-process .
Outcome: The proposed framework improves search query and dialog response generation and human written instance-level feedback brings further gains over GPT-3.5 written feedback.
Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for assessing the validity of explanations for NLI are time-consuming and prone to logical errors.
Approach: They propose a framework that integrates Large Language Models and Theorem Provers to verify and refine natural language explanations through crowd-sourcing . they propose to use TPs to generate and formalise explanatory sentences and suggest potential inference strategies for NLI.
Outcome: The proposed framework generates and formalises explanatory sentences and suggests potential inference strategies for NLI.
InfoMetIC: An Informative Metric for Reference-free Image Caption Evaluation (2023.acl-long)

Copied to clipboard

Challenge: Existing image captioning metrics provide a single score to measure caption qualities, which are less explainable and informative.
Approach: They propose an Informative Metric for Reference-free Image Caption evaluation to support this feedback . they propose to provide a text precision score, a vision recall score and an overall quality score .
Outcome: The proposed method improves on existing metrics on multiple benchmarks and compares coarse-grained scores with human judgements.
Denoising Neural Network for News Recommendation with Positive and Negative Implicit Feedback (2022.findings-naacl)

Copied to clipboard

Challenge: Existing work on news recommendation only used positive and negative implicit feedback and suffered from the noise impact.
Approach: They propose a denoising neural network for news recommendation with positive and negative implicit feedback, named DRPN.
Outcome: The proposed method improves on the real-world large-scale dataset.
BiKT: Enabling Bidirectional Knowledge Transfer Between Pretrained Models and Sequential Downstream Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing frameworks adapt from initial pretrained model to each downstream task directly, but ignore sequential nature of downstream tasks and feedback effect on pretrained models.
Approach: They propose a framework to enable bidirectional knowledge transfer between pretrained models and downstream tasks in rounds.
Outcome: The proposed framework improves on 9 GLUE datasets and 6 SuperGLUEs.
RAG-Critic: Leveraging Automated Critic-Guided Agentic Workflow for Retrieval Augmented Generation (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated remarkable performance across a wide range of downstream tasks.
Approach: They propose a framework that leverages a critic-guided agentic workflow to improve RAG capabilities autonomously.
Outcome: The proposed framework improves RAG capabilities autonomously by leveraging a critic-guided agentic workflow.
Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems (N18-1)

Copied to clipboard

Challenge: Existing methods for learning task-oriented dialogues include applying reinforcement learning with user feedback on supervised pre-training models.
Approach: They propose a hybrid imitation and reinforcement learning method that integrates user feedback and reinforcement training to improve the agent's performance.
Outcome: The proposed method can learn from the mistake it makes via imitation learning from user teaching and feedback.
Speak to your Parser: Interactive Text-to-SQL with Natural Language Feedback (2020.acl-main)

Copied to clipboard

Challenge: a natural language interface (NLI) can be used to correct semantic parsing errors . human correction accuracy is 81.5%, but the best model achieves only 25.1% .
Approach: They propose a task where humans can provide free-form natural language feedback to correct a system when it generates an inaccurate interpretation of an initial utterance.
Outcome: The proposed model improves the parsing accuracy while maintaining flexibility of natural language interaction.
Apeiron: A Scalable LLM-agentic Framework for Autonomous Full-lifecycle Demand-optimized Application Synthesis (2026.findings-acl)

Copied to clipboard

Challenge: Traditional, rigid, 'one-size-fits-all' apps are struggling in the contemporary landscape.
Approach: They propose a scalable and extensible framework for addressing *amorphous* user demands through autonomous, full-lifecycle application synthesis.
Outcome: The proposed framework outperforms baselines in CUA ratings and user-demand task scores across 300 app scenarios, 2,400 personas, and 46,338 demands.
Improving Automatic Grammatical Error Annotation for Chinese Through Linguistically-Informed Error Typology (2025.coling-main)

Copied to clipboard

Challenge: In educational settings, GEC systems provide immediate and consistent feedback to both native (L1) and non-native (L2) language learners.
Approach: They propose a framework that provides detailed feedback on 12-16% of all errors by identifying them under a new error typology, specific enough to uncover subtle differences in error patterns between L1 and L2 writings.
Outcome: The proposed framework can provide detailed feedback on 12-16% of all errors, revealing subtle differences in error patterns between L1 and L2 writings.
Reinforced Training Data Selection for Domain Adaptation (P19-1)

Copied to clipboard

Challenge: Existing approaches to learn domains with massive data are not easy to implement and require a predefined threshold.
Approach: They propose a framework that searches for training instances relevant to the target domain and learns better representations for them.
Outcome: The proposed framework is effective in data selection and representation, but generalized to accommodate different NLP tasks.
Unlocking Large Audio-Language Models for Interactive Language Learning (2026.findings-eacl)

Copied to clipboard

Challenge: Computer-Assisted Pronunciation Training (CAPT) systems provide unintuitive feedback that lacks actionable guidance.
Approach: They propose to use audio-language models to provide more user-friendly feedback for pronunciation training.
Outcome: The proposed model outperforms baselines on mispronunciation detection and suggestion generation.
Learning to Learn Semantic Parsers from Natural Language Supervision (D18-1)

Copied to clipboard

Challenge: Existing logical forms require a user to be familiar with the underlying structure to learn a semantic parser.
Approach: They propose a method for training semantic parsers from natural language feedback . they use natural language inputs to parse feedback to leverage it as a form of supervision .
Outcome: The proposed algorithm learns a semantic parser from users’ corrections expressed in natural language.
MindAgent: Emergent Gaming Interaction (2024.findings-naacl)

Copied to clipboard

Challenge: Large foundation models (LFMs) can perform complex scheduling in a multi-agent system and can coordinate agents to complete complex tasks that require extensive collaboration.
Approach: They propose a gaming-based infrastructure that evaluates LFMs' planning and coordination capabilities in the context of gaming interaction.
Outcome: The proposed infrastructure can be deployed in a customized VR version of Cuisineworld and adapted in the “Minecraft” domain.
CoTKR: Chain-of-Thought Enhanced Knowledge Rewriting for Complex Knowledge Graph Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge rewriting methods may include irrelevant information, omit crucial details, or fail to align with the question’s semantics.
Approach: They propose a new rewriting method CoTKR for generating reasoning traces and corresponding knowledge in an interleaved manner, thereby mitigating the limitations of single-step knowledge rewrite.
Outcome: The proposed method mitigates the limitations of single-step knowledge rewriting and bridges the preference gap between the knowledge reactor and the question answering (QA) model.
Curriculum Learning Meets Weakly Supervised Multimodal Correlation Learning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies have used the correlation information stored in samples for self-supervised learning, but they feed the training pairs in a random order without consideration of difficulty.
Approach: They propose to inject curriculum learning into weakly supervised multimodal correlation learning by scoring and feeding pairs according to difficulty.
Outcome: The proposed model achieves state-of-the-art on multimodal sentiment analysis without human annotation.
MPO: Boosting LLM Agents with Meta Plan Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for interactive planning tasks suffer from planning hallucinations and require retraining for each new agent.
Approach: They propose a framework that leverages explicit guidance through meta plans to assist agent planning and enables continuous optimization based on feedback from the agent’s task execution.
Outcome: The proposed framework outperforms existing baselines on two representative tasks and significantly improves task completion efficiency and generalization capabilities.
Gradient Imitation Reinforcement Learning for Low Resource Relation Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods to extract relation facts from limited labeled corpora are laborintensive to obtain . Existing approaches use self-training to generate pseudo labels that will cause gradual drift problem or leverage meta-learning scheme which does not solicit feedback explicitly.
Approach: They propose a Gradient Imitation Reinforcement Learning method to encourage pseudo label data to imitate gradient descent direction on labeled data and bootstrap its optimization capability through trial and error.
Outcome: The proposed method handles two major scenarios in low-resource relation extraction when no unlabeled data is available.
Dataset and Baseline for Automatic Student Feedback Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Currently, student feedback is collected manually, but it does not indicate the student's opinion on different aspects of the teaching/learning process.
Approach: They propose to annotate student feedback corpus which contains 3000 instances . they propose a hierarchical taxonomy for aspect categorization, which covers all areas .
Outcome: The proposed model can be used for aspects analysis, document level sentiment analysis and document level analysis.
Learning Improvised Chatbots from Adversarial Modifications of Natural Language Feedback (2020.findings-emnlp)

Copied to clipboard

Challenge: Currently, user feedback contains extraneous sequences hindering their usefulness as a training sample.
Approach: They propose a generative adversarial model that converts noisy feedback into a plausible natural response in a conversation and fools the discriminator which distinguishes feedback from natural responses.
Outcome: The proposed model improves the original chatbot performance from 69.94%to 75.96% in ranking correct responses on the PERSONACHATdataset.
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Recent work has shown promise by incorporating pixel-level visual information into the reasoning process, enabling VLMs to access high-resolution visual details during their thought process.
Approach: They propose a framework that dynamically determines necessary pixel-level operations based on the input query.
Outcome: The proposed model achieves 73.4% accuracy on HR-Bench 4K while maintaining a tool usage ratio of only 20.1%, improving accuracy and reducing tool usage by 66.5% compared to the previous methods.
PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling (2024.emnlp-main)

Copied to clipboard

Challenge: Prompt optimization aims to find the best prompt to a large language model (LLM) for a given task.
Approach: They propose a method to optimize prompts for LLM-driven multi-step tasks using a human-designed feedback rule.
Outcome: The proposed method outperforms human-engineered prompts and several other prompt optimization methods on 11 representative multi-step tasks.
Multi-Level Feedback Generation with Large Language Models for Empowering Novice Peer Counselors (2024.acl-long)

Copied to clipboard

Challenge: Existing mechanisms of providing feedback rely on human supervision . existing mechanisms of delivering feedback largely rely only on human oversight .
Approach: They propose to leverage large language models to provide contextualized feedback to peer counselors . they construct a publicly available dataset with detailed feedback annotations of 400 conversations .
Outcome: The proposed method minimizes the risk of potentially harmful and low-quality feedback generation.
Sentence Level Human Translation Quality Estimation with Attention-based Neural Networks (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for assessing translation quality rely on manual features and external knowledge.
Approach: They propose to use a neural model without feature engineering to detect which parts in sentence pairs are most relevant for assessing quality.
Outcome: The proposed model outperforms feature-based methods on a large human annotated dataset.
Improving In-Context Learning with Prediction Feedback for Sentiment Analysis (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved promising results in sentiment analysis through the in-context learning paradigm.
Approach: They propose a framework that incorporates prior predictions and feedback to improve sentiment understanding by incorporating prior feedback and leveraging a feedback-driven prompt.
Outcome: The proposed framework improves on nine sentiment analysis datasets with an average improvement of 5.95% over conventional methods.
Improving Conversational Recommendation Systems via Bias Analysis and Language-Model-Enhanced Data Augmentation (2023.findings-emnlp)

Copied to clipboard

Challenge: Conversational Recommendation System (CRS) is a rapidly growing research area, along with advancements in language modelling techniques.
Approach: They propose to use a benchmark dataset to develop CRS models and address biases arising from feedback loop inherent in multi-turn interactions to enhance model performance while mitigating biase.
Outcome: The proposed strategies improve on ReDial and TG-ReDial benchmark datasets and offer additional insights on addressing multiple newly formulated biases.
Negative language transfer in learner English: A new dataset (2021.naacl-main)

Copied to clipboard

Challenge: This dataset contains annotated error causes for learner writing errors that tie learner mistakes to structures from their first language.
Approach: They propose a learner English dataset enhanced with annotated error causes and concrete examples of learner errors that relate to their first languages.
Outcome: The proposed dataset will be used to analyze learner errors related to language transfer from the learners’ first language.
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback (2024.emnlp-main)

Copied to clipboard

Challenge: Large Multimodal Models excel at comprehending human instructions and demonstrate remarkable results across a broad spectrum of tasks.
Approach: They propose an algorithm that alters REinforcement Learning and Supervised Fine-Tuning to refine large multimodal models with specific preferences.
Outcome: The proposed algorithm achieves 70% win rate compared to baseline models judged by GPT-4o.
Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney (2023.emnlp-main)

Copied to clipboard

Challenge: Generating images with Text-to-Image models often requires multiple trials, where human users iteratively update their prompt based on feedback, namely the output image.
Approach: They compile a dataset of iterative interactions of human users with Midjourney and analyze the dynamics of the user prompts along these iterations.
Outcome: The proposed model produces better images for a specific language style than other models.
Unified Demonstration Retriever for In-Context Learning (2023.acl-long)

Copied to clipboard

Challenge: In-context learning is a new learning paradigm where a language model conditions on a few input-output pairs (demonstrations) and a test input, and directly outputs the prediction.
Approach: They propose a single model to retrieve demonstrations for a wide range of tasks by combining training signals from various tasks into a unified list-wise ranking formulation by language model’s feedback.
Outcome: The proposed model outperforms baselines on 30+ tasks across 13 task families and multiple data domains.
Learning to Reason from Feedback at Test-Time (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to utilizing feedback are expensive and lack the time to perform iterative interactions with the environment.
Approach: They propose a novel paradigm that formulates feedback utilization as an optimization problem at test time and a learnable test-time optimizer to effectively exploit feedback.
Outcome: The proposed paradigm improves scalability and performance on two large language models across four reasoning datasets.
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations (2024.findings-naacl)

Copied to clipboard

Challenge: supervised systems have not replaced dedicated supervised models for machine translation tasks.
Approach: They propose to guide LLMs to post-edit MT with feedback from MQM annotations . they then fine-tune the LLM to improve its ability to exploit the feedback .
Outcome: The proposed model improves TER, BLEU and COMET scores on Chinese-English, English-German and English-Russian data.
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text (2025.acl-long)

Copied to clipboard

Challenge: Qualitative analysis of experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity).
Approach: They hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions.
Outcome: The annotators who frequently use LLMs for writing tasks outperform commercial and open-source detectors even without evasion tactics like paraphrasing and humanization.
KidsArtBench: Multi-Dimensional Children’s Art Evaluation with Attribute-Aware MLLMs (2026.eacl-long)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) show impressive capabilities across visual–language tasks, but their capacity to evaluate artistic expression remains limited.
Approach: They propose an attribute-specific multi-LoRA approach where each attribute corresponds to a distinct evaluation dimension in the scoring rubric.
Outcome: The proposed approach increases correlation from 0.468 to 0.653 on Qwen2.5-VL-7B, with the largest gains on perceptual dimensions and narrowed gaps on higher-order attributes.
Collaborative Policy Learning for Open Knowledge Graph Reasoning (D19-1)

Copied to clipboard

Challenge: Existing models of knowledge graph reasoning suffer from limited performance when working on sparse and incomplete graphs due to the lack of evidential paths that can reach target entities.
Approach: They propose a framework to train two collaborative agents to reason for missing facts over a graph augmented by a text corpus.
Outcome: Experiments on two public datasets show the proposed approach is effective on a knowledge graph reasoning task.
Exploiting Edited Large Language Models as General Scientific Optimizers (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for solving optimization problems in scientific scenarios use observational feedback as additional textual descriptions, but these methods struggle to utilize it effectively.
Approach: They propose a generalized approach to boost mathematical optimization in scientific scenarios by using observational feedback from LLMs as additional textual descriptions.
Outcome: The proposed method outperforms existing state-of-the-art methods on six different tasks using six different LLM backbones.
Neural Quality Estimation of Grammatical Error Correction (D18-1)

Copied to clipboard

Challenge: Grammatical error correction systems are expected to correct most learners’ writing errors, but in practice they often produce spurious corrections and fail to correct many errors, thereby misleading learners.
Approach: They propose to use supervised learning to estimate the quality of GEC output sentences to help instructors decide whether to correct the errors or ignore them altogether.
Outcome: The proposed model improves on a feature-based baseline and shows that the state-of-the-art system can be improved when quality scores are used as features for re-ranking the N-best candidates.
Lightweight Grammatical Annotation in the TEI: New Perspectives (L18-1)

Copied to clipboard

Challenge: a small set of descriptive devices have been made available for lightweight linguistic annotation . merit of a predefined TEI tagset is the homogeneity of tagging and better interoperability of simple linguistic resources encoded in the TE.
Approach: They propose a new attribute class that would gather token-level attributes facilitating simple linguistic annotation.
Outcome: The proposed attribute class addresses community feedback on the lack of a specific tagset for lightweight linguistic annotation within the TEI.
MAGID: An Automated Pipeline for Generating Synthetic Multi-modal Datasets (2024.naacl-long)

Copied to clipboard

Challenge: Existing approaches to augment textual dialogues with retrieved images pose privacy, diversity, and quality constraints.
Approach: They propose a framework to augment text-only dialogues with diverse and high-quality images by using a diffusion model and a feedback loop.
Outcome: The proposed framework is comparable to or better than baselines, with significant improvements in human evaluation, especially against retrieval baselines where the image database is small.
Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) demonstrate exceptional instruct-following ability to complete downstream tasks.
Approach: They propose a black-box combinatorial optimization framework that iteratively improves lexical choices in prompts by a search strategy related to word influence.
Outcome: The proposed framework recovers the model's ability to instruct-follow and solve downstream tasks even when the variations are imperceptible to humans.
Test-time Corpus Feedback: From Retrieval to RAG (2026.findings-eacl)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) pipelines treat retrieval and reasoning as isolated components, limiting performance on complex tasks.
Approach: They propose to integrate large language models with retrieval to improve query quality . they also propose to use feedback to improve the query, retrieved context, or document pool .
Outcome: The proposed methods bridge IR and NLP perspectives and highlight retrieval as a dynamic, learnable component of end-to-end RAG systems.
UCFE: A User-Centric Financial Expertise Benchmark for Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have expanded their potential applications in finance.
Approach: They propose a framework to evaluate the ability of large language models to handle financial tasks using human expert evaluations and task-specific interactions.
Outcome: The proposed framework evaluates the ability of large language models to handle complex financial tasks and combines human expert evaluations with dynamic, task-specific interactions to simulate the complexities of evolving financial scenarios.
Query Rewriting in Retrieval-Augmented Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on adapting either the retriever or the reader, but this approach is more focused on adaptation of the query itself.
Approach: They propose a new framework for retrieval-augmented Large Language Models . they propose rewrite-retrieve-read instead of retrieve-then-read .
Outcome: The proposed framework improves performance on downstream tasks, open-domain QA and multiple-choice QA.
JRE-L: Journalist, Reader, and Editor LLMs in the Loop for Science Journalism for the General Audience (2025.naacl-long)

Copied to clipboard

Challenge: Science journalism reports current scientific discoveries to non-specialists, aiming to enable public comprehension of the state of the art.
Approach: They propose a framework that integrates three LLMs mimicking the writing-reading-feedback-revision loop.
Outcome: The proposed framework generates articles that are more accessible than existing methods, including prompting single advanced models such as GPT-4 and other LLM-collaboration strategies.
Multi-Granularity Optimization for Non-Autoregressive Translation (2022.emnlp-main)

Copied to clipboard

Challenge: Non-autoregressive machine translation suffers severe performance deterioration due to the naive independence assumption.
Approach: They propose a method which collects model behaviours on translation segments of various granularities and integrates feedback for backpropagation to reduce latency.
Outcome: Experiments on four benchmark datasets show that the proposed method outperforms baseline models trained with cross-entropy loss and achieves the best performance on WMT’16 EnRo and highly competitive results on WTM’14 EnDe.
Adapting LLM Agents with Universal Communication Feedback (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated potential for LLM agents.
Approach: They propose a universal buffer and iterative pipeline to store feedback and itersative pipelines to enable LLM agents to explore and update their policy in an environment.
Outcome: The proposed approach outperforms supervised instruction fine-tuning baselines on four datasets.
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a new system that leverages the encyclopedic knowledge and linguistic reasoning capabilities of Large Language Models (LLMs) is proposed to enhance the productivity of researchers . a researcher's research idea generation process involves problem identification, method development, experiment design and iterative revision .
Approach: They propose a system that leverages encyclopedic knowledge and linguistic reasoning capabilities of Large Language Models to assist researchers in their work.
Outcome: The proposed system generates novel ideas based on human and model-based evaluations . it leverages encyclopedic knowledge and linguistic reasoning capabilities of Large Language Models based systems .
Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to write grounded, long-form articles have limited planning capacity and require extensive research and planning in the pre-writing stage.
Approach: They propose a system for the Synthesis of Topic Outlines throughRetrieval and Multi-perspective Question Asking that models the pre-writing stage by (1) discovering diverse perspectives in researching the given topic, (2) simulating conversations where writers carrying different perspectives pose questions to a topic expert grounded on trusted Internet sources, (3) curating the collected information to create an outline.
Outcome: The proposed system is based on a dataset of high-quality Wikipedia articles and evaluates the pre-writing stage.
Simulating Bandit Learning from User Feedback for Extractive Question Answering (2022.acl-long)

Copied to clipboard

Challenge: Explicit feedback from users can be used to continually improve system performance.
Approach: They study the potential of learning from user feedback for extractive question answering by simulating feedback using supervised data.
Outcome: The proposed model improves on a few examples and can be deployed in new domains without any data annotation effort.
Logical DA: Enhancing Data Augmentation for Logical Reasoning via a Multi-Agent System (2025.findings-acl)

Copied to clipboard

Challenge: Existing data augmentation paradigms isolate data synthesis from label validation, thereby reducing their utility for complex reasoning tasks.
Approach: They propose a framework for enhancing reasoning-focused data augmentation in few-shot learning scenarios that integrates four agents through two synergistic phases: diverse data generation and label verification.
Outcome: The proposed framework achieves the highest average improvement in task accuracy in both fine-tuning and in-context learning paradigms.
Learning from Relevant Subgoals in Successful Dialogs using Iterative Training for Task-oriented Dialog Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Task-oriented Dialog (ToD) systems have to solve multiple subgoals to accomplish user goals, whereas feedback is often obtained only at the end of the dialog.
Approach: They propose an iterative training approach that uses subgoals to improve task-oriented dialog systems.
Outcome: The proposed approach improves on a popular ToD benchmark by combining fine-tuning and preference learning steps.
Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations (2026.findings-acl)

Copied to clipboard

Challenge: Recent MT metrics like xCOMET, Met-ricX, and Remedy have strong correlations with human preferences, but they are black boxes, revealing little insight into why a translation is good or bad.
Approach: They propose a reasoning-driven generative MT metric trained with reinforcement learning from pairwise translation preferences without requiring error-span annotations or distillation from closed LLMs.
Outcome: The proposed reasoning-driven generative MT metric produces step-by-step analyses of accuracy, fluency, and completeness, enabling more interpretable assessments.
MIAPARLE: Online training for the discrimination of stress contrasts (L18-1)

Copied to clipboard

Challenge: Second language learners tend to imprint the prosody of their mother language onto the second language (L2) . this can hamper communication between learners and natives, and can also affect the credibility of learners and how they are evaluated by others.
Approach: They propose a tool that focuses on stress perception for speakers whose L1 is a fixed-stress language, such as French.
Outcome: The tool is particularly useful for speakers whose L1 is a fixed-stress language, such as French.
ESCRITO - An NLP-Enhanced Educational Scoring Toolkit (L18-1)

Copied to clipboard

Challenge: Existing implementations are very specific to specific use cases and datasets.
Approach: ESCRITO is a toolkit for scoring student writings using NLP techniques . authors propose teachers and NLP researchers to use APIs for scoring pipelines .
Outcome: ESCRITO is a toolkit for scoring student writings using NLP techniques . it addresses two main user groups: teachers and NLP researchers .
EPO: Hierarchical LLM Agents with Environment Preference Optimization (2024.emnlp-main)

Copied to clipboard

Challenge: Long-horizon decision-making tasks require extensive planning over multiple steps, maintaining coherence and goal orientation, which is difficult for LLMs that are typically designed for more immediate and localized predictions.
Approach: They propose a hierarchical framework that decomposes complex tasks into manageable subgoals, utilizing separate LLMs for subgoal prediction and low-level action generation.
Outcome: The proposed framework achieves first place on the ALFRED public leaderboard and demonstrates its potential to improve long-horizon decision-making in diverse environments.
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models for automatic dialogue tutoring fail to provide accurate feedback or reveal solutions to students too early.
Approach: They propose a framework to generate one-to-one teacher-student tutoring dialogues by pairing human teachers with a Large Language Model (LLM) they use scaffolding questions and annotations to fine-tune models to be more effective tutors .
Outcome: The proposed framework can generate 3k one-to-one teacher-student tutoring dialogues grounded in multi-step math reasoning problems.
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews (2024.acl-long)

Copied to clipboard

Challenge: Existing systems that can interpret complex writing feedback and edit documents in response are limited on the most demanding writing tasks.
Approach: They propose to use peer feedback to revise scientific papers based on peer feedback . they provide labels linking each reviewer comment to the specific paper edits made by the author .
Outcome: The proposed model fails to identify which edits correspond to a comment and the original paper.
PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to Automated Essay Scoring (AES) treat scoring and feedback as separate components, resulting in fragmentation.
Approach: They propose a psychometrically-aware framework that integrates diagnostic assessment with instructional scaffolding through a shared latent ability representation.
Outcome: The proposed framework integrates diagnostic assessment with instructional scaffolding through a shared latent ability representation.
Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data (2023.emnlp-main)

Copied to clipboard

Challenge: Despite the promising potential of chat models, they are only accessible through restricted APIs, creating barriers for new research and progress in the field.
Approach: They propose a pipeline that can automatically generate a high-quality multi-turn chat corpus by leveraging ChatGPT to engage in a conversation with itself.
Outcome: The proposed pipeline generates a high-quality multi-turn chat corpus by leveraging ChatGPT to engage in a conversation with itself, simulating both user and AI responses.
Distilling ChatGPT for Explainable Automated Student Answer Assessment (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing automated student answer assessment models lack explainable and faithful feedback.
Approach: They propose a framework that leverages ChatGPT for student answer scoring and rationale generation.
Outcome: The proposed method improves the overall QWK score by 11% compared to ChatGPT.
Rhetoric, Logic, and Dialectic: Advancing Theory-based Argument Quality Assessment in Natural Language Processing (2020.coling-main)

Copied to clipboard

Challenge: Existing work on argument quality (AQ) focuses on overall quality, but there is no large-scale theory-based corpus and corresponding computational models.
Approach: They propose to use a large-scale English multi-domain argumentative writing corpus annotated with theory-based AQ scores to assess argument quality.
Outcome: The proposed methods improve argument quality in three domains and can be used as strong baselines for future work.
Few-shot Subgoal Planning with Language Models (2022.naacl-main)

Copied to clipboard

Challenge: Pre-trained language models have shown successful progress in many text understanding benchmarks.
Approach: They propose a strategy to re-rank language model predictions based on interaction and feedback from the environment.
Outcome: The proposed approach shows competitive performance on subgoal prediction and task completion in the ALFRED benchmark compared to prior methods that assume more subgoals supervision.
Dual-Feedback Knowledge Retrieval for Task-Oriented Dialogue Systems (2023.emnlp-main)

Copied to clipboard

Challenge: Current approaches to task-oriented dialogue systems integrate knowledge retrieval and response generation, which poses scalability challenges when dealing with extensive knowledge bases.
Approach: They propose a retriever-generator architecture that harnesses a retrieval and a generator to generate system responses by using feedback from the generator as pseudo-labels.
Outcome: The proposed architecture shows superior performance on three benchmark datasets.
Learning to Retrieve Iteratively for In-Context Learning (2024.emnlp-main)

Copied to clipboard

Challenge: In-context learning is a powerful tool for learning large language models.
Approach: They propose an iterative retrieval framework that empowers retrievers to make iterable decisions through policy optimization.
Outcome: The proposed framework outperforms existing methods on semantic parsing datasets with 4M additional parameters for state encoding.
A Simple Log-based Loss Function for Ordinal Text Classification (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for ordinal text classification do not incorporate ordinal character into their feedback.
Approach: They propose a new ordinal log-loss loss function that incorporates ordinal character into its feedback.
Outcome: The proposed loss function outperforms state-of-the-art methods on four benchmark text classification datasets.
CERET: Cost-Effective Extrinsic Refinement for Text Generation (2024.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) generate incomplete, biased or misleading outputs in their initial attempts.
Approach: They propose a method for refining text generation that takes into account semantic stability, entailment and inter-sample uncertainty measures.
Outcome: The proposed method outperforms self-consistency and self-rerank baselines under various task setups by 1.6% and 3.5% respectively.
Finding Support Examples for In-Context Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: In-context learning is a new learning paradigm where a language model observes a few examples and directly outputs the test input’s prediction.
Approach: They propose a method to find “support examples” for in-context learning by filtering a training dataset and a progressive filtering process to filter out uninformative examples.
Outcome: The proposed method outperforms baselines and shows that each component contributes critically to the improvements.
ComPO: Community Preferences for Language Model Personalization (2025.naacl-long)

Copied to clipboard

Challenge: Current methods for training language models with human feedback rely on subjective preferences that are assumed to account for an "average" user . however, annotating preferences is inherently subjective and results in generic models that generate outputs not preferred by many user groups.
Approach: They propose a method to personalize preference optimization in LMs by contextualizing the probability distribution of model outputs with the preference provider.
Outcome: The proposed method improves performance by focusing on group-level preferences rather than individual feedback.
From Relevance to Utility: Evidence Retrieval with Feedback for Fact Verification (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing evidence retrieval models are based on probability ranking principle . existing models do not align with retrieval-enhanced verification frameworks .
Approach: They propose a feedback-based evidence retriever that optimizes the evidence retrieval process by incorporating feedback from the claim verifier.
Outcome: Empirical studies show that the proposed method is superior to baseline methods.
A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support (2020.emnlp-main)

Copied to clipboard

Challenge: Empathy measurement has predominantly occurred in synchronous, face-to-face settings, and may not translate to asynchronous, text-based contexts.
Approach: They propose a computational approach to understanding how empathy is expressed in online mental health platforms.
Outcome: The proposed model can identify empathic conversations and extract rationales from them.
Self-Critique and Refinement for Faithful Natural Language Explanations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work has demonstrated that Large Language Models (LLMs) can self-critique and refine their initial outputs, but this capability remains unexplored for improving explanation faithfulness.
Approach: They propose a framework that enables models to improve the faithfulness of their own explanations through an iterative critique and refinement process without external supervision.
Outcome: The proposed framework reduces unfaithfulness rates in three datasets and four state-of-the-art LLMs by 36% compared to 54.81% for baseline.
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for reweighting data mixtures rely on manual designation with certain heuristics based on intuition or empirical results.
Approach: They propose a model-based framework that learns to re-weight domains by reinforcement learning on large quantities of data mixing trajectories with corresponding feedback from an evaluation environment.
Outcome: The proposed framework outperforms baselines in achieving balanced performance across source and target fields and domain spaces without retraining.
Unified Low-Resource Sequence Labeling by Sample-Aware Dynamic Sparse Finetuning (2023.emnlp-main)

Copied to clipboard

Challenge: Named Entity Recognition, Relation Extraction, Semantic Role Labeling are examples of sequence labeling problems that require finetuning to the target format.
Approach: They propose a dynamic sparse finetuning strategy that selectively focuses on a fraction of parameters, informed by feedback from highly regressing examples.
Outcome: The proposed approach improves performance in low-resource settings and in extreme low-level settings.
Enhancing Recommendation Explanations through User-Centric Refinement (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing explanations for user reviews often fail to meet user-centric aspects, reducing their usefulness to users.
Approach: They propose a paradigm that refines initial explanations generated by existing models during the inference stage to enhance their quality in multiple aspects.
Outcome: The proposed model improves explanations generated by existing models during the inference stage to enhance their quality in multiple aspects.
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)

Copied to clipboard

Challenge: Language is a powerful means of communication and should be regarded as more than just a collection of tokens.
Approach: They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices.
Outcome: The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices.
PACE: Improving Prompt with Actor-Critic Editing for Large Language Model (2024.findings-acl)

Copied to clipboard

Challenge: Prompt with Actor-Critic Editing (PACE) for LLMs improves performance of different human-written prompts, resulting in significant performance discrepancies.
Approach: They propose to use LLMs as actors and critics to enable automatic prompt editing by taking feedback from both actors performing prompt and criticizing response into account.
Outcome: The proposed model improves the performance of human-written prompts by 98% and compares to high-quality human-writing prompts.
IEvoAgent: Evolving Conversational Agent based on User Implicit Feedback (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to optimize conversational agents often rely on explicit preference pairs and expert evaluations.
Approach: They propose a conversational agent framework that leverages the structured dependency between agent responses and user reactions to extract implicit feedback.
Outcome: The proposed framework improves on MT-Bench-101, WildBench, and FB-Bech, and shows that mining implicit feedback supports better multi-turn alignment under evolving user preferences.
SaFeRDialogues: Taking Feedback Gracefully after Conversational Safety Failures (2022.acl-long)

Copied to clipboard

Challenge: Existing open-domain conversational models can easily be made to talk in inadequate ways.
Approach: They propose a task and dataset of graceful responses to safety feedback . they collect 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging feedback.
Outcome: The proposed model improves on a dataset of 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging the feedback.
The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects (2025.acl-long)

Copied to clipboard

Challenge: Recent large-scale T2I models like DALLE-3 have made progress in reducing gender stereotypes when generating single-person images.
Approach: They propose a framework that queries T2I models to depict two individuals with gender-stereotyped social identities to evaluate gender biases.
Outcome: The proposed framework reduces gender stereotypes when generating images with more than one person.
Dynamic Uncertainty Ranking: Enhancing Retrieval-Augmented In-Context Learning for Long-Tail Knowledge in LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Prior work has shown that in-context learning (ICL) with retriever augmentation can help LLMs better capture long-tail knowledge, reducing their reliance on pre-trained data.
Approach: They propose a reinforcement learning-based dynamic uncertainty ranking method that accounts for the varying impact of each retrieved sample on LLM predictions.
Outcome: The proposed method outperforms baseline models on question-answering datasets by 2.76% and 5.96% on long-tail questions that elude zero-shot inference.
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods that optimize for scalar scores or ranking reward ignore multi-dimensional nature of human preferences.
Approach: They propose to extend the preference of Direct Preference Optimization to two dimensions: segments and aspects.
Outcome: The proposed framework decomposes the overall objective into multi-segment and multi-aspect objectives.
LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that LLMs struggle with instructions containing multiple constraints.
Approach: They propose a self-correction pipeline that decomposes the original instruction into a list of constraints and uses a Critic model to decide when and where the LLM’s response needs refinement.
Outcome: The proposed model outperforms GPT-4 on RealInstruct and IFEval even with weak feedback.
The Niki and Julie Corpus: Collaborative Multimodal Dialogues between Humans, Robots, and Virtual Agents (L18-1)

Copied to clipboard

Challenge: Niki and Julie corpus contains more than 600 dialogues between humans and robots . corpus includes audio and video recordings, results of ranking tasks, questionnaire responses .
Approach: the corpus contains more than 600 dialogues between human participants and a robot . the dialogues are part of a collaborative item-ranking task designed to measure influence .
Outcome: the corpus contains more than 600 dialogues between human participants and a robot or virtual agent . the dialogues contain conversational errors by the robot, which simulates typical of modern automated agents .
Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction (2025.findings-naacl)

Copied to clipboard

Challenge: Recent approaches to ensembling neural machine translation models require inference across all candidate models, leading to significant computational overhead.
Approach: They propose a reinforcement learning-based strategy that improves the CSB by selecting a small, fixed number of candidates and identifying optimal groups to pass to the fusion block for each input sentence.
Outcome: The proposed approach improves the CSB by selecting a small, fixed number of candidates and identifying optimal groups to pass to the fusion block for each input sentence.
CARD: Cross-modal Agent Framework for Generative and Editable Residential Design (2025.emnlp-main)

Copied to clipboard

Challenge: Architectural design automation has made significant progress, but the complexity of open-world environments makes residential design a challenging task.
Approach: They propose a framework that leverages a system of specialized cross-modal agents to adapt to open-world residential design.
Outcome: The proposed framework enables users to generate and edit residential design without requiring specialized expertise.
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models for dialog tutoring fail to detect student errors and tailor their feedback to them.
Approach: They propose to build dialog tutoring models to scaffold students' problem-solving and verify student solutions by using automatic and human evaluation.
Outcome: The proposed model improves the quality of the tutor response generation by detecting student errors and adjusting the feedback to the errors.
BERT Learns to Teach: Knowledge Distillation with Meta Learning (2022.acl-long)

Copied to clipboard

Challenge: Existing knowledge distillation methods are based on teacher model, but have drawbacks . a teacher model is fixed during training, but meta learning can improve student performance .
Approach: They propose a meta learning framework that allows the teacher network to learn to better transfer knowledge to the student network.
Outcome: Experiments show that MetaDistil can improve on existing methods and is less sensitive to student capacity and hyperparameters.
Think&Cite: Improving Attributed Text Generation with Self-Guided Tree Search and Progress Reward Modeling (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are prone to hallucinations and producing factually incorrect information.
Approach: They propose a framework that allows LLMs to generate citations that provide evidence for any statement.
Outcome: The proposed framework outperforms baseline approaches on three datasets and significantly outperformed baseline approaches.
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to enhance code generation performance include integrating compiler feedback.
Approach: They propose a method that integrates compiler feedback to improve one-off code generation performance.
Outcome: The proposed method improves one-off code generation performance on three benchmarks and can be applied to other domains that focus on final results and require long reasoning paths.
DORA: Dynamic Optimization Prompt for Continuous Reflection of LLM-based Agent (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have shown that reflection can enhance performance, but our investigation reveals an undesirable pattern in reflection framework: effective self-reflection occurs primarily at the beginning of iterations, with subsequent attempts failing to produce further improvements.
Approach: They propose a framework that generates task-adaptive reflection advice using an external open-source small language model.
Outcome: The proposed framework generates task-adaptive and diverse reflection advice in MiniWoB++ and Alfworld environments.
Can Active Label Correction Improve LLM-based Modular AI Systems? (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are powerful zero or few-shot learners and can generalize to a wide range of tasks without any model fine-tuning.
Approach: They propose to use LLM annotations to train smaller task-specific improved models that can replace LLMs.
Outcome: The proposed method can improve oracle performance with feedback on 17-24% fewer examples than the number of noisy examples in the dataset across three different NLP tasks.
TempCompass: Do Video LLMs Really Understand Videos? (2024.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks on video large language models lack a comprehensive feedback on temporal perception ability . current models cannot distinguish between different temporal aspects and are limited in task formats .
Approach: They propose a benchmark to evaluate temporal perception ability of video large language models . they construct conflicting videos that share the same static content but differ in a specific temporal aspect .
Outcome: The proposed benchmarks show that video large language models exhibit poor temporal perception ability.
Language Technology for Multilingual Europe: An Analysis of a Large-Scale Survey regarding Challenges, Demands, Gaps and Needs (L18-1)

Copied to clipboard

Challenge: a survey titled "Language Technology for Multilingual Europe" was conducted between May and June 2017 . 634 participants in 52 countries responded to the survey .
Approach: a large-scale survey was conducted to assess the best multilingual technologies in Europe. a total of 634 participants in 52 countries responded to the survey.
Outcome: The study aims to identify the biggest challenges, obstacles and gaps in European language technology . participants were encouraged to share concrete suggestions and recommendations on how present challenges can be turned into opportunities .
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement Learning from Hindsight Simulation (RLHF) can cause severe misalignment in generative AI, but it is not a universal method for fine-tuning large language models.
Approach: They propose a method that uses evaluator feedback to decouple alignment signal from potentially compromised predictions.
Outcome: The proposed method significantly outperforms RLHF in comparisons with baselines and human evaluations.
Sample Efficient Alignment Learning With Episodic Control (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing parametric methods for aligning large language models with task objectives are limited.
Approach: They propose a non-parametric framework that aligns large language models with task objectives . they use a key-value memory to store associations between generated text and its corresponding values .
Outcome: The proposed framework outperforms state-of-the-art baselines on harmless, helpful, and summarization tasks.
Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has developed algorithms for reinforcement learning from human feedback and AI-generated feedback.
Approach: They propose a method for reinforcement learning from human feedback and AI-generated feedback that incorporates weighting, ranking, and constraining to handle disparate rewards.
Outcome: The proposed method reduces disparity and enhances stability among rewards . empirical results show that the proposed method is efficient and straightforward .
Learning Efficient Dialogue Policy from Demonstrations through Shaping (2020.acl-main)

Copied to clipboard

Challenge: Using reinforcement learning to learn dialogue policy requires a large volume of interactions with users.
Approach: They propose a task-oriented dialogue agent that efficiently learns dialogue policy from demonstrations . they use an imitation model to distill knowledge from demonstration and reward shaping .
Outcome: The proposed agent efficiently learns dialogue policy from demonstrations through policy shaping and reward shaping.
Understanding Client Reactions in Online Mental Health Counseling (2023.acl-long)

Copied to clipboard

Challenge: Communication success relies heavily on reading participants’ reactions, but little research is on how listeners' reactions shape trajectories and outcomes of conversations.
Approach: They propose to use client reactions to predict counseling outcomes by using an annotation framework that encompasses counselors’ strategies and client reaction behaviors.
Outcome: The proposed framework can predict counselors' strategies and client reaction behaviors against a large-scale text-based counseling dataset.
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for improving large language models have focused on improving model responses rather than judgment capabilities, resulting in rapid saturation during iterative training.
Approach: They propose an iterative Meta-Rewarding step where the model judges its own judgements and uses that feedback to refine its judgment skills.
Outcome: The proposed model improves Llama-3-8B-Instruct from 22.9% to 39.4% on AlpacaEval 2 and 20.6% to 29.1% on Arena-Hard.
Zero-shot Visual Question Answering with Language Model Feedback (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for knowledge-based visual question answering are based on pre-trained language models.
Approach: They propose a language model guided captioning approach that leverages a pre-trained language model to generate captions for an image to help answer a visual question.
Outcome: The proposed method outperforms several competing methods on the knowledge-based VQA task and achieves comparable results to a fine-tuned VLP model.
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset (2025.naacl-long)

Copied to clipboard

Challenge: Learning from human feedback (LHF) has been used to mitigate the harms of large language models (LLMs) but the quality of this feedback and its effectiveness as a safety mitigation technique remain unclear.
Approach: They audit the Helpful and Harmless (HH) dataset by Anthropic and examine how conceptualization failures and quality issues identified in the dataset can create additional harms .
Outcome: The findings highlight the need for more nuanced, context-sensitive approaches to safety mitigation in large language models.
ACUEval: Fine-grained Hallucination Evaluation and Correction for Abstractive Summarization (2024.findings-acl)

Copied to clipboard

Challenge: Recent-proposed evaluation metrics for large language models have a preference-bias . however, such metrics often lack interpretability and only offer a single score .
Approach: They propose a metric that leverages the power of large language models to perform two sub-tasks: decomposing summaries into atomic content units and validating them against the source document.
Outcome: The proposed metric improves faithfulness scores on three summarization evaluation benchmarks by 3% compared to the next-best metric.
More Insightful Feedback for Tutoring: Enhancing Generation Mechanisms and Automatic Evaluation (2024.emnlp-main)

Copied to clipboard

Challenge: Incorrect student answers can be valuable learning opportunities provided that the student understands where they went wrong and why.
Approach: They propose to use a KL regularization term to achieve more targeted input representations and a preference optimization step to encourage student answer-adaptive feedback generation.
Outcome: The proposed model outperforms existing models in 3.3 METEOR points.
Domain Adaptation for Conversational Query Production with the RAG Model Feedback (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have focused on human-annotated search queries but they can not cover conversations of various domains.
Approach: They propose a domain adaptation framework that uses retrieval-augmented generation to improve the model's robustness.
Outcome: The proposed model is more robust and performs significantly better in a more challenging setting over strong baselines.
Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing reinforcement learning systems lack verifiable reward mechanisms for long-form question answering . current systems lack reliable long-term answers due to lack of factual content .
Approach: They propose a framework for reinforced verifiable informativeness optimization . it defines informativeness as measurable and externally verifier objective for RL .
Outcome: Experiments show that RioRAG achieves higher factual recall and faithfulness . the proposed framework is based on a framework that uses nugget-centric verification with cross-source checks .
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback (2026.findings-acl)

Copied to clipboard

Challenge: Recent research emphasizes the generation of high-quality feedback that provides justification and actionable guidance.
Approach: They propose an LLM-based framework for evaluating LLM feedback along three dimensions: specificity, helpfulness, and validity.
Outcome: The proposed framework evaluates LLM-generated feedback along three dimensions: specificity, helpfulness, and validity.
Feedback to Reasoning: LLM-Assisted Molecular Optimization with Domain Feedback and Historical Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for molecular optimization do not leverage domain feedback and historical knowledge with reasoning traces and chemical insights.
Approach: They propose a conversational molecular optimization pipeline that enables LLMs to accumulate and retrieve past actions, rationales, and feedback.
Outcome: The proposed framework transforms LLMs from passive text generators into agentic experts that learn both actions and reasoning from experience.
Beyond Screenshots: Evaluating VLMs’ Understanding of UI Animations (2026.findings-acl)

Copied to clipboard

Challenge: Recent studies of Vision Language Models (VLMs) for UI understanding have focused primarily on static screenshots, leaving it unclear how well these models handle dynamic UI animations.
Approach: They evaluate UI animation models' ability to perceive animation effects and interpret animation meaning . they use motion, context, and perceptual cues to probe factors affecting VLM performance .
Outcome: The proposed model can detect primitive motion, but its interpretation is inconsistent . the proposed model is based on 300 annotated UI animation videos .
Multi-Faceted Self-Consistent Preference Alignment for Query Rewriting in Conversational Search (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to rewrite ambiguous queries ignore feedback from query rewriting, passage retrieval and response generation in the rewritten process.
Approach: They propose to construct self-consistent preference alignment data to generate more diverse rewritten queries.
Outcome: The proposed method is effective in both in- and out-of-distribution scenarios.
Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System Improvement (2022.emnlp-main)

Copied to clipboard

Challenge: Using simulated feedback, our system (called TeachMe) continually improves with time, and without model retraining.
Approach: They propose to augment a QA model with a dynamic memory of user feedback, containing user-supplied corrections toerroneous model beliefs that users identify during interaction.
Outcome: The proposed system improves with time and without model retraining, and with real users, by 15% on a hidden test set after teaching.
Logic-driven Indirect Supervision: An Application to Crisis Counseling (2023.acl-long)

Copied to clipboard

Challenge: Text-based crisis counseling services are increasingly adopted by people seeking confidential mental health support.
Approach: They propose an inexpensive method that exploits declaratively stated structural dependencies between both levels of annotation to improve utterance modeling.
Outcome: The proposed method improves utterance modeling by 3.5% over a strong multitask baseline.
What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal Context (2026.acl-long)

Copied to clipboard

Challenge: Existing preference-alignment approaches rely on binary pairwise comparisons, overlooking preference intensity and temporal context.
Approach: They propose a unified preference optimization framework that maps both explicit and implicit feedback into a common preference signal and constructs adaptive reward margins that jointly account for preference intensity and interaction recency.
Outcome: The proposed framework outperforms state-of-the-art recommendations while maintaining behavioral patterns aligned with human decision-making.
Learning from a Friend: Improving Event Extraction via Self-Training with Feedback from Abstract Meaning Representation (2023.findings-acl)

Copied to clipboard

Challenge: Existing data scarcity hinders the progress of event extraction, authors say . ACE-052 has 10 of the 33 event types with less than 80 annotations, authors claim .
Approach: They propose a self-training with feedback framework that leverages large-scale unlabeled data to acquire feedback for each new event prediction from the unlabed data.
Outcome: The proposed framework improves event extraction models even when unlabeled data are unavailable.
LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback (2025.findings-acl)

Copied to clipboard

Challenge: Large Action Models (LAMs) face challenges due to the need for high-quality training data, especially for multi-steps tasks that involve planning, executing tool calls, and responding to feedback.
Approach: They propose a framework for online exploration of agentic tasks with high-quality feedback . they use a dynamic task query generator and an extensive collection of tools to create a high-level feedback environment for LLM Agents.
Outcome: The proposed framework achieves 49.3% performance improvement over baselines on toolbench and CRMArena.
PerfCoder: Large Language Models for Interpretable Code Performance Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have advanced automatic code generation, but their ability to produce high-performance code remains limited.
Approach: They propose a family of large language models that generate performance-enhanced code through interpretable and customized optimization strategies.
Outcome: The proposed model outperforms existing models on the PIE code performance benchmark and produces interpretable feedback that can guide larger LLMs in a planner–optimizer workflow.
Uncovering Intervention Opportunities for Suicide Prevention with Language Model Assistants (2026.acl-long)

Copied to clipboard

Challenge: Using language models, annotators can help develop novel suicide interventions . 85% of cases where LM predictions disagree with existing annotations are analyzed .
Approach: They propose a human-in-the-loop algorithm that leverages language models as an assistant to annotators and experts to facilitate data-driven insights from NVDRS data.
Outcome: The proposed algorithm can be used to support the development of novel suicide interventions . it finds that LM predictions match existing data annotations about 85% of the time .
MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics (2026.acl-long)

Copied to clipboard

Challenge: MalruleLib is a learning-science-grounded framework that translates documented misconceptions into executable procedures and generates step-by-step traces of malrule-consistent student reasoning.
Approach: They propose a learning-science-grounded framework that translates documented misconceptions into executable procedures and generates step-by-step traces of malrule-consistent student reasoning.
Outcome: The framework translates misconceptions into executable procedures and generates step-by-step traces of malrule-consistent student reasoning.
Small But Funny: A Feedback-Driven Approach to Humor Distillation (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been used to transfer knowledge from LLMs to smaller, smaller language models (SLMs).
Approach: They propose to assign a dual role to the LLM as a “teacher” generating data, as well as evaluating the student’s performance.
Outcome: The proposed approach narrows the performance gap between LLMs and larger models by incorporating feedback into the data.
Divide-Verify-Refine: Can LLMs Self-align with Complex Instructions? (2025.findings-acl)

Copied to clipboard

Challenge: Existing research shows LLMs struggle with complex instructions involving multiple constraints.
Approach: They propose a framework to divide complex instructions into single constraints and prepare appropriate tools to verify responses.
Outcome: The proposed framework doubles Llama3.1-8B’s constraint adherence and triples Mistral-7B’ s performance.
ACE: A LLM-based Negotiation Coaching System (2024.emnlp-main)

Copied to clipboard

Challenge: The rapid progress of LLMs has led to the development of more sophisticated AI tutoring systems.
Approach: They develop an LLM-based assistant for coaching negotiation that provides users with targeted feedback for improvement.
Outcome: The proposed system improves negotiation performance significantly compared to a system that doesn’t provide feedback and one which uses an alternative method.
Retrieval Enhanced Feedback via In-context Neural Error-book (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for learning from errors lack a structured framework for analyzing and mitigating errors, especially in Multimodal Large Language Models (MLLMs).
Approach: They propose a teacher-student framework that systematically structures errors to deliver targeted feedback for multimodal reasoning.
Outcome: The proposed framework improves inference efficiency, token usage, and scalability by building a query-based structure that prioritizes visual information, diagnoses failure points, and guides corrective actions.
SummIt: Iterative Text Summarization via ChatGPT (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing text summarization systems generate summaries in a single step, but are often inadequate due to the issue of hallucination and the lack of accuracy.
Approach: They propose an iterative text summarization framework based on large language models like ChatGPT that refines the generated summary iterativly through self-evaluation and feedback.
Outcome: The proposed framework refines the generated summary iteratively through self-evaluation and feedback, closely resembling the iteration humans undertake when drafting and revising summaries.
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing Vid-LLMs lack robust mechanisms for maintaining grounded spatiotemporal beliefs under conversational feedback.
Approach: They propose a negation-based gaslighting evaluation framework and introduce a benchmark to investigate spatiotemporal sycophancy.
Outcome: The proposed framework evaluates state-of-the-art Vid-LLMs across video understanding tasks.
Quality Scoring of Source Words in Neural Translation Models (2022.emnlp-main)

Copied to clipboard

Challenge: Recent approaches to improving word-level quality scores on input source sentences require training special word-scoring models or require repeated invocation of the translation model.
Approach: They propose to reason how well each word is explained by the target sentence as against the source language model and use it to translate into an unfamiliar target language.
Outcome: The proposed method provides up to five points higher F1 scores and is significantly faster than the state of the art methods on three language pairs.
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation (2024.findings-acl)

Copied to clipboard

Challenge: Existing vision-and-language navigation methods do not incorporate environmental feedback into their decision-making processes.
Approach: They propose a framework that incorporates environmental feedback into decision-making and a 3D simulator that renders realistic scenarios using Unreal Engine 5.
Outcome: The proposed framework outperforms existing vision-and-language navigation methods in a zero-shot multi-task setting by 28.1% on average.
Grasping the Essentials: Tailoring Large Language Models for Zero-Shot Relation Extraction (2024.emnlp-main)

Copied to clipboard

Challenge: Existing Relation extraction models require extensive annotated training data, which is costly and labor-intensive to collect.
Approach: They propose a new zero-shot RE task where only relation definitions are provided instead of seen-unseen relation instances.
Outcome: The proposed task significantly improves cost-effective zero-shot performance by large margins.
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? (2025.acl-long)

Copied to clipboard

Challenge: Pairwise feedback is widely used to evaluate and provide feedback to large language models (LLMs).
Approach: They propose a tool-using agentic system to provide higher quality feedback on three challenging response domains: long-form factual, math and code tasks.
Outcome: The proposed system can provide higher quality pairwise comparisons on three domains, independent of the LLM’s internal knowledge and biases.
Mutual-Taught for Co-adapting Policy and Reward Models (2025.acl-long)

Copied to clipboard

Challenge: Experimental results show that this iterative approach leads to consistent improvements in both the policy model and reward model.
Approach: They propose a method that iteratively improves both the policy model and reward model without requiring additional human annotation.
Outcome: The proposed method improves both the policy model and reward model without human annotation.
InfFeed: Influence Functions as a Feedback to Improve the Performance of Subjective Tasks (2024.lrec-main)

Copied to clipboard

Challenge: InfFeed uses influence functions to compute the influential instances for a target instance.
Approach: They propose an apparatus that uses influence functions to compute the influential instances for a target instance.
Outcome: The proposed model outperforms the state-of-the-art baselines by 4% for hate speech classification, 3.5% for stance classification, and 3% for irony and 2% for sarcasm detection.
LCR-RAG: Enhancing Logical Consistency in Retrieval-Augmented Generation via Neuro-symbolic Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) is widely used to ground large language models in external knowledge and improve factual accuracy.
Approach: They propose a framework that integrates neuro-symbolic verification with reinforcement learning to optimize logical consistency.
Outcome: The proposed framework outperforms strong RAG baselines on hotpotQA, ASQA, and TriviaQA.
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: The application scope of large language models (LLMs) is expanding . however, evaluating whether models can respond to user feedback has not been thoroughly analyzed.
Approach: They propose a benchmark to assess whether large language models can respond to refuting feedback and adhere to user demands throughout the conversation.
Outcome: The proposed benchmark covers tasks such as question answering, machine translation, and email writing.
Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving (2025.findings-acl)

Copied to clipboard

Challenge: Recent methods for AI reasoning require applying variants of reinforcement learning (RL) on rolled out trajectories, even for step-wise rewards, or large quantities of human-annotated trajectory data.
Approach: They propose a verifier-in-the-loop design that uses an automated verifier to give intermediate feedback at each step of the reasoning process.
Outcome: The proposed model improves on the Automatic Theorem Proving task using Lean as the verifier.
CPT-Agent: A Cognitive Process Theory-driven Framework for Student Simulation in Writing Development (2026.acl-long)

Copied to clipboard

Challenge: Existing LLMs model overly capable learners who over-apply feedback, resulting in pedagogically implausible behavior.
Approach: They propose a framework that decouples cognitive ability from writing proficiency and models their interaction during writing and revision.
Outcome: The proposed model produces distinguishable proficiency levels and is consistent with instructional theories.
Re-ReST: Reflection-Reinforced Self-Training for Language Agents (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to fine tune language agents with reasoning-action trajectories require high-quality model-generated samples, which are hard to obtain for challenging language agent tasks.
Approach: They propose a method to employ reflection during inference without ground-truth feedback to improve agents more autonomously.
Outcome: The proposed method improves self-training performance on open-source language agents by 7.6% and 14.1% respectively.
Prospector: Improving LLM Agents with Self-Asking and Trajectory Ranking (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing LLMs are limited in their ability to incorporate feedback from an environment.
Approach: They propose an LLM agent that consists of an Actor and a Critic.
Outcome: The proposed agent outperforms existing LLMs on benchmark environments and shows that it can generate diverse trajectories and pick the most rewarding trajectory.
CESRec: Constructing Pseudo Interactions for Sequential Recommendation via Conversational Feedback (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing Sequential Recommendation Systems (SRS) rely on collaborative filtering signals and fail to capture real-time user preferences.
Approach: They propose a framework that integrates the long-term preference modeling of SRS with the real-time preference elicitation of CRS.
Outcome: The proposed framework integrates the long-term preference modeling of SRS with the real-time preference elicitation of CRS.
AskQE: Question Answering as Automatic Evaluation for Machine Translation (2025.findings-acl)

Copied to clipboard

Challenge: Existing MT error detection and quality estimation (QE) techniques do not address this practical scenario.
Approach: They propose a question generation and answering framework that detects critical MT errors and provides actionable feedback to help users decide whether to accept or reject MT outputs even without the knowledge of the target language.
Outcome: The proposed framework has higher Kendall’s Tau correlation and decision accuracy with human ratings compared to other QE metrics.
Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple Rewards (2024.emnlp-main)

Copied to clipboard

Challenge: Existing reinforcement learning (RL) applications in AES are limited to classification models despite associated performance degradation.
Approach: They propose to integrate actual evaluation schemes into the training process by designing QWK-based rewards with a mean-squared error penalty for multi-trait AES.
Outcome: The proposed scoring-aware multi-reward reinforcement learning integrates actual evaluation schemes into the training process.
LM2: A Simple Society of Language Models Solves Complex Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that providing guidance via decomposing the original question into multiple subproblems elicits more robustness in LLM reasoning.
Approach: They propose a language-based decomposition, solution and verification framework that modularizes the decomposer, solution, and verification into three different language models.
Outcome: The proposed model outperforms existing methods on in- and out-domain reasoning problems, outperforming the best baselines by 8.1% on MATH, 7.71% on JEEBench, and 9.7% on MedQA problems.
Learning to love diligent trolls: Accounting for rater effects in the dialogue safety task (2023.findings-emnlp)

Copied to clipboard

Challenge: Xu et al., 2018: chatbots generate offensive utterances, which must be avoided . he proposes a solution that can learn from user interactions in a way that is robust to trolls .
Approach: They propose a method to learn from user feedback in a way that is robust to trolls . they propose multiple users rate each utterance, then perform latent class analysis to infer correct labels.
Outcome: The proposed solution can infer training labels with high accuracy when trolls are consistent, even when a majority are trolled.
Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents (2025.emnlp-main)

Copied to clipboard

Challenge: Visual Language Models (VLMs) have shown strong performance in tasks like radiology report generation but struggle with hallucinations, vague descriptions, Inconsistent logic and poor localization.
Approach: They propose a framework for medical visual reasoning based on Visual Guidance and Self-Reward paradigms and Monte Carlo Tree Search to improve the model's visual reasoning capabilities.
Outcome: The proposed framework outperforms existing models on multiple medical VQA benchmarks.
Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Reward hacking is a problem in reinforcement learning where the ability to specify the desired behavior of a reward function is difficult.
Approach: They propose to use feedback as a potential-based shaping function to solicit and apply feedback from large language models to improve convergence speed and policy returns.
Outcome: The proposed method improves convergence speed and policy returns over baselines even with significant ranking errors and eliminates the need for complex post-processing of reward functions.
LLM-Evolve: Evaluation for LLM’s Evolving Capability on Benchmarks (2024.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for large language models evaluate LLMs on i.i.d. tasks, overlooking their ability to learn iteratively from past experiences.
Approach: They propose a framework which extends established benchmarks to sequential problem-solving settings and provides feedback after each round to build a demonstration memory that the models can query in future tasks.
Outcome: The proposed framework can improve performance of LLMs by learning from past interactions and improve models' performance over time.
Selective Prompting Tuning for Personalized Conversations with LLMs (2024.findings-acl)

Copied to clipboard

Challenge: Personalization in conversational AI requires persona profiles and contextual understanding to create meaningful conversations.
Approach: They propose a method that softly prompts LLMs for personalized conversations in a selective way.
Outcome: The proposed approach improves response diversity by up to 90% on the CONVAI2 dataset.
FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge Flow (2026.acl-long)

Copied to clipboard

Challenge: FlowSearch is a multi-agent framework that actively constructs and evolves a dynamic structured knowledge flow to drive subtask execution and reasoning.
Approach: They propose a multi-agent framework that actively constructs and evolves a dynamic structured knowledge flow to drive subtask execution and reasoning.
Outcome: The proposed framework achieves competitive performance on GAIA, HLE, GPQA and TRQA benchmarks and is available to download.
RePair: Automated Program Repair with Process-based Feedback (2024.findings-acl)

Copied to clipboard

Challenge: Commercial-scale language models (LMs) have taken APR to unprecedented levels, but they are limited by parameters and humans interact with them through explicit prompts.
Approach: They propose a method that utilizes process supervision to improve program repair by allowing users to input feedback from compilers and test cases.
Outcome: The proposed method outperforms large outcome-based generation methods and is inspired by strategies used in programming competitions.
Optimizing Reasoning for Text-to-SQL with Execution Feedback (2025.findings-acl)

Copied to clipboard

Challenge: Large language models excel in many reasoning tasks, but their ability to leverage Chain-of-Thought (CoT) reasoning remains underexplored.
Approach: They propose a framework that iteratively optimizes open-source LLMs by combining CoT reasoning with off-policy and on-poly DPO, relying solely on execution accuracy as feedback.
Outcome: The proposed framework improves execution accuracy on BIRD and Spider datasets.
Multimodal Behaviour in an Online Environment: The GEHM Zoom Corpus Collection (2024.lrec-main)

Copied to clipboard

Challenge: Several studies have discussed pros and cons of videoconferencing for group meetings, international conference organisation and teaching.
Approach: They propose to use 12 video recordings of Zoom meetings held in English by an international group of researchers from September 2021 to March 2023 to study group communication in a reallife setting.
Outcome: The proposed corpus was developed under the auspices of the international network on Gesture and Head Movement in Language (GEHM) it shows that the participants' speech transcription and visual keypoint values can be visualised to see how gestural behaviour supports feedback words during the interaction.
See2Refine: Vision-Language Feedback Improves LLM-Based eHMI Action Designers (2026.acl-long)

Copied to clipboard

Challenge: External Human-Machine Interfaces (eHMIs) are emerging as promising solutions to address this communication gap.
Approach: They propose a framework that uses vision-language models (VLMs) for perceptual evaluation as automated visual feedback to improve an LLM-based eHMI action designer.
Outcome: The proposed framework outperforms prompt-only LLM designers and manually specified baselines in three eHMI modalities and multiple LLM model sizes.
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that AI-assisted research methods can improve research efficiency . a closed-loop framework is used to enhance the automation level of scientific research .
Approach: They propose a closed-loop LLM-driven framework to enhance the automation level of scientific research.
Outcome: The proposed framework improves the efficiency of scientific research by improving data analysis, accelerating computation, and fostering novel idea generation.
DUET: Joint Exploration of User–Item Profiles in Recommendation System (2026.findings-acl)

Copied to clipboard

Challenge: Existing LLMs are opaque and difficult to interpret, resulting in limited interpretability.
Approach: They propose an interaction-aware profile generator that jointly produces user and item profiles conditioned on both user history and item evidence.
Outcome: The proposed model outperforms baselines on three real-world datasets.
Diagnosing Failures in Large Language Models’ Answers: Integrating Error Attribution into Evaluation Framework (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation models lack error attribution capability due to their proprietary nature.
Approach: They propose a misattribution framework with 6 primary and 15 secondary categories to facilitate in-depth analysis.
Outcome: The proposed framework is based on a dataset specifically designed for error attribution, along with the corresponding scores and feedback.
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for instruction-following lack data coverage and oversimplified pairwise evaluation paradigms that misalign with model optimization scenarios.
Approach: They propose a meta-evaluation benchmark for instruction-following that covers diverse instruction and constraint types and a preference graph for each instruction.
Outcome: Extensive experiments on IF-RewardBench show that the proposed benchmark achieves a stronger positive correlation with downstream task performance compared to existing benchmarks.
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have demonstrated significant potential in single-turn reasoning tasks.
Approach: They propose a tool-memory based self-evolving agentic framework that integrates planning with execution.
Outcome: The proposed framework is able to extract explicit knowledge from historical data and leverage inter-trajectory correlations to densify reward signals.
Towards Aligning Language Models with Textual Feedback (2024.emnlp-main)

Copied to clipboard

Challenge: Using textual feedback, language models can be trained to learn from textual inputs.
Approach: They propose an approach that aligns language models with user preferences expressed in text.
Outcome: The proposed approach outperforms PPO on toxicity reduction, summarization, and dialog response tasks while achieving the same performance with only 20% of the samples.
AMPO: Automatic Multi-Branched Prompt Optimization (2024.emnlp-main)

Copied to clipboard

Challenge: Existing prompt engineering techniques are limited to producing single flow instructions, struggling with handling diverse patterns.
Approach: They propose an automatic prompt optimization method that iteratively develops a multi-branched prompt using failure cases as feedback.
Outcome: The proposed method achieves the best results across five tasks and demonstrates significant optimization efficiency due to adoption of a minimal search strategy.
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks (2025.acl-long)

Copied to clipboard

Challenge: Inference-Time Scaling is critical to the success of recent models such as OpenAI o1 and DeepSeek R1 . however, many techniques require tasks to have answers that can be verified .
Approach: They use data to train dedicated Feedback and Edit Models capable of inference-time scaling for open-ended tasks.
Outcome: The proposed model can reach SoTA performance on Arena Hard at 92.7 as of 5 Mar 2025.
Coffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Code (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have made great progress in code generation, however, they still produce errors.
Approach: They propose a RL environment that provides feedback on code editing by analyzing the performance of the revised code in unit tests.
Outcome: The proposed model outperforms baselines in enhancing open-source code LLMs’ code editing, making them comparable with closed-source LLM.
Do Language Models Have Semantics? On the Five Standard Positions (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are trained to solve the so-called cloze task . solving clozing tasks is essentially a memorization task, says a recent study .
Approach: They propose to use five positions to determine whether large language models exhibit semantic understanding . large language model is trained to solve the so-called cloze task .
Outcome: The proposed theory is based on a pairwise comparison of five positions on semantic understanding in large language models and chatbots.
Iterative Repair with Weak Verifiers for Few-shot Transfer in KBQA with Unanswerability (2025.findings-acl)

Copied to clipboard

Challenge: Existing models for KBQA with unanswerable questions are inadequate for real-world applications.
Approach: They propose a task of few-shot transfer for KBQA with unanswerable questions that extends FuSIC-KBQA to include feedback for unanswered questions.
Outcome: The proposed model outperforms suitable adaptations of multiple LLM-based and supervised SoTA models on the task while establishing a new performance for answerable few-shot transfer as well.
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback (2025.findings-acl)

Copied to clipboard

Challenge: Existing static benchmarks that measure task performance often rely on a simple input-output configuration.
Approach: They propose an evaluation pipeline that evaluates code models with different feedback types in an interactive setting.
Outcome: The proposed evaluation pipeline compares model-user collaboration with static benchmarks by obfuscating inputs to a simulated user.
CEAES: Bidirectional Reinforcement Learning Optimization for Consistent and Explainable Essay Assessment (2025.acl-long)

Copied to clipboard

Challenge: Current automated essay quality assessment systems treat score prediction and feedback generation as separate tasks.
Approach: They propose a bidirectional reinforcement learning framework that jointly optimizes score prediction and feedback generation.
Outcome: The proposed framework outperforms current state-of-the-art models in both scoring and feedback quality.
Steering Away from Refusal: A Black-box Jailbreak Method Based on First-Token Distribution (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to analyze black-box jailbreaks lack direct optimization signals to refine adversarial prompts.
Approach: They propose a distribution-jailbreak attack method that selects effective jailbreak templates and iteratively optimizes adversarial suffixes by maximizing the KL divergence from the standard refusal distribution.
Outcome: The proposed method achieves state-of-the-art Attack Success Rate (ASR) on all tested open-source models and delivers over 94% ASR on GPT-4.1.
Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries (2025.emnlp-main)

Copied to clipboard

Challenge: a typical human-assistant conversation is lengthy and shows significant diversity in topics, intents, and requirements across turns.
Approach: They propose a framework that leverages pertinent linguistic concepts of dialog-acts and maxims to improve the accuracy of LLM-judges on preference data with complex, multi-turn conversational context.
Outcome: The proposed framework improves on 4 challenging datasets showing that humans frequently change their intents from one turn of the conversation to the next.
Understanding the Dark Side of LLMs’ Intrinsic Self-Correction (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that LLMs’ intrinsic self-correction fails without oracle labels as feedback.
Approach: They propose to use one simple task and three complex tasks with state-of-the-art LLMs like ChatGPT, Llama, and DeepSeek to interpret LLM's intrinsic self-correction.
Outcome: The proposed methods reveal the dark side of LLMs’ intrinsic self-correction for different tasks, especially for those failure cases.
AgentDrug: Utilizing Large Language Models in an Agentic Workflow for Zero-Shot Molecular Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Molecular optimization is a fundamental task in drug discovery.
Approach: They propose an agentic workflow that leverages LLMs in a structured refinement process to achieve significantly higher accuracy.
Outcome: The proposed workflow improves on single- and multi-property optimization tasks under loose and strict thresholds.
ProToM: Promoting Prosocial Behaviour via Theory of Mind-Informed Feedback (2026.findings-acl)

Copied to clipboard

Challenge: ProToM provides targeted, context-sensitive feedback to individual agents, achieving a higher success rate, shorter task completion times, and is consistently preferred by human users.
Approach: They propose a Theory of Mind-informed facilitator that provides targeted, context-sensitive feedback to individual agents.
Outcome: The proposed system provides targeted, context-sensitive feedback to promote prosocial behaviour, even when not directly aligned with one’s own goals.
Leveraging Unpaired Feedback for Long-Term LLM-based Recommendation Tuning (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study highlights unpaired feedback as a key challenge for long-term LLM-based recommenders . unpaired user feedback is crucial for improving LLMs in dynamic user environments, authors say .
Approach: They propose a framework that incorporates unpaired feedback into LLMs to improve long-term recommendation performance.
Outcome: The proposed framework improves long-term recommendation performance by incorporating unpaired feedback without requiring paired supervision.
Understanding Impact of Human Feedback via Influence Functions (2025.acl-long)

Copied to clipboard

Challenge: In reinforcement learning from human feedback, human feedback can be noisy, inconsistent or biased . this variability can lead to misaligned reward signals, potentially causing unintended side effects .
Approach: They propose an approximation method that measures the impact of human feedback on the performance of reward models.
Outcome: The proposed method detects common labeler biases in human feedback datasets and guides labelers in refining their strategies to better align with expert feedback.
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization (2025.acl-long)

Copied to clipboard

Challenge: Existing studies focus on building text-only agents in synthetic environments where the reward signals are clearly defined.
Approach: They propose a multimodal web agent that can autonomously conduct real-world exploration and improve itself after each iteration.
Outcome: The proposed agent improves itself after each iteration, demonstrating strong performance across multiple test sets.
SynthFix: Adaptive Neuro-Symbolic Code Vulnerability Repair (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) struggle with complex semantic and structural correctness required for automated code repair.
Approach: They propose a hybrid neural-symbolic framework that unifies code synthesis with compiler-informed symbolic feedback to improve LLM-based vulnerability repair.
Outcome: The proposed framework improves code repair accuracy and efficiency over strong SFT and RFT training strategies on the FixJS and CodeFlaws benchmarks.
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Recent work on LLM-as-a-Judge has reported higher correlations with human judgments due to its static nature.
Approach: They propose a framework that leverages multi-turn interactions where the LLM interviewer actively provides feedback on responses and poses follow-up questions to the evaluated LLM.
Outcome: The proposed framework evaluates six models on reasoning, factuality and instruction-following tasks.
Training Language Model to Critique for Better Refinement (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have remarkable evaluation and critique capabilities, providing insightful feedback and identifying flaws in various tasks.
Approach: They propose a framework to train critic models using refinement signals to generate feedback loops where critiques guide the model in refining its responses.
Outcome: The proposed framework outperforms traditional methods and open-source models in terms of critique quality and refinement outcomes.
Enhancing Goal-oriented Proactive Dialogue Systems via Dynamic Multi-dimensional Consistency Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing work on goal-oriented proactive dialogue systems failed to address the multi-dimensional consistency issue between generated responses and key contextual elements.
Approach: They propose a Dynamic Multi-dimensional Consistency Reinforcement Learning framework which measures the impact of each consistency dimension on overall dialogue quality and provides feedback to improve response quality.
Outcome: The proposed framework significantly improves the consistency of generated responses on two datasets.
CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model (2025.acl-long)

Copied to clipboard

Challenge: Existing multimodal models for motion related tasks have shown significant progress.
Approach: They propose a reference-based model that analyzes the differences between a learner’s motion and a physical reference under temporal and physical aspects.
Outcome: The proposed model outperforms GPT-4o on figure skating and boxing by 31.6% and 58.3% respectively.
Feedback Adaptation for Retrieval-Augmented Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation protocols focus on overall accuracy and fail to capture how systems adapt after feedback is introduced.
Approach: They propose to use feedback adaptation as a problem setting for RAG systems . they propose a minimal inference-time instantiation that incorporates feedback without retraining .
Outcome: The proposed evaluations show that training-based approaches exhibit a trade-off between delayed correction and reliable adaptation.
Theorem Prover as a Judge for Synthetic Data Generation (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that large language models are increasingly capable of tackling mathematical problems.
Approach: They propose an approach that iteratively refines theorem prover formalisation to mitigate errors.
Outcome: The proposed method increases execution rate on the Lean prover from 60% to 87%, while human annotation is replaced with theorem prover feedback.
Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches for personalizing large language models require modifying parameters.
Approach: They propose a lightweight approach to personalizing large language models via retrieval augmentation . relevance serves as an unreliable proxy for utility, they argue .
Outcome: The proposed framework outperforms strong heuristic and retrieval-augmented baselines on nine personalization tasks.
GPT-4 as a Homework Tutor Can Improve Student Engagement and Learning Outcomes (2025.acl-long)

Copied to clipboard

Challenge: a recent study has shown that homework is never graded or is done superficially.
Approach: They propose a prompting strategy that enables GPT-4 to conduct interactive homework sessions for high school students learning English as a second language.
Outcome: The proposed solution improves homework in high school students learning English as a second language with minimal effort in content preparation, one of the key challenges of alternative methods.
Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop (2026.acl-long)

Copied to clipboard

Challenge: Existing synthetic training loops for large language models cause performance drops and induce emerging biases . a large amount of generated content is posted to coding platforms, social media platforms and other platforms on the internet .
Approach: They propose a self-consuming retraining loop where models are trained on their own outputs . they use a control loop to isolate and analyze feedback-driven bias evolution .
Outcome: The proposed model increases preference bias and decreases disparate bias.
ThinkTuning: Instilling Cognitive Reflections without Distillation (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning.
Approach: They propose a GRPO-based interactive training approach that augments the rollouts of a student model with the guidance of . a teacher poses a problem, lets the student try an answer, then gives corrective feedback–enough to point the mind in the right direction and then show the correct solution.
Outcome: The proposed method shows 3.69% improvement over zero-shot baselines and 2.08% and 3.99% improvement over the vanilla-GRPO baselines.
Automated Knowledge Component Generation and Interpretable Knowledge Tracing in Coding Problems (2026.findings-acl)

Copied to clipboard

Challenge: Existing solutions to automate KC generation and tagging for open-ended programming problems are highly labor-intensive and prone to bias and errors.
Approach: They propose an automated pipeline for KC generation and tagging for open-ended programming problems using large language models.
Outcome: The proposed method outperforms existing ones and outperfies human-written KCs on future student response prediction.
QUIDS: Query Intent Description for Exploratory Search via Dual Space Modeling (2025.emnlp-main)

Copied to clipboard

Challenge: Using QUIDS, we generate user-facing query intent descriptions that surface what the search engine likely inferred the query to mean based on post-retrieval evidence.
Approach: They propose a method that leverages dual-space contrastive learning to isolate intent-relevant information while suppressing irrelevant content.
Outcome: The proposed method outperforms state-of-the-art methods across ROUGE, BERTScore, and human/LLM evaluations.
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents.
Approach: They propose to integrate human-provided information, feedback, or control into the agent system to enhance system performance, reliability, and safety.
Outcome: The proposed systems improve system performance, reliability, and safety by integrating human-provided information, feedback, or control into the agent system.
ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback (2026.findings-acl)

Copied to clipboard

Challenge: Unlike chatbots, autonomous agents act directly on external environments, making tool invocation safety critical for reliable deployment.
Approach: They develop a benchmark for step-level tool invocation safety detection in LLM agents and a guardrail model that proactively detects unsafe tool invoking actions before execution using multi-task reinforcement learning.
Outcome: The proposed model reduces harmful tool invocations of ReAct-style agents by 65% on average and improves benign task completion by 10% under prompt injection attacks.
Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) can call tools effectively, but they remain brittle in multi-turn execution.
Approach: They propose a framework that converts execution errors into on-policy corrective supervision within the RL training loop.
Outcome: The proposed framework improves the error recovery rate of Qwen3-8B by 5.7% absolute and overall accuracy by 4.0% on BFCL v4 Multi-Turn.
Feedback Is The Key for Automated Survey Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a promising foundation for literature surveys, but guiding them to generate accurate, reliable content remains a fundamental challenge.
Approach: They propose a feedback-driven framework that incorporates feedback across three dimensions: outline feedback for structural clarity, citation feedback for evidence validation, and content feedback for readability and analytical depth.
Outcome: The proposed framework significantly improves both citation and content quality, demonstrating feedback as the critical mechanism for automatic survey generation.
CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment (2026.acl-long)

Copied to clipboard

Challenge: Recent schema-based TOD frameworks improve generalization by decoupling task logic from language understanding, but their reliance on neural or generative models obscures how task schemas influence behaviour and hence impair interpretability.
Approach: They propose a framework that converts a predefined task schema to a structured heterogeneous graph and then to popular programmatic LLM guardrailing code, such as NVIDIA’s Colang.
Outcome: The proposed framework achieves state-of-the-art performance on the widely used benchmark datasets while providing inherent interpretability in the design.
DarwinTOD: LLM-Driven Lifelong Self-evolution for Task-oriented Dialog Systems (2026.acl-long)

Copied to clipboard

Challenge: Continual learning approaches fail to achieve autonomy lifelong improvement in dynamic environments . current task-oriented dialog systems are static, unable to learn from ongoing interactions .
Approach: They propose a lifelong self-evolving dialog framework that integrates evolutionary computation and LLM driven self-improvement into a single framework.
Outcome: The proposed framework surpasses state-of-the-art methods and exhibits continuous performance gains throughout evolution.
Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs (2026.acl-long)

Copied to clipboard

Challenge: Existing datasets for evaluating text-to-image generation focus mostly on real-life images, which poses challenges for assessing academic figure generation given real scientific captions.
Approach: They propose a dataset that first provides a Holistic Evaluation for Academic caption-to-Figure Generation (HE4AFG) they collect real figure captions from 8 scientific domains and generate 3,900 evaluation samples .
Outcome: The proposed model provides high-quality human ratings in terms of three aspects—scientific aesthetic (SA), topic relevance (TR), and attribute correctness (AC).
FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) often produce factually incorrect responses.
Approach: They propose a new method that adapts across domains without retraining and leverages structured feedback to generate a correction.
Outcome: The proposed method outperforms baseline methods on a VELI5 dataset and several popular long-form factuality datasets.
Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have enhanced capabilities in complex reasoning through step-by-step trace generation.
Approach: They propose a generative verifier that dynamically allocates compute between rapid fast thinking and deliberative slow thinking.
Outcome: The proposed solution outperforms GenPRM-32B on ProcessBench while requiring 2.3x fewer TFLOPS and 15x less training data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations