Papers with Feedback

15 papers
NextGen AML: Distributed Deep Learning based Language Technologies to Augment Anti Money Laundering Investigation (P18-4)

Copied to clipboard

Challenge: Money laundering (AML) is the process of transferring criminal and illegal proceeds into ostensibly legitimate assets.
Approach: They propose a framework that uses deep learning to augment AML monitoring and investigation . money laundering is the process of transferring criminal and illegal proceeds into ostensibly legitimate assets .
Outcome: The proposed framework reduces time and cost by 30% compared to existing methods . money laundering is the world's third largest "industry"
An Online Readability Leveled Arabic Thesaurus (2020.coling-demos)

Copied to clipboard

Challenge: a small minority of dictionaries specify the readability level of their words, let alone their lexical relations with other words.
Approach: They propose to use Arabic lemmas, roots, English glosses, related Arabic words and phrases to provide a readability leveled Arabic thesaurus interface.
Outcome: The proposed system provides the user with lemmas, roots, English glosses, related Arabic words and phrases, and readability on a five-level readability scale.
FEAT-writing: An Interactive Training System for Argumentative Writing (2025.coling-demos)

Copied to clipboard

Challenge: Argumentative writing is a critical skill for academic success, but many students struggle to develop these skills.
Approach: They developed an online system that provides students with automated feedback and exercises for argumentative writing.
Outcome: The proposed system improves argumentative writing quality among native English speakers and english-as-a-foreign-language learners.
A Dataset for Investigating the Impact of Feedback on Student Revision Outcome (2020.lrec-1)

Copied to clipboard

Challenge: Despite numerous studies on the kinds of feedback that can best promote learning, this question remains an open debate in the area of Second Language Acquisition (SLA).
Approach: They annotate a corpus of student-written sentences with teacher feedback provided for the errors.
Outcome: The proposed annotation scheme and the teacher feedback dataset are based on student-written sentences in their original and revised versions with teacher feedback provided for the errors.
Synthesizing Human Gaze Feedback for Improved NLP Performance (2023.eacl-main)

Copied to clipboard

Challenge: Prior work on eye tracking and NLP reveals that human scanpaths can aid in understanding and performance of NLP models.
Approach: They propose a model for generating human scanpaths over text that approximates meaningful cognitive signals in human gaze patterns.
Outcome: The proposed model can approximate meaningful cognitive signals in human gaze patterns.
A Survey of Post-Training Scaling in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated proficiency in understanding and generating human natural languages.
Approach: They propose a framework for scaling large language models using supervised fine-tuning, RLxF and test-time compute methodologies.
Outcome: The proposed model can be used to understand and generate human natural languages.
Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data (2023.emnlp-main)

Copied to clipboard

Challenge: Despite the promising potential of chat models, they are only accessible through restricted APIs, creating barriers for new research and progress in the field.
Approach: They propose a pipeline that can automatically generate a high-quality multi-turn chat corpus by leveraging ChatGPT to engage in a conversation with itself.
Outcome: The proposed pipeline generates a high-quality multi-turn chat corpus by leveraging ChatGPT to engage in a conversation with itself, simulating both user and AI responses.
RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs (2023.acl-long)

Copied to clipboard

Challenge: Despite their success, even the largest language models make mistakes.
Approach: They propose a framework where one language model can generate critiques to improve its peer's performance.
Outcome: The proposed framework improves the performance of a fixed model 200 times its size by 10% over other models.
Feedback to Reasoning: LLM-Assisted Molecular Optimization with Domain Feedback and Historical Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for molecular optimization do not leverage domain feedback and historical knowledge with reasoning traces and chemical insights.
Approach: They propose a conversational molecular optimization pipeline that enables LLMs to accumulate and retrieve past actions, rationales, and feedback.
Outcome: The proposed framework transforms LLMs from passive text generators into agentic experts that learn both actions and reasoning from experience.
Learning from a Friend: Improving Event Extraction via Self-Training with Feedback from Abstract Meaning Representation (2023.findings-acl)

Copied to clipboard

Challenge: Existing data scarcity hinders the progress of event extraction, authors say . ACE-052 has 10 of the 33 event types with less than 80 annotations, authors claim .
Approach: They propose a self-training with feedback framework that leverages large-scale unlabeled data to acquire feedback for each new event prediction from the unlabed data.
Outcome: The proposed framework improves event extraction models even when unlabeled data are unavailable.
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement (2024.findings-acl)

Copied to clipboard

Challenge: OpenCodeInterpreter-33B provides a high level of performance for code generation, executing, and iterative refinement.
Approach: They propose a family of open-source code systems for generating, executing, and iteratively refining code.
Outcome: The OpenCodeInterpreter-33B performs well on humanEval, MBPP, and EvalPlus benchmarks.
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks (2025.acl-long)

Copied to clipboard

Challenge: Inference-Time Scaling is critical to the success of recent models such as OpenAI o1 and DeepSeek R1 . however, many techniques require tasks to have answers that can be verified .
Approach: They use data to train dedicated Feedback and Edit Models capable of inference-time scaling for open-ended tasks.
Outcome: The proposed model can reach SoTA performance on Arena Hard at 92.7 as of 5 Mar 2025.
Iterative Repair with Weak Verifiers for Few-shot Transfer in KBQA with Unanswerability (2025.findings-acl)

Copied to clipboard

Challenge: Existing models for KBQA with unanswerable questions are inadequate for real-world applications.
Approach: They propose a task of few-shot transfer for KBQA with unanswerable questions that extends FuSIC-KBQA to include feedback for unanswered questions.
Outcome: The proposed model outperforms suitable adaptations of multiple LLM-based and supervised SoTA models on the task while establishing a new performance for answerable few-shot transfer as well.
Creating Grammar Teaching Material for Endangered Languages with Hybrid Grammar Induction (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for creating grammar lessons are labor-intensive and often fall to teachers who lack formal training in grammar.
Approach: They propose a hybrid grammar-induction method that uses typological priors, Bayesian inference, constrained LLM reasoning and retrieval from sparse corpora to generate topic-specific grammar lessons.
Outcome: The proposed method can produce coherent and useful lessons with better quality when modest explanatory evidence is available.
Feedback Is The Key for Automated Survey Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a promising foundation for literature surveys, but guiding them to generate accurate, reliable content remains a fundamental challenge.
Approach: They propose a feedback-driven framework that incorporates feedback across three dimensions: outline feedback for structural clarity, citation feedback for evidence validation, and content feedback for readability and analytical depth.
Outcome: The proposed framework significantly improves both citation and content quality, demonstrating feedback as the critical mechanism for automatic survey generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations