Papers with human-in-the-loop

9 papers
Quality-Aware Adversarial Ensemble for Singer Identification in 1960s Tamil Film Music (2026.eacl-srw)

Copied to clipboard

Challenge: 1960s Tamil film music lacks adequate metadata identifying playback singers in archival recordings.
Approach: They propose a quality-aware adversarial ensemble approach based on variable audio degradation and instrumentation leakage confounding singer-specific features.
Outcome: The proposed approach achieves 96.2% accuracy and 2.0% EER on a held-out test set of 52 clips.
Active Learning for New Domains in Natural Language Understanding (N19-2)

Copied to clipboard

Challenge: Existing approaches to improve the accuracy of new domains are lacking annotated live utterances.
Approach: They propose an algorithm called Majority-CRF that uses an ensemble of classification models to guide the selection of relevant utterances and a sequence labeling model to prioritize informative examples.
Outcome: The proposed algorithm achieves 6.6%-9% error rate reduction and statistically significant improvements on six new domains.
First-AID: the first Annotation Interface for grounded Dialogues (2025.acl-demo)

Copied to clipboard

Challenge: Existing tools to fine-tune Large Language Models for specific tasks are limited due to financial constraints and limited availability of human experts.
Approach: They propose a human-in-the-loop framework for the knowledge-driven generation of synthetic dialogues using LLM prompting that implements different strategies of data collection that require different user intervention during dialogue generation.
Outcome: The proposed framework reduces post-editing efforts and improves quality of generated dialogues.
AnnoHID: LLM-Assisted Annotation Framework for Low-Resource Medical Texts (2026.acl-demo)

Copied to clipboard

Challenge: Social media platforms are a popular way to communicate with medical experts and improve health literacy.
Approach: They introduce a semi-automated annotation framework for medical texts in low-resource languages . they use large language models for pre-annotation and human validation to support efficient annotation .
Outcome: The proposed framework is applied to medical social media texts in Bahasa Indonesia . it yields higher inter-annotator agreement and human review improves output . future work focuses on mitigating pre-annotation bias and reducing annotation overhead .
Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness (2024.acl-long)

Copied to clipboard

Challenge: BSDetector is a method for detecting bad and speculative answers from a pretrained Large Language Model.
Approach: They propose a method for detecting bad and speculative answers from a pretrained Large Language Model by estimating a confidence score for any output it generated.
Outcome: Experiments on open-form Question-Answer benchmarks show that BSDetector more accurately identifies incorrect LLM responses than alternative uncertainty estimation procedures.
Enabling Interactive Transcription in an Indigenous Community (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for manual transcription are often in isolation from the speech community, and so we miss out on the opportunity to take advantage of the interests and skills of local people.
Approach: They propose a transcription workflow which combines spoken term detection and human-in-the-loop to support speech transcription in almost-zero resource settings.
Outcome: The proposed workflow is based on two endangered languages with zero-resource datasets.
CORWA: A Citation-Oriented Related Work Annotation Dataset (2022.naacl-main)

Copied to clipboard

Challenge: Academic research is an exploratory activity to discover new solutions to problems . prior work focused on the sentence as the basic unit of generation, neglecting that related work sections consist of variable length text fragments derived from different information sources.
Approach: They propose a Citation Oriented Related Work Annotation dataset that labels citation text fragments . they propose linguistically-motivated framework for human-in-the-loop, abstractive related work generation .
Outcome: The proposed framework is based on a Citation Oriented Related Work Annotation dataset . it automatically tags unlabeled related work sections on the dataset based upon the proposed model .
Exploring Data Augmentation Strategies for Hate Speech Detection in Roman Urdu (2022.lrec-1)

Copied to clipboard

Challenge: a number of social media platforms are generating hateful content, a new study finds . augmentation techniques are needed to improve the performance of the models .
Approach: They evaluate different data augmentation techniques for the improvement of hate speech detection in Roman Urdu.
Outcome: The proposed techniques improve hate speech detection in Roman Urdu on two datasets.
STORM-BORN: A Challenging Mathematical Derivations Dataset Curated via a Human-in-the-Loop Multi-Agent Framework (2025.findings-acl)

Copied to clipboard

Challenge: Existing datasets suffer from outdated and insufficient challenging content, neglecting human-like reasoning, and limited reliability due to single-LLM generation.
Approach: They propose a human-in-the-loop, multi-agent data generation framework that integrates reasoning-dense filters, multiagent collaboration, and human mathematicians’ evaluations to ensure the reliability and quality of the dataset.
Outcome: The proposed framework improves accuracy and quality of the 2,000-synthesized datasets by integrating reasoning-dense filters, multi-agent collaboration, and human mathematicians’ evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations