Challenge: Existing approaches for optimizing human annotation efforts are limited . et al., 2015) suggest that densely annotated image captions improve vision-language alignment .
Approach: They propose an AI-in-the-loop methodology to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints.
Outcome: The proposed method improves annotation speed and retrieval performance over the parallel method.

Similar Papers

HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have shown that high-quality video captions can improve MLLMs' performance on videos involving human actions.
Approach: They propose a data annotation pipeline to collect videos featuring clear human actions from the Internet and annotate them in a standardized caption format that uses human attributes to distinguish individuals.
Outcome: The proposed pipeline combines two datasets to evaluate human action understanding.
CaBSALLM: Efficient Context-Aware Batch Annotation of Conversational Streams with Large Language Models (2026.acl-short)

Copied to clipboard

Challenge: Large-scale annotations of subjective, discourse-dependent social interactions remain a critical bottleneck in computational social science.
Approach: They propose a pipeline that incorporates lightweight conversational context and a dynamic batching method to improve throughput and scalability.
Outcome: The proposed pipeline improves throughput and scalability while preserving interpretive depth essential to complex social annotations.
Discourse-Centric Evaluation of Document-level Machine Translation with a New Densely Annotated Parallel Corpus of Novels (2023.acl-long)

Copied to clipboard

Challenge: Several recent papers claim to have achieved human parity at sentence-level machine translation.
Approach: They propose to use a dataset with rich discourse annotations to evaluate MT performance . they find that MT outputs differ fundamentally from human translations in terms of latent discourse structures.
Outcome: The proposed dataset builds upon the large-scale parallel corpus BWB . it covers 15,095 entity mentions in both languages and compares them to human translations .
IC3: Image Captioning by Committee Consensus (2023.emnlp-main)

Copied to clipboard

Challenge: Traditionally, image captioning models are trained to generate a single “best’ (most like a reference) image caption.
Approach: They propose a method to generate a single caption that captures high-level details from several annotator viewpoints.
Outcome: The proposed method outperforms baseline SOTA models and improves the performance of automated recall systems by up to 84%.
GPTs Are Multilingual Annotators for Sequence Generation Tasks (2024.findings-eacl)

Copied to clipboard

Challenge: Existing methods of data annotation are time-consuming and expensive . complexity of crowdsourcing increases when dealing with low-resource languages .
Approach: They propose an autonomous method to gather unlabeled data and label them using large language models.
Outcome: The proposed method is cost-efficient and applicable for low-resource language annotation.
Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies (2026.findings-acl)

Copied to clipboard

Challenge: Simultaneous machine translation requires high-quality translations under strict real-time constraints.
Approach: They extend the action space of simultaneous machine translation with four adaptive actions . they adapt these actions in a large language model framework and construct training references .
Outcome: The proposed framework improves semantic metrics and achieves lower delay compared to reference translations and salami-based baselines.
Multimodal Pretraining for Dense Video Captioning (2020.aacl-main)

Copied to clipboard

Challenge: a billion hours of videos are being watched on YouTube every day . videos are difficult to skim through, making it harder to quickly target the relevant part(s) of a video.
Approach: They propose to use a video timeline tag dataset to generate time-stamped annotations for videos . they propose to pretrain and finetune captioning models using YouCook2 and ViTT .
Outcome: The proposed model generalizes well and is robust over a wide variety of instructional videos.
Mitigating Open-Vocabulary Caption Hallucinations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for image captioning ignore the long-tailed nature of hallucinations . a new framework is proposed to address hallucines in image captions in the open-vocabulary setting .
Approach: They propose a framework to address hallucinations in image captioning in the open-vocabulary setting.
Outcome: The proposed framework surpasses the CHAIR benchmark in diversity and accuracy in open-vocabulary captioning.
MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for supervised visual captioning require large scale of images or videos paired with descriptions in a specific language.
Approach: They propose a zero-shot approach that generates captions for different scenarios without labeling . they use concept prompts to retrieve concepts and auto-encode them to learn writing styles .
Outcome: The proposed approach generates captions for different scenarios and languages without labeled vision-caption pairs.
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains (2025.acl-long)

Copied to clipboard

Challenge: Existing vision-language models struggle to disentangle information scattered across complex visual inputs, leading to performance degradation.
Approach: They propose a focus-centric visual chain paradigm that enhances VLMs’ perception, comprehension, and reasoning abilities in multi-image scenarios.
Outcome: The proposed approach achieves average performance gains of 3.16% and 2.24% across two distinct model architectures, without compromising the general vision-language capabilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations