Challenge: Existing methods to “vet” labels from noisy captions for weakly-supervised object detection are limited for object detection.
Approach: They propose a technique to “vet” labels extracted from noisy captions and use them for weakly-supervised object detection without any bounding boxes.
Outcome: The proposed method improves WSOD without label vetting by 30% on five datasets.

Similar Papers

Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning Approach (D19-1)

Copied to clipboard

Challenge: Recent work on image captioning has made impressive progress . however, the results are limited and the model is difficult to train .
Approach: They propose a semi-supervised framework for training an image captioning model by assigning pseudo-labels to unpaired samples via Generative Adversarial Networks.
Outcome: The proposed framework is compared to baselines when the number of paired samples is scarce.
Learning from Noisy Labels for Entity-Centric Information Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Recent information extraction approaches can easily overfit noisy labels and suffer from performance degradation.
Approach: They propose a co-regularization framework for entity-centric information extraction that optimizes neural models with task-specific losses and regularizes them to generate similar predictions based on agreement loss.
Outcome: The proposed framework is optimized with task-specific losses and generates similar predictions based on agreement loss.
Learning Named Entity Tagger using Domain-Specific Dictionary (D18-1)

Copied to clipboard

Challenge: Existing methods to build reliable named entity recognition systems require large amounts of manually-labeled training data.
Approach: They propose a revised fuzzy CRF layer to handle tokens with multiple possible labels to address noisy distant supervision.
Outcome: The proposed model can handle tokens with multiple possible labels under the traditional framework and improves on the existing model with a new Tie or Break scheme.
Weakly-Supervised Learning of Visual Relations in Multimodal Pretraining (2023.emnlp-main)

Copied to clipboard

Challenge: Recent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations.
Approach: They propose two approaches to contextualise visual entities in a multimodal setup by using verbalised scene graphs and masked relation prediction.
Outcome: The proposed models can learn better representations from weakly-supervised relations data.
Automatically Identifying Words That Can Serve as Labels for Few-Shot Text Classification (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to few-shot text classification require domain expertise and an understanding of the language model's abilities to define the mapping between words and labels.
Approach: They propose a method that converts textual inputs to cloze questions that contain some form of task description and processes them with a pretrained language model to map the predicted words to labels.
Outcome: The proposed approach performs almost as well as hand-crafted label-to-word mappings for a number of tasks with small amounts of training data.
Learning to Relate from Captions and Bounding Boxes (P19-1)

Copied to clipboard

Challenge: Existing methods for classifying images without supervision are limited.
Approach: They propose a top-down attention mechanism to align entities in captions to objects in the image and leverage the syntactic structure of captions for alignment.
Outcome: The proposed model achieves a recall@50 of 15% and recall@100 of 25% on the relationships present in the image and predicts relations that are not present in captions.
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets for abusive language detection are expensive and lack of knowledge about the target is a challenge.
Approach: They propose to build models cheaply for a new target label set and/or language, using only a few training examples of the target domain.
Outcome: The proposed model improves monolingually and across languages using existing datasets and only a few-shots of the target domain.
Denoising Multi-Source Weak Supervision for Neural Text Classification (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent years have witnessed the rapid development of deep neural networks (DNNs) for text classification problems.
Approach: They propose a label denoiser which estimates the source reliability using a conditional soft attention mechanism and reduces label noise by aggregating rule-annotated weak labels.
Outcome: The proposed model outperforms state-of-the-art methods on sentiment, topic, and relation classifications and achieves comparable performance with fully-supervised methods even without labeled data.
Defoiling Foiled Image Captions (N18-2)

Copied to clipboard

Challenge: Existing models for vision-to-language tasks do not understand images sufficiently, despite their impressive performance.
Approach: They propose to use explicit object information to solve foiled caption detection problem . they propose to replace a word in caption with a semantically similar word .
Outcome: The proposed model achieves state-of-the-art on a recently published dataset with scores exceeding those achieved by humans on the task.
Enhancing Multi-Label Text Classification under Label-Dependent Noise: A Label-Specific Denoising Framework (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing noisy multi-label text classification methods rely on the class-conditional noise assumption, but in practice, noisy labels exhibit a certain degree of correlation with the true labels.
Approach: They propose a label-specific denoising framework to counteract label-dependent noise by evaluating loss information, ranking information, and feature centroid.
Outcome: The proposed framework significantly improves over existing state-of-the-art models under both synthetic and real-world noise conditions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations