| Challenge: | Existing methods to learn correspondence between visual segments and texts require temporal coordinates for training, which leads to high costs of annotation. |
| Approach: | They propose weakly supervised language localization networks to detect events in untrimmed videos . they train with only video-sentence pairs without accessing to temporal locations of events . |
| Outcome: | Experiments on ActivityNet Captions and DiDeMo show that WSLLN performs state-of-the-art. |
Similar Papers
Scene-robust Natural Language Video Localization via Learning Domain-invariant Representations (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have focused on improving performance with the assumption of independently identical data distribution while ignoring out-of-distribution data. |
| Approach: | They propose a scene-robust NLVL problem and a generalizable framework to learn a robust model. |
| Outcome: | The proposed model learns generalizable domain-invariant representations by alignment and decomposition. |
Natural Language Video Localization with Learnable Moment Proposals (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for video moment localization have poor performance due to predefined rules. |
| Approach: | They propose a model with a fixed set of learnable moment proposals with 'border-aware loss' they propose to localize the video moment corresponding to the query by locating the start and end timestamps in an untrimmed video. |
| Outcome: | The proposed model outperforms state-of-the-art models on two challenging benchmarks. |
Exploiting Intrinsic Multilateral Logical Rules for Weakly Supervised Natural Language Video Localization (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for WS-NLVL rarely consider complex temporal relations enclosing the language query, yielding illogical predictions. |
| Approach: | They propose a plug-and-play method to exploit temporal relations and logical rules for WS-NLVL. |
| Outcome: | The proposed method is able to retrieve the moment corresponding to a language query in a video with only video-language pairs utilized during training. |
Weakly Supervised Word Segmentation for Computational Language Documentation (2022.acl-long)
Copied to clipboard
| Challenge: | a recent paper aims to improve the effectiveness of unsupervised language analysis techniques in low resource settings. |
| Approach: | They propose to use a weak supervision to improve linguistic segmentation in low resource languages . they propose to provide linguists with LTs that can be used to create interactive annotation tools . |
| Outcome: | The proposed models can be used to improve the quality of language segmentation in low resource languages. |
Weakly-Supervised Spoken Video Grounding via Semantic Interaction Learning (2023.acl-long)
Copied to clipboard
| Challenge: | Recent work on spoken video grounding challenges extracting semantic information from speech . previous studies focused on textual queries, but recent work focuses on spoken queries . |
| Approach: | They propose a framework for weakly-supervised spoken video grounding to represent cross-modal semantics without expensive temporal annotations. |
| Outcome: | The proposed framework is more efficient than existing methods. |
Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence Localization (2023.acl-long)
Copied to clipboard
| Challenge: | Existing zero-shot pipelines generate event proposals and then generate a pseudo query for each event proposal. |
| Approach: | They propose a Structure-based Pseudo Label generation (SPL) that generates free-form interpretable pseudo queries before constructing query-dependent event proposals. |
| Outcome: | The proposed method learns with only video data without any annotation . it generates free-form interpretable pseudo queries before constructing query-dependent event proposals . |
Unsupervised Cross-Lingual Representation Learning (P19-4)
Copied to clipboard
| Challenge: | a comprehensive survey of cutting-edge weakly-supervised and unsupervised cross-lingual word representations is presented . |
| Approach: | This tutorial provides a comprehensive survey of recent work on weakly-supervised and unsupervised cross-lingual word representations. |
| Outcome: | This tutorial provides a comprehensive survey of cutting-edge weakly-supervised and unsupervised word representations. |
Low-resource Cross-lingual Event Type Detection via Distant Supervision with Minimal Effort (C18-1)
Copied to clipboard
| Challenge: | Currently, few or no language processing tools or resources exist for most languages . a problem is that there is not enough available training data even in resource-rich languages if the task is complex. |
| Approach: | They propose to use a bilingual dictionary to train machine learning in a resource-poor language . they also explore adversarial training of bilingual word representations . |
| Outcome: | The proposed approach gives similar performance in event-type detection tasks. |
Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video (P19-1)
Copied to clipboard
| Challenge: | Existing techniques for weakly-supervised spatio-temporally grounding natural sentence in video are lacking . |
| Approach: | They propose a weakly-supervised task for spatially grounding sentences in video . they extract instances from video and encode them using attentive interactor . results demonstrate superiority of their proposed task over baseline approaches . |
| Outcome: | The proposed model outperforms baseline approaches in a weakly-supervised task . it can characterize reliable instance-sentence pairs and penalize unreliable ones . |
Weakly Supervised Vision-and-Language Pre-training with Relative Representations (2023.acl-long)
Copied to clipboard
| Challenge: | Weakly supervised vision-and-language pre-training (WVLP) uses only local descriptions of images as cross-modal anchors to construct weakly-aligned image-text pairs for pre- training. |
| Approach: | They propose to take a small number of aligned image-text pairs as anchors and represent each unaligned image and text by its similarities to these anchors. |
| Outcome: | The proposed model reduces the cost of pre-training while maintaining decent performance on downstream tasks. |