Papers with YouTube

21 papers
Multifaceted Domain-Specific Document Embeddings (2021.naacl-demos)

Copied to clipboard

Challenge: Current document embeddings require large training corpora but fail to learn high-quality representations when confronted with a small number of domain-specific documents and rare terms.
Approach: They propose a faceted domain encoder that transforms each document into a single embedding vector . they use a Siamese neural network architecture to leverage knowledge graphs to enhance the embeddables .
Outcome: The proposed model achieves the same embedding quality as state-of-the-art models while requiring only a tiny fraction of training data.
ORANGE: Text-video Retrieval via Watch-time-aware Heterogeneous Graph Contrastive Learning (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for text-video retrieval focus on informative representations and delicate matching mechanisms, but real-world scenarios often involve brief, ambiguous queries and low-quality videos.
Approach: They propose a novel method to learn informative embeddings for queries and videos . they use a watch-time-aware contrastive learning paradigm to capture dependencies .
Outcome: The proposed method is effective in a real-world video-search service.
entity-linkings: A Unified Library for Entity Linking (2026.eacl-demo)

Copied to clipboard

Challenge: Entity linking (EL) is the task of mapping named entities in text to canonical entries in a knowledge base.
Approach: They propose a unified library for using and developing entity linking systems . a strong emphasis is placed on usability, making it highly extensible .
Outcome: a new library aims to disambiguate named entities in text by mapping them to canonical entries in a knowledge base.
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing workflows for pretraining large language models are cumbersome, fragmented and inaccessible.
Approach: They propose an open-source library for editing, inspection, and analysis of large language model datasets.
Outcome: TokenSmith is an open-source library for editing, inspection, and analysis of large language model datasets.
Multi-source Multi-domain Sentiment Analysis with BERT-based Models (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis is a widely studied task in natural language processing.
Approach: They propose to improve BERT-based models for sentiment analysis on italian corpora and evaluate their performance on the basis of eight corpors.
Outcome: The proposed model is evaluated over eight sentiment analysis corpora from different domains and sources on the prediction of positive, negative and neutral classes.
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension (2024.eacl-long)

Copied to clipboard

Challenge: Existing datasets for machine reading comprehension tasks in Vietnamese focus on written documents, such as Wikipedia articles, online newspapers, or textbooks.
Approach: They propose to capture Vietnamese spoken language in natural settings and use it to create a machine-learning corpus for machine reading comprehension tasks.
Outcome: The proposed corpus consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube .
AssameseBackTranslit: Back Transliteration of Romanized Assamese Social Media Text (2024.lrec-main)

Copied to clipboard

Challenge: a novel dataset capturing native text composed in the Roman/Latin script is presented . the dataset comprises 60,312 Roman-native parallel transliterated sentences .
Approach: They propose a back transliteration dataset capturing native text composed in the Roman/Latin script and its corresponding representation in the native Assamese script.
Outcome: The proposed dataset outperforms baseline models in terms of word-level transliteration evaluation benchmarks and performance assessments.
MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform (2025.emnlp-main)

Copied to clipboard

Challenge: 108K drug overdose deaths in 2022, according to NIDA .
Approach: They propose a large-scale study of OUD-related myths on YouTube with clinical experts to validate 8 pervasive myths and release an expert-labeled video dataset.
Outcome: The proposed model reduces annotation time and cost by over 76% compared to experts and full LLM labeling.
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild (2025.findings-naacl)

Copied to clipboard

Challenge: EgoSpeak predicts when an agent should begin speaking based on egocentric streaming video.
Approach: They propose a framework for real-time speech initiation prediction in egocentric streaming video by modeling the conversation from the camera wearer's first-person perspective.
Outcome: The proposed framework outperforms random and silence-based baselines in real time and highlights the importance of multimodal input and context length in effectively deciding when to speak.
Can Language Models Laugh at YouTube Short-form Videos? (2023.emnlp-main)

Copied to clipboard

Challenge: Existing datasets that focus on verbal cues and focus on short-form funny videos focus on focusing on verbs and visual cue.
Approach: They curate a user-generated dataset of 10K multimodal funny videos from YouTube and annotate each video with timestamps and explanations for funny moments.
Outcome: The proposed dataset improves the ability of large language models to understand humor.
A Dataset of Offensive Language in Kosovo Social Media (2022.lrec-1)

Copied to clipboard

Challenge: Social media are a central part of people’s lives but are rife with bullying and offensive language, creating an unsafe environment for their users.
Approach: They propose to use user-generated comments on Facebook and YouTube from selected Kosovo news platforms to annotate offensive language in Albanian.
Outcome: The proposed system improves on Danish but not Albanian, on offensive language recognition and distinguishing targeted and untargeted offence.
Dataset for Identification of Homophobia and Transphobia for Telugu, Kannada, and Gujarati (2024.lrec-main)

Copied to clipboard

Challenge: There has been a rise in homophobic and transphobic content targeting LGBT+ individuals on social media platforms.
Approach: They propose to use a dataset to automatically identify homophobic and transphobic content within comments collected from YouTube for three languages.
Outcome: The proposed dataset will identify homophobic and transphobic content within comments collected from YouTube in Telugu, Kannada, and Gujarati.
Automatic In-the-wild Dataset Annotation with Deep Generalized Multiple Instance Learning (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to label large datasets that resemble real life situations are prohibitive due to the cost of manual labeling.
Approach: They propose to automate the annotation process by using end-to-end differentiable neural networks to label large datasets that resemble real life conditions.
Outcome: The proposed method can label a large dataset in the wild without human intervention without any cost.
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)

Copied to clipboard

Challenge: 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms.
Approach: They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
Outcome: The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English.
GREENER: Graph Neural Networks for News Media Profiling (2022.emnlp-main)

Copied to clipboard

Challenge: a new method for profiling news media on the Web addresses the factuality of reporting and bias problem . a recent study has focused on text features but has focused primarily on text .
Approach: They propose a model that models the similarity between media outlets based on their audience overlap . they propose GREENER, which builds a graph of inter-media connections based upon audience overlap.
Outcome: The proposed model improves on state-of-the-art models on two datasets.
YouMakeup: A Large-Scale Domain-Specific Multimodal Dataset for Fine-Grained Semantic Comprehension (D19-1)

Copied to clipboard

Challenge: Multimodal semantic comprehension has attracted increasing research interest recently such as visual question answering and caption generation.
Approach: They propose to use a large-scale multimodal instructional video dataset to support fine-grained comprehension research in specific domain.
Outcome: The proposed dataset contains 2,800 videos from YouTube, spanning more than 420 hours in total.
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos (2024.findings-acl)

Copied to clipboard

Challenge: Using a dataset of 931 videos with 4021 code-mixed Hindi-English utterances, we find that video content with multiple modalities is more accurate and more accurate than textual content.
Approach: They propose to use a dataset to analyze toxic content in video content in non-English languages by leveraging language models.
Outcome: The proposed framework achieves an Accuracy and Weighted F1 score of 94.29% and 94.35% for the first time in its class.
Paying Attention to Deflections: Mining Pragmatic Nuances for Whataboutism Detection in Online Discourse (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on whataboutism have focused on tracking "what about" phrases, but they neglect the unique challenges to its detection.
Approach: They propose to use attention weights to distinguish the ‘what about’ lexical construct from whataboutism by using Twitter/X and YouTube datasets.
Outcome: The proposed method improves by 4% and 10% over previous state-of-the-art methods in Twitter and YouTube datasets.
A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection (2020.lrec-1)

Copied to clipboard

Challenge: Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals.
Approach: They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments .
Outcome: The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube.
Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting (2025.emnlp-main)

Copied to clipboard

Challenge: Existing toxic speech datasets lack demographic context and age data are limited . funk and its subsidiary accounts target users aged 14-29 .
Approach: a german project introduces a large-scale toxic speech dataset annotated for toxicity . the dataset includes 3,024 human-annotated and 30,024 LLM-annnotated comments . researchers used human expertise and state-of-the-art language models to label comments based on toxic keywords .
Outcome: The study combines human expertise with state-of-the-art language models to identify toxic speech categories.
Personalized Video Comment Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Generating personalized responses in video poses a unique challenge for language models.
Approach: They propose a new automatic metric based on Large Language Models with few-shot in-context learning that measures quality from the aspects of emotion, language style and content relevance.
Outcome: The proposed metric measures quality from emotion, language style and content relevance with human evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations