Papers with YouTube
Multifaceted Domain-Specific Document Embeddings (2021.naacl-demos)
Copied to clipboard
| Challenge: | Current document embeddings require large training corpora but fail to learn high-quality representations when confronted with a small number of domain-specific documents and rare terms. |
| Approach: | They propose a faceted domain encoder that transforms each document into a single embedding vector . they use a Siamese neural network architecture to leverage knowledge graphs to enhance the embeddables . |
| Outcome: | The proposed model achieves the same embedding quality as state-of-the-art models while requiring only a tiny fraction of training data. |
ORANGE: Text-video Retrieval via Watch-time-aware Heterogeneous Graph Contrastive Learning (2023.emnlp-industry)
Copied to clipboard
Yucheng Lin, Tim Chang, Yaning Chang, Jianqiang Ma, Donghui Li, Ting Peng, Zang Li, Zhiyi Zhou, Feng Wang
| Challenge: | Existing methods for text-video retrieval focus on informative representations and delicate matching mechanisms, but real-world scenarios often involve brief, ambiguous queries and low-quality videos. |
| Approach: | They propose a novel method to learn informative embeddings for queries and videos . they use a watch-time-aware contrastive learning paradigm to capture dependencies . |
| Outcome: | The proposed method is effective in a real-world video-search service. |
entity-linkings: A Unified Library for Entity Linking (2026.eacl-demo)
Copied to clipboard
| Challenge: | Entity linking (EL) is the task of mapping named entities in text to canonical entries in a knowledge base. |
| Approach: | They propose a unified library for using and developing entity linking systems . a strong emphasis is placed on usability, making it highly extensible . |
| Outcome: | a new library aims to disambiguate named entities in text by mapping them to canonical entries in a knowledge base. |
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability (2025.emnlp-demos)
Copied to clipboard
Mohammad Aflah Khan, Ameya Godbole, Johnny Wei, Ryan Yixiang Wang, James Flemings, Krishna P. Gummadi, Willie Neiswanger, Robin Jia
| Challenge: | Existing workflows for pretraining large language models are cumbersome, fragmented and inaccessible. |
| Approach: | They propose an open-source library for editing, inspection, and analysis of large language model datasets. |
| Outcome: | TokenSmith is an open-source library for editing, inspection, and analysis of large language model datasets. |
Multi-source Multi-domain Sentiment Analysis with BERT-based Models (2022.lrec-1)
Copied to clipboard
| Challenge: | Sentiment analysis is a widely studied task in natural language processing. |
| Approach: | They propose to improve BERT-based models for sentiment analysis on italian corpora and evaluate their performance on the basis of eight corpors. |
| Outcome: | The proposed model is evaluated over eight sentiment analysis corpora from different domains and sources on the prediction of positive, negative and neutral classes. |
VlogQA: Task, Dataset, and Baseline Models for Vietnamese Spoken-Based Machine Reading Comprehension (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing datasets for machine reading comprehension tasks in Vietnamese focus on written documents, such as Wikipedia articles, online newspapers, or textbooks. |
| Approach: | They propose to capture Vietnamese spoken language in natural settings and use it to create a machine-learning corpus for machine reading comprehension tasks. |
| Outcome: | The proposed corpus consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube . |
AssameseBackTranslit: Back Transliteration of Romanized Assamese Social Media Text (2024.lrec-main)
Copied to clipboard
| Challenge: | a novel dataset capturing native text composed in the Roman/Latin script is presented . the dataset comprises 60,312 Roman-native parallel transliterated sentences . |
| Approach: | They propose a back transliteration dataset capturing native text composed in the Roman/Latin script and its corresponding representation in the native Assamese script. |
| Outcome: | The proposed dataset outperforms baseline models in terms of word-level transliteration evaluation benchmarks and performance assessments. |
MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform (2025.emnlp-main)
Copied to clipboard
| Challenge: | 108K drug overdose deaths in 2022, according to NIDA . |
| Approach: | They propose a large-scale study of OUD-related myths on YouTube with clinical experts to validate 8 pervasive myths and release an expert-labeled video dataset. |
| Outcome: | The proposed model reduces annotation time and cost by over 76% compared to experts and full LLM labeling. |
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild (2025.findings-naacl)
Copied to clipboard
Junhyeok Kim, Min Soo Kim, Jiwan Chung, Jungbin Cho, Jisoo Kim, Sungwoong Kim, Gyeongbo Sim, Youngjae Yu
| Challenge: | EgoSpeak predicts when an agent should begin speaking based on egocentric streaming video. |
| Approach: | They propose a framework for real-time speech initiation prediction in egocentric streaming video by modeling the conversation from the camera wearer's first-person perspective. |
| Outcome: | The proposed framework outperforms random and silence-based baselines in real time and highlights the importance of multimodal input and context length in effectively deciding when to speak. |
Can Language Models Laugh at YouTube Short-form Videos? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets that focus on verbal cues and focus on short-form funny videos focus on focusing on verbs and visual cue. |
| Approach: | They curate a user-generated dataset of 10K multimodal funny videos from YouTube and annotate each video with timestamps and explanations for funny moments. |
| Outcome: | The proposed dataset improves the ability of large language models to understand humor. |
A Dataset of Offensive Language in Kosovo Social Media (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media are a central part of people’s lives but are rife with bullying and offensive language, creating an unsafe environment for their users. |
| Approach: | They propose to use user-generated comments on Facebook and YouTube from selected Kosovo news platforms to annotate offensive language in Albanian. |
| Outcome: | The proposed system improves on Danish but not Albanian, on offensive language recognition and distinguishing targeted and untargeted offence. |
Dataset for Identification of Homophobia and Transphobia for Telugu, Kannada, and Gujarati (2024.lrec-main)
Copied to clipboard
| Challenge: | There has been a rise in homophobic and transphobic content targeting LGBT+ individuals on social media platforms. |
| Approach: | They propose to use a dataset to automatically identify homophobic and transphobic content within comments collected from YouTube for three languages. |
| Outcome: | The proposed dataset will identify homophobic and transphobic content within comments collected from YouTube in Telugu, Kannada, and Gujarati. |
Automatic In-the-wild Dataset Annotation with Deep Generalized Multiple Instance Learning (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to label large datasets that resemble real life situations are prohibitive due to the cost of manual labeling. |
| Approach: | They propose to automate the annotation process by using end-to-end differentiable neural networks to label large datasets that resemble real life conditions. |
| Outcome: | The proposed method can label a large dataset in the wild without human intervention without any cost. |
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)
Copied to clipboard
Ritesh Kumar, Shyam Ratan, Siddharth Singh, Enakshi Nandi, Laishram Niranjana Devi, Akash Bhagat, Yogesh Dawer, Bornini Lahiri, Akanksha Bansal, Atul Kr. Ojha
| Challenge: | 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms. |
| Approach: | They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. |
| Outcome: | The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English. |
GREENER: Graph Neural Networks for News Media Profiling (2022.emnlp-main)
Copied to clipboard
| Challenge: | a new method for profiling news media on the Web addresses the factuality of reporting and bias problem . a recent study has focused on text features but has focused primarily on text . |
| Approach: | They propose a model that models the similarity between media outlets based on their audience overlap . they propose GREENER, which builds a graph of inter-media connections based upon audience overlap. |
| Outcome: | The proposed model improves on state-of-the-art models on two datasets. |
YouMakeup: A Large-Scale Domain-Specific Multimodal Dataset for Fine-Grained Semantic Comprehension (D19-1)
Copied to clipboard
| Challenge: | Multimodal semantic comprehension has attracted increasing research interest recently such as visual question answering and caption generation. |
| Approach: | They propose to use a large-scale multimodal instructional video dataset to support fine-grained comprehension research in specific domain. |
| Outcome: | The proposed dataset contains 2,800 videos from YouTube, spanning more than 420 hours in total. |
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos (2024.findings-acl)
Copied to clipboard
| Challenge: | Using a dataset of 931 videos with 4021 code-mixed Hindi-English utterances, we find that video content with multiple modalities is more accurate and more accurate than textual content. |
| Approach: | They propose to use a dataset to analyze toxic content in video content in non-English languages by leveraging language models. |
| Outcome: | The proposed framework achieves an Accuracy and Weighted F1 score of 94.29% and 94.35% for the first time in its class. |
Paying Attention to Deflections: Mining Pragmatic Nuances for Whataboutism Detection in Online Discourse (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on whataboutism have focused on tracking "what about" phrases, but they neglect the unique challenges to its detection. |
| Approach: | They propose to use attention weights to distinguish the ‘what about’ lexical construct from whataboutism by using Twitter/X and YouTube datasets. |
| Outcome: | The proposed method improves by 4% and 10% over previous state-of-the-art methods in Twitter and YouTube datasets. |
A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection (2020.lrec-1)
Copied to clipboard
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-gyo Jung, Bernard J. Jansen, Joni Salminen
| Challenge: | Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals. |
| Approach: | They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments . |
| Outcome: | The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube. |
Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing toxic speech datasets lack demographic context and age data are limited . funk and its subsidiary accounts target users aged 14-29 . |
| Approach: | a german project introduces a large-scale toxic speech dataset annotated for toxicity . the dataset includes 3,024 human-annotated and 30,024 LLM-annnotated comments . researchers used human expertise and state-of-the-art language models to label comments based on toxic keywords . |
| Outcome: | The study combines human expertise with state-of-the-art language models to identify toxic speech categories. |
Personalized Video Comment Generation (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Generating personalized responses in video poses a unique challenge for language models. |
| Approach: | They propose a new automatic metric based on Large Language Models with few-shot in-context learning that measures quality from the aspects of emotion, language style and content relevance. |
| Outcome: | The proposed metric measures quality from emotion, language style and content relevance with human evaluations. |