BehanceCC: A ChitChat Detection Dataset For Livestreaming Video Transcripts (2022.lrec-1)
Copied to clipboard
| Challenge: | livestreaming videos contain a considerable amount of off-topic content, causing noises and data load to downstream applications. |
| Approach: | They propose a human-annotated benchmark dataset for off-topic detection in livestreaming video transcripts. |
| Outcome: | The proposed dataset reveals the complexity of chitchat detection in livestreaming videos . livestreams tend to be longer than pre-recorded videos and have fewer verbal pauses . |
Similar Papers
BehancePR: A Punctuation Restoration Dataset for Livestreaming Video Transcript (2022.findings-naacl)
Copied to clipboard
| Challenge: | a growing number of livestreaming videos provide useful knowledge with exceptional visual demonstrations. |
| Approach: | They propose a human-annotated corpus for punctuation restoration in livestreaming video transcripts . they show popular natural language processing tools underperform on sentence boundary detection . |
| Outcome: | The proposed dataset shows that natural language processing tools underperform on sentence boundary detection on livestreaming video transcripts. |
BehanceQA: A New Dataset for Identifying Question-Answer Pairs in Video Transcripts (2022.lrec-1)
Copied to clipboard
| Challenge: | Question-Answer (QA) is an effective method for storing knowledge . prior QA identification systems have been limited to formal written documents . a large-scale QA dataset annotated by human over 500 hours of video transcripts is a challenge . |
| Approach: | They present a large-scale QA identification dataset annotated by human over 500 hours of video transcripts. |
| Outcome: | The proposed dataset presents unique challenges for existing methods . it shows that the annotated dataset presents challenges for new methods - the results will be released . |
Event Extraction in Video Transcripts (2022.coling-1)
Copied to clipboard
| Challenge: | Existing EE datasets are limited to formally written documents such as news articles or scientific papers . existing EE methods and datasets cannot be used in informal and noisy texts . |
| Approach: | They propose to use video transcripts as a dataset for event extraction . they demonstrate that existing state-of-the-art EE methods cannot achieve adequate performance . |
| Outcome: | The proposed dataset evaluates state-of-the-art EE methods on streamed videos on Behance . it shows that such systems cannot achieve adequate performance on the proposed dataset . |
StreamHover: Livestream Transcript Summarization and Annotation (2021.emnlp-main)
Copied to clipboard
Sangwoo Cho, Franck Dernoncourt, Tim Ganter, Trung Bui, Nedim Lipka, Walter Chang, Hailin Jin, Jonathan Brandt, Hassan Foroosh, Fei Liu
| Challenge: | StreamHover is a framework for annotating and summarizing livestream transcripts . the problem is that there is n't enough annotated datasets to summarize livestreams based on the informal nature of spoken language . |
| Approach: | They propose a framework for annotating and summarizing livestream transcripts using a text preview. |
| Outcome: | The proposed model generalizes better and improves over strong baselines. |
A Stacking-based Efficient Method for Toxic Language Detection on Live Streaming Chat (2022.emnlp-industry)
Copied to clipboard
| Challenge: | Existing methods for toxic language detection are based on deep learning, but they are not scalable considering inference speed and computational resources. |
| Approach: | They propose a method for toxic language detection that is aware of real-world scenarios by partial stacking partial stacks that feeds initial results with low confidence to meta-classifier. |
| Outcome: | The proposed method achieves faster inference speed than BERT-based models with comparable performance. |
A Survey on Detection of LLMs-Generated Content (2024.findings-emnlp)
Copied to clipboard
Xianjun Yang, Liangming Pan, Xuandong Zhao, Haifeng Chen, Linda Petzold, William Yang Wang, Wei Cheng
| Challenge: | Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance . |
| Approach: | They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy. |
| Outcome: | The proposed model will be able to detect human-written content in real time. |
LifeQA: A Real-life Dataset for Video Question Answering (2020.lrec-1)
Copied to clipboard
Santiago Castro, Mahmoud Azab, Jonathan Stroud, Cristina Noujaim, Ruoyao Wang, Jia Deng, Rada Mihalcea
| Challenge: | Existing video question answering datasets consist of movies and TV shows, but they are not representative of our day-to-day lives. |
| Approach: | They propose a benchmark dataset for video question answering that focuses on day-to-day situations. |
| Outcome: | The proposed dataset analyzes the challenging but realistic aspects of LifeQA . it consists of video clips and over 2.3k multiple-choice questions . |
LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming (2023.acl-long)
Copied to clipboard
| Challenge: | a recent study shows that open-domain dialogue systems are not able to perform well in fast-growing scenarios such as live streaming due to the domain gap between online-post constructed data and those required in downstream conversational tasks. |
| Approach: | They propose to train a conversational agent based on large social media datasets with multiple domains to improve response in live streaming scenarios. |
| Outcome: | The proposed model improves response modeling and addressee recognition in live open-domain scenarios. |
TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection (L18-1)
Copied to clipboard
| Challenge: | Detecting novelty of an entire document is an AI frontier problem . present state-of-the-art text matching techniques are unable to process such redundancy. |
| Approach: | They propose a document-level novelty detection resource that can be used to benchmark techniques . they crawl news documents across several domains and use it to find out whether they contain new information . |
| Outcome: | The proposed dataset is compared with a standard system for document novelty detection . the proposed system can detect elements that have not appeared before, or new or original . |
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties (P18-1)
Copied to clipboard
| Challenge: | a lack of understanding of the properties of sentence embeddings is limiting the use of the techniques. |
| Approach: | They propose 10 probing tasks designed to capture simple linguistic features of sentences . they use three different encoders to train embeddings in eight different ways . |
| Outcome: | The proposed tasks capture key linguistic features of sentences, but they are difficult to infer from them. |