Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu, Rezvaneh Rezapour, Hamed Bonab, Maria Eskevich, Gareth Jones, Jussi Karlgren, Ben Carterette, Rosie Jones
| Challenge: | Podcasts are a large and growing repository of spoken audio. |
| Approach: | They propose to use podcasts as a resource for speech processing and linguistics . they use a corpus of 100,000 podcasts to study the complexity of the domain . |
| Outcome: | The Spotify Podcast Dataset is the largest corpus of transcribed speech data . the dataset contains 60,000 hours of podcasts, with a range of genres and styles . |
Similar Papers
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus (2025.acl-long)
Copied to clipboard
| Challenge: | a dataset of over 1.1M podcast transcripts is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020. |
| Approach: | They propose to build a large-scale open dataset of podcast transcripts that includes metadata, speaker roles, audio features and speaker turns for a subset of 370K episodes. |
| Outcome: | The proposed dataset is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020. |
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented. |
| Approach: | They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods. |
| Outcome: | The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents. |
Identifying Narrative Content in Podcast Transcripts (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing methods to study narrativity in novels, social media and patient records are limited. |
| Approach: | They propose to process podcast transcripts and extract narrative content from podcasts . they use annotations to enable future research into narrativity within a large corpus of podcast episodes. |
| Outcome: | The proposed methods compare to existing methods and can enable future research into narrativity within a large corpus of approximately 100,000 podcast episodes. |
A Repository of Corpora for Summarization (L18-1)
Copied to clipboard
| Challenge: | Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task. |
| Approach: | They propose a repository containing corpora available to train and evaluate automatic summarization systems. |
| Outcome: | The proposed system is based on a repository of corpora available for summarization tasks. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Khan Academy Corpus: A Multilingual Corpus of Khan Academy Lectures (2024.lrec-main)
Copied to clipboard
| Challenge: | a dataset of 10122 hours in 87394 recordings is presented in a new journal . 43% of recordings have human-written subtitles, covering a total of 137 languages. |
| Approach: | They present a Khan Academy corpus with 10122 hours in 87394 recordings . 43% of recordings have human-written subtitles, and 137 languages are included . |
| Outcome: | The dataset can be used to train multilingual speech recognition and translation models. |
MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing datasets for dialogue summarization are limited to their small sizes and are built from a narrow domain. |
| Approach: | They propose a large-scale media interview dataset consisting of 463.6K transcripts with abstractive summaries. |
| Outcome: | The proposed dataset is larger and contains multi-party conversations from multiple domains. |
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Podcast script generation is a challenging task for large language models, but evaluation resources are limited. |
| Approach: | They propose a benchmark to evaluate podcast script generation using a multifaceted evaluation framework . PodBench is a prototype that integrates quantitative constraints with LLM-based quality assessment . |
| Outcome: | The proposed framework integrates quantitative constraints with LLM-based quality assessment. |
WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and Discourse (D18-1)
Copied to clipboard
| Challenge: | a corpus of 43 million atomic edits is available for Wikipedia edit history . edits are instances in which a human editor has inserted a single contiguous phrase into, or deleted a contigous phrase from, an existing sentence. |
| Approach: | They use Wikipedia edit history to mine atomic edits across 8 languages . they find edits contain instances in which a human editor has inserted a single phrase into, or deleted a contiguous phrase from, an existing sentence. |
| Outcome: | The data show that edits differ from the language observed in standard corpora and that models trained on edits encode different aspects of semantics and discourse than models trained in raw text. |
The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning: data updates, training and evaluation tools (2022.lrec-1)
Copied to clipboard
Gaëlle Laperrière, Valentin Pelloin, Antoine Caubrière, Salima Mdhaffar, Nathalie Camelin, Sahar Ghannay, Bassam Jabaian, Yannick Estève
| Challenge: | a growing number of studies address the spoken language understanding domain through a simple task like speech intent detection. |
| Approach: | They focus on the french MEDIA SLU dataset, which is distributed since 2005 . they propose a recipe for its use, including data preparation, training and evaluation scripts . |
| Outcome: | The MEDIA SLU dataset is used as a benchmark dataset for a large number of research projects. |