Papers by Ann Clifton
Transforming Podcast Preview Generation: From Expert Models to LLM-Based Systems (2025.acl-industry)
Copied to clipboard
| Challenge: | Podcasts, videos, and other long-form talk content requires significant time investment to assess their relevance. |
| Approach: | They propose an LLM-based approach for generating podcast episode previews and deploy it at scale, serving hundreds of thousands of podcast previews in a real-world application. |
| Outcome: | The proposed approach outperforms a baseline built on top of various ML expert models and offers a 4.6% increase in user engagement with preview content and a 5x boost in processing efficiency. |
100,000 Podcasts: A Spoken English Document Corpus (2020.coling-main)
Copied to clipboard
Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu, Rezvaneh Rezapour, Hamed Bonab, Maria Eskevich, Gareth Jones, Jussi Karlgren, Ben Carterette, Rosie Jones
| Challenge: | Podcasts are a large and growing repository of spoken audio. |
| Approach: | They propose to use podcasts as a resource for speech processing and linguistics . they use a corpus of 100,000 podcasts to study the complexity of the domain . |
| Outcome: | The Spotify Podcast Dataset is the largest corpus of transcribed speech data . the dataset contains 60,000 hours of podcasts, with a range of genres and styles . |