Papers by Jussi Karlgren

2 papers
100,000 Podcasts: A Spoken English Document Corpus (2020.coling-main)

Copied to clipboard

Challenge: Podcasts are a large and growing repository of spoken audio.
Approach: They propose to use podcasts as a resource for speech processing and linguistics . they use a corpus of 100,000 podcasts to study the complexity of the domain .
Outcome: The Spotify Podcast Dataset is the largest corpus of transcribed speech data . the dataset contains 60,000 hours of podcasts, with a range of genres and styles .
Challenging the Assumption of Structure-based embeddings in Few- and Zero-shot Knowledge Graph Completion (2022.lrec-1)

Copied to clipboard

Challenge: Existing work on Knowledge Graph completion only uses textual descriptive data . knowledge graphs are incomplete because not every relation has been observed at the time of their construction.
Approach: They propose to use textual descriptive data to enrich benchmark data sets for Few- and Zero-shot Knowledge Graph completion tasks.
Outcome: The proposed task improves for Few- and Zero-shot scenarios with up to twofold increase in the Zero- shot setting.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations