Challenge: a new approach to search for sound in large archives is being developed . speech and speech technology researchers struggle to access large amounts of data .
Approach: They propose a method for fast and efficient non-sequential browsing of sound in large archives that we know little about . they combine audio browsing through massively multi-object sound environments and an unsupervised dimensionality reduction algorithm to search for sound in public archives.
Outcome: The proposed method is shown to combine well, resulting in rapid and interpretable observations.

Similar Papers

100,000 Podcasts: A Spoken English Document Corpus (2020.coling-main)

Copied to clipboard

Challenge: Podcasts are a large and growing repository of spoken audio.
Approach: They propose to use podcasts as a resource for speech processing and linguistics . they use a corpus of 100,000 podcasts to study the complexity of the domain .
Outcome: The Spotify Podcast Dataset is the largest corpus of transcribed speech data . the dataset contains 60,000 hours of podcasts, with a range of genres and styles .
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus (2025.acl-long)

Copied to clipboard

Challenge: a dataset of over 1.1M podcast transcripts is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020.
Approach: They propose to build a large-scale open dataset of podcast transcripts that includes metadata, speaker roles, audio features and speaker turns for a subset of 370K episodes.
Outcome: The proposed dataset is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
A database of German definitory contexts from selected web sources (L18-1)

Copied to clipboard

Challenge: a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts.
Approach: They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction .
Outcome: The proposed method is based on a web corpus and a robust pattern-based extraction method.
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
Synthetic Data Made to Order: The Case of Parsing (D18-1)

Copied to clipboard

Challenge: supervised dependency parsing is a core task in natural language processing, but unsupervised parsers can hardly produce useful parses.
Approach: They propose to permute the constituents of an existing dependency treebank so that its surface part-of-speech statistics approximately match those of the target language.
Outcome: The proposed method improves the parsing accuracy of a target language . the proposed method is based on a distribution of gold POS bigrams .
Improved Transcription and Indexing of Oral History Interviews for Digital Humanities Research (L18-1)

Copied to clipboard

Challenge: Existing methods to improve transcription and indexing quality of Oral History interviews are not available.
Approach: They propose to use a German Oral History test-set to improve transcription and indexing quality . they propose to combine acoustic modeling techniques with sophisticated neural networks .
Outcome: The proposed system reduces word error rate by 28.3% on German Oral History test-set compared to baseline system . the Fraunhofer IAIS Audio Mining system can process long audio-files to automatically create time-aligned transcriptions.
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
A Large Resource of Patterns for Verbal Paraphrases (L18-1)

Copied to clipboard

Challenge: Xu et al., 2015: paraphrases play an important role in natural language understanding . he says it is difficult to propose a paraphrasing relation for natural language processing systems .
Approach: They propose a resource of such paraphrases that can be used to identify hidden paraphrase pairs . they propose to use the resource to identify paraphrase relationships between two words .
Outcome: The proposed resource contains tens of thousands of such pairs and is available for academic purposes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations