CoANZSE Audio: Creation of an Online Corpus for Linguistic and Phonetic Analysis of Australian and New Zealand Englishes (2024.lrec-main)
Copied to clipboard
| Challenge: | CoANZSE Audio is a searchable online corpus of 195 million words of geo-located YouTube transcripts of local government channels. |
| Approach: | They describe the methods used to create the corpus from open-source tools and the architecture of the CoANZSE Audio website. |
| Outcome: | The corpus contains 195-million-word transcripts of local government channels . it is one of the first large, free, fully searchable online corpora containing data suitable for acoustic phonetic analyses in addition to lexical, grammatical, and discourse properties of Australian and New Zealand Englishes. |
Similar Papers
100,000 Podcasts: A Spoken English Document Corpus (2020.coling-main)
Copied to clipboard
Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu, Rezvaneh Rezapour, Hamed Bonab, Maria Eskevich, Gareth Jones, Jussi Karlgren, Ben Carterette, Rosie Jones
| Challenge: | Podcasts are a large and growing repository of spoken audio. |
| Approach: | They propose to use podcasts as a resource for speech processing and linguistics . they use a corpus of 100,000 podcasts to study the complexity of the domain . |
| Outcome: | The Spotify Podcast Dataset is the largest corpus of transcribed speech data . the dataset contains 60,000 hours of podcasts, with a range of genres and styles . |
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)
Copied to clipboard
| Challenge: | Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Approach: | They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora. |
| Outcome: | The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Open-source Multi-speaker Corpora of the English Accents in the British Isles (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a dataset of high-quality audio, the authors examine the accents of 120 volunteers in the British Isles. |
| Approach: | They present a dataset of high-quality audio of English sentences recorded by volunteers with different accents of the British Isles. |
| Outcome: | The transcribed audio includes pronunciations of global locations, major airlines and common personal names in different accents. |
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)
Copied to clipboard
| Challenge: | Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues. |
| Approach: | They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms. |
| Outcome: | The proposed corpus includes parallel text and speech data of 21 Japanese dialects. |
Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech (2024.lrec-main)
Copied to clipboard
| Challenge: | The Dutch Dialect Database contains dialectal variations of Dutch recorded in the second half of the twentieth century. |
| Approach: | They propose to create a corpus containing audio recordings and orthographic transcriptions of Dutch dialects recorded in the second half of the 20th century. |
| Outcome: | The Dutch Dialect Database contains dialectal variations recorded all over the Netherlands in the second half of the twentieth century. |
VoxCommunis: A Corpus for Cross-linguistic Phonetic Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Until recently, the movement towards large-scale cross-linguistic phonetic research has been limited. |
| Approach: | They propose to use the VoxCommunis Corpus to facilitate cross-linguistic phonetic research . corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . |
| Outcome: | The VoxCommunis Corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . the corpus is free to download and use under a CC0 license . |
Seshat: a Tool for Managing and Verifying Annotation Campaigns of Audio Data (2020.lrec-1)
Copied to clipboard
Hadrien Titeux, Rachid Riad, Xuan-Nga Cao, Nicolas Hamilakis, Kris Madden, Alejandrina Cristia, Anne-Catherine Bachoud-Lévi, Emmanuel Dupoux
| Challenge: | Seshat is a software for the automated management of annotation campaigns for audio/speech data. |
| Approach: | They propose a system for the automated management of annotation campaigns for audio/speech data which addresses these challenges. |
| Outcome: | The proposed system computes an associated inter-annotator agreement with the gamma measure taking into account the categorisation and segmentation discrepancies. |
The ACQDIV Corpus Database and Aggregation Pipeline (2020.lrec-1)
Copied to clipboard
| Challenge: | ACQDIV corpus database and aggregation pipeline aims to identify universal cognitive processes that allow children to acquire any language. |
| Approach: | They present the ACQDIV corpus database and aggregation pipeline . the tool aims to identify universal cognitive processes that allow children to acquire any language . |
| Outcome: | The ACQDIV corpus database and aggregation pipeline is a tool developed by the European Research Council . the database represents 15 corpora from 14 typologically maximally diverse languages . |
AET: Web-based Adjective Exploration Tool for German (L18-1)
Copied to clipboard
| Challenge: | AET enables research on the modificational behavior of German adjectives and adverbs . currently available online corpus query tools for German do not lend themselves specifically to research on adjectives - e.g., syntactic relationships or morphological properties. |
| Approach: | They propose a web-based corpus query tool that can be used to query German corpus . they extracted modifiers and modifiees from a print media corpus and stored them in a database . |
| Outcome: | The proposed tool can be transferred to other languages and modification phenomena. |