Towards Processing of the Oral History Interviews and Related Printed Documents (L18-1)
Copied to clipboard
Zbyněk Zajíc, Lucie Skorkovská, Petr Neduchal, Pavel Ircing, Josef V. Psutka, Marek Hrúz, Aleš Pražák, Daniel Soutner, Jan Švec, Lukáš Bureš, Luděk Müller
| Challenge: | a project aims to create an integrated archive of the recordings, scanned documents and photographs from totalitarian regimes in Czechoslovakia . the archive will be accessible online and provide multifaceted search capabilities . |
| Approach: | They propose to use automatic speech recognition and optical character recognition to build an archive of the interviews, scanned documents and photographs. |
| Outcome: | The proposed archive will be accessible online and provide multifaceted search capabilities. |
Similar Papers
Improved Transcription and Indexing of Oral History Interviews for Digital Humanities Research (L18-1)
Copied to clipboard
| Challenge: | Existing methods to improve transcription and indexing quality of Oral History interviews are not available. |
| Approach: | They propose to use a German Oral History test-set to improve transcription and indexing quality . they propose to combine acoustic modeling techniques with sophisticated neural networks . |
| Outcome: | The proposed system reduces word error rate by 28.3% on German Oral History test-set compared to baseline system . the Fraunhofer IAIS Audio Mining system can process long audio-files to automatically create time-aligned transcriptions. |
A Bird’s-eye View of Language Processing Projects at the Romanian Academy (L18-1)
Copied to clipboard
| Challenge: | a recent article outlines five projects that address contemporary Romanian language . the authors argue that a constant accumulation of human expertise is needed to develop complex projects. |
| Approach: | a new article gives a general overview of five AI language-related projects at the Romanian Academy . they focus on the creation of a contemporary Romanian language text and speech corpus and language related applications . |
| Outcome: | a new article gives an overview of five AI language-related projects at the Romanian Academy . the projects address contemporary Romanian language, as well as language related applications . |
A CLARIN Transcription Portal for Interview Data (2020.lrec-1)
Copied to clipboard
| Challenge: | a transcription portal for audio files based on automatic speech recognition (ASR) is implemented in the CLARIN resources research network and intended for use by non-technical scholars. |
| Approach: | They propose a transcription portal for audio files based on automatic speech recognition in various languages. |
| Outcome: | The proposed transcription portal is implemented in the CLARIN resources research network and intended for use by non-technical scholars. |
A Recorded Debating Dataset (L18-1)
Copied to clipboard
Shachar Mirkin, Michal Jacovi, Tamar Lavee, Hong-Kwang Kuo, Samuel Thomas, Leslie Sager, Lili Kotlerman, Elad Venezian, Noam Slonim
| Challenge: | Existing research in computational argumentation and debating technologies focuses on argumentation mining, but other tasks are being addressed as well. |
| Approach: | They describe a dataset of debating speeches in English that is used for research . they use an automatic speech recognition system to produce a more "nLP-friendly" text . |
| Outcome: | The proposed dataset contains 60 speeches on various controversial topics, each in five formats corresponding to different stages in production. |
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible. |
| Approach: | They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries. |
| Outcome: | The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country. |
Text Mining for History: first steps on building a large dataset (L18-1)
Copied to clipboard
| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
Automatic Orality Identification in Historical Texts (2020.lrec-1)
Copied to clipboard
| Challenge: | a set of general linguistic features are used to identify conceptually-oral historical texts . linguists recognize that there is also a lot of variation within discourse modes . |
| Approach: | They propose to use general linguistic features to identify conceptually-oral historical texts . they find they are useful for determining conceptuality of historical data as for modern data . |
| Outcome: | The proposed features are used to identify conceptually-oral historical German texts . the features are useful in determining conceptuality of historical data as they are for modern data . |
An Application for Building a Polish Telephone Speech Corpus (L18-1)
Copied to clipboard
| Challenge: | Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance. |
| Approach: | They propose to build a tool for speech corpus collection of a specific domain content. |
| Outcome: | The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks. |
Large Corpus of Czech Parliament Plenary Hearings (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Czech parliament plenary sessions is a valuable resource for future research . only a few public datasets are available in the Czech language . end-to-end approaches require extensive training data to produce competitive results . |
| Approach: | They present a corpus of Czech parliament plenary sessions which is a large corpus . they combine a traditional approach with a more traditional approach . |
| Outcome: | The proposed model architectures can be used to train and evaluate speech recognition systems on a large corpus of speech data and transcripts. |
Creating a Data Set of Abstractive Summaries of Turn-labeled Spoken Human-Computer Conversations (2022.lrec-1)
Copied to clipboard
| Challenge: | Digital recorded written and spoken dialogues are becoming more available due to the growing popularity of online messenger services and chatbots. |
| Approach: | They propose to use Dutch spoken human-computer conversations, an annotation layer of turn labels, and conversational abstractive summaries of user answers to build a conversational agent. |
| Outcome: | The proposed system can be integrated into a conversational agent. |