Papers by Michael Gref

5 papers
A Study on the Ambiguity in Human Annotation of German Oral History Interviews for Perceived Emotion Recognition and Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Sentiment analysis and emotion recognition can help research in audiovisual interview archives . however, humans perceive sentiments and emotions ambiguously and subjectively .
Approach: They investigate human perceptions of emotions and sentiments in oral history interviews . they show that human perception for different emotions is ambiguous and subjective . authors propose deep learning as a way to categorize and search emotions .
Outcome: The proposed techniques can be used to search and index audiovisual interviews . the authors show that human perceptions differ for different emotions .
Multitask Learning for Grapheme-to-Phoneme Conversion of Anglicisms in German Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Anglicisms are a challenge in German speech recognition due to their irregular pronunciation compared to native German words.
Approach: They propose a multitask sequence-to-sequence approach for grapheme-tophoneme conversion to improve the phonetization of Anglicisms.
Outcome: The proposed model reduces the word error rate by 1 % and the Anglicism error rate, while still maintaining the accuracy of the baseline model.
Improved Transcription and Indexing of Oral History Interviews for Digital Humanities Research (L18-1)

Copied to clipboard

Challenge: Existing methods to improve transcription and indexing quality of Oral History interviews are not available.
Approach: They propose to use a German Oral History test-set to improve transcription and indexing quality . they propose to combine acoustic modeling techniques with sophisticated neural networks .
Outcome: The proposed system reduces word error rate by 28.3% on German Oral History test-set compared to baseline system . the Fraunhofer IAIS Audio Mining system can process long audio-files to automatically create time-aligned transcriptions.
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible.
Approach: They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries.
Outcome: The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country.
Multi-Staged Cross-Lingual Acoustic Model Adaption for Robust Speech Recognition in Real-World Applications - A Case Study on German Oral History Interviews (2020.lrec-1)

Copied to clipboard

Challenge: Current automatic speech recognition systems show remarkable performance when adequate data is used for training.
Approach: They propose to perform a robust acoustic model adaption to a target domain in a cross-lingual manner.
Outcome: The proposed approach reduces word error rate by more than 30% on German oral history interviews compared to a model trained from scratch on the target domain and 6-7% on same-language out-of-domain training data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations