Papers by Jan Deriu
On the Effectiveness of Automated Metrics for Text Generation Systems (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation methods lack a sound theoretical foundation for evaluation campaigns . imperfect automated metrics and insufficiently sized test sets are some of the factors that cause uncertainty. |
| Approach: | They propose a theoretical framework that incorporates different sources of uncertainty, such as imperfect automated metrics and insufficiently sized test sets. |
| Outcome: | The proposed model can be leveraged to improve evaluation protocols regarding reliability, robustness, and significance of the evaluation outcome. |
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)
Copied to clipboard
Michel Plüss, Manuela Hürlimann, Marc Cuny, Alla Stöckli, Nikolaos Kapotis, Julia Hartmann, Malgorzata Anna Ulasik, Christian Scheller, Yanick Schraner, Amit Jain, Jan Deriu, Mark Cieliebak, Manfred Vogel
| Challenge: | Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it. |
| Approach: | They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems . |
| Outcome: | The dataset allows for training speech translation, dialect recognition, and speech synthesis systems. |
A Methodology for Creating Question Answering Corpora Using Inverse Data Annotation (2020.acl-main)
Copied to clipboard
Jan Deriu, Katsiaryna Mlynchyk, Philippe Schläpfer, Alvaro Rodrigo, Dirk von Grünigen, Nicolas Kaiser, Kurt Stockinger, Eneko Agirre, Mark Cieliebak
| Challenge: | Existing methods to efficiently construct corpus for question answering over structured data are time-consuming and cost-intensive. |
| Approach: | They propose a method to efficiently construct a corpus for question answering over structured data. |
| Outcome: | The proposed method triples the annotation speed while maintaining complexity of queries. |
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems (2020.emnlp-main)
Copied to clipboard
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak
| Challenge: | Lack of time efficient and reliable evalu-ation methods is hampering the development of conversational dialogue systems (chatbots). |
| Approach: | They propose a framework that replaces human-bot conversations with conversations between bots and an annotation tool that ranks chatbots based on their ability to mimic human behaviour. |
| Outcome: | The proposed evaluation framework replaces human-bot conversations with bot conversations and allows for frequent evaluations of chatbots during their evaluation cycle. |
SB-CH: A Swiss German Corpus with Sentiment Annotations (L18-1)
Copied to clipboard
| Challenge: | Using sentiment annotations, we find no corpus for written Swiss German, which is considered low-resourced due to its non-official status and phonetic differences. |
| Approach: | They propose to annotate a Swiss German corpus with sentiment annotations for sentiment analysis using Facebook comments and online chats. |
| Outcome: | The proposed corpus consists of more than 200,000 phrases and 1843 phrases with labels positive, negative, or neutral. |
Dialect Transfer for Swiss German Speech Translation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a study of Swiss German speech translation systems focuses on dialect diversity and differences between Swiss German and Standard German. |
| Approach: | They focus on the impact of dialect diversity and differences between Swiss German and Standard German . they first review the Swiss German dialect landscape and the differences to Standard German. |
| Outcome: | The proposed model is based on the Swiss German dialect landscape and differences to Standard German. |
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)
Copied to clipboard
Michel Plüss, Jan Deriu, Yanick Schraner, Claudio Paonessa, Julia Hartmann, Larissa Schmidt, Christian Scheller, Manuela Hürlimann, Tanja Samardžić, Manfred Vogel, Mark Cieliebak
| Challenge: | We present a corpus of Swiss German speech annotated with Standard German text at the sentence level. |
| Approach: | They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them . |
| Outcome: | The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date . |
DoQA - Accessing Domain-Specific FAQs via Conversational QA (2020.acl-main)
Copied to clipboard
| Challenge: | a dataset of 2,437 dialogues and 10,917 QA pairs is used to access domain-specific FAQ information. |
| Approach: | They present a dataset with 2,437 dialogues and 10,917 QA pairs for FAQs . they use the Wizard of Oz method with crowdsourcing to create dialogues using the original post and the original reply. |
| Outcome: | The proposed system can access domain-specific FAQ information without training data. |
Correction of Errors in Preference Ratings from Automated Metrics for Text Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation methods are over-confident in assigning significant differences between systems . Currently, the most reliable evaluation methods for text generation are human-based evaluations. |
| Approach: | They propose to combine human ratings with automated ratings to reduce the amount of human ratings needed to arrive at robust results. |
| Outcome: | The proposed evaluation protocol reduces the amount of human ratings by 50% while yielding the same evaluation outcome as the pure human evaluation in 95% of cases. |
Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation (2024.acl-long)
Copied to clipboard
| Challenge: | Generative AI systems are becoming ubiquitous for all kinds of modalities . evaluation of generated outputs is increasingly difficult due to cost and complexity of human evaluations. |
| Approach: | They propose to evaluate preference ratings on sign accuracy and favoritism . they propose to use automated metrics to assess generated outputs . |
| Outcome: | The proposed evaluations of preference ratings rely on correlation to human judgments or sign accuracy scores, but this does not tell the whole story. |
Probing the Robustness of Trained Metrics for Conversational Dialogue Systems (2022.acl-short)
Copied to clipboard
| Challenge: | Existing methods for evaluating conversational dialogue systems have been shown to be inefficient and instabile. |
| Approach: | They propose an adversarial method to stress-test trained metrics for evaluation of conversational dialogue systems using Reinforcement Learning. |
| Outcome: | The proposed method outperforms existing methods and can be applied to stress-test trained metrics for conversational dialogue systems. |
Error-preserving Automatic Speech Recognition of Young English Learners’ Language (2024.acl-long)
Copied to clipboard
| Challenge: | State-of-the-art speech recognition models are often trained on adult read-aloud data by native speakers and do not transfer well to young language learners’ speech. |
| Approach: | They propose to use an automated speech recognition module to train language learners' speaking skills on spontaneous speech by young language learners. |
| Outcome: | The proposed model improves on 85 hours of English audio spoken by Swiss learners and preserves their mistakes. |