Papers by Mark Cieliebak

19 papers
TRANSLIT: A Large-scale Name Transliteration Resource (2020.lrec-1)

Copied to clipboard

Challenge: Transliteration is the process of expressing a proper name from a source language in the characters of a target language.
Approach: They present a large-scale corpus of transliterated names in 180 languages . they use machine learning to train automatic transliteration .
Outcome: The proposed system achieves 92% accuracy on identification of transliterated pairs.
ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos (2025.emnlp-main)

Copied to clipboard

Challenge: Existing efforts in misinformation detection focus on written text, leaving a significant gap in addressing the complexity of spoken text in video transcripts.
Approach: They propose to annotate video transcripts in three languages and six topics using a custom annotation tool.
Outcome: The proposed tool shows strong cross-validation performance but challenges for generalization to unseen topics.
On the Effectiveness of Automated Metrics for Text Generation Systems (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation methods lack a sound theoretical foundation for evaluation campaigns . imperfect automated metrics and insufficiently sized test sets are some of the factors that cause uncertainty.
Approach: They propose a theoretical framework that incorporates different sources of uncertainty, such as imperfect automated metrics and insufficiently sized test sets.
Outcome: The proposed model can be leveraged to improve evaluation protocols regarding reliability, robustness, and significance of the evaluation outcome.
CEASR: A Corpus for Evaluating Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are increasingly needed for research and practical applications.
Approach: They propose to use public speech corpora to evaluate the quality of automatic speech recognition (ASR) they calculate an average Word Error Rate (WER) per corpus, per system and per corpor-system pair .
Outcome: The proposed corpus evaluates the quality of automatic speech recognition systems using public speech corpora and transcripts generated by state-of-the-art systems.
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it.
Approach: They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems .
Outcome: The dataset allows for training speech translation, dialect recognition, and speech synthesis systems.
A Methodology for Creating Question Answering Corpora Using Inverse Data Annotation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to efficiently construct corpus for question answering over structured data are time-consuming and cost-intensive.
Approach: They propose a method to efficiently construct a corpus for question answering over structured data.
Outcome: The proposed method triples the annotation speed while maintaining complexity of queries.
A Measure of the System Dependence of Automated Metrics (2025.acl-short)

Copied to clipboard

Challenge: Recent advances in machine translation evaluations are expensive and time-intensive.
Approach: They propose a method to evaluate the correlation between human and metric scores . they argue that it is equally important to ensure that metrics treat all systems fairly and consistently.
Outcome: The proposed method ignores a central requirement of the evaluation process, and ignores the need for a thorough evaluation procedure.
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems (2020.emnlp-main)

Copied to clipboard

Challenge: Lack of time efficient and reliable evalu-ation methods is hampering the development of conversational dialogue systems (chatbots).
Approach: They propose a framework that replaces human-bot conversations with conversations between bots and an annotation tool that ranks chatbots based on their ability to mimic human behaviour.
Outcome: The proposed evaluation framework replaces human-bot conversations with bot conversations and allows for frequent evaluations of chatbots during their evaluation cycle.
SB-CH: A Swiss German Corpus with Sentiment Annotations (L18-1)

Copied to clipboard

Challenge: Using sentiment annotations, we find no corpus for written Swiss German, which is considered low-resourced due to its non-official status and phonetic differences.
Approach: They propose to annotate a Swiss German corpus with sentiment annotations for sentiment analysis using Facebook comments and online chats.
Outcome: The proposed corpus consists of more than 200,000 phrases and 1843 phrases with labels positive, negative, or neutral.
Dialect Transfer for Swiss German Speech Translation (2023.findings-emnlp)

Copied to clipboard

Challenge: a study of Swiss German speech translation systems focuses on dialect diversity and differences between Swiss German and Standard German.
Approach: They focus on the impact of dialect diversity and differences between Swiss German and Standard German . they first review the Swiss German dialect landscape and the differences to Standard German.
Outcome: The proposed model is based on the Swiss German dialect landscape and differences to Standard German.
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)

Copied to clipboard

Challenge: We present a corpus of Swiss German speech annotated with Standard German text at the sentence level.
Approach: They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them .
Outcome: The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date .
DoQA - Accessing Domain-Specific FAQs via Conversational QA (2020.acl-main)

Copied to clipboard

Challenge: a dataset of 2,437 dialogues and 10,917 QA pairs is used to access domain-specific FAQ information.
Approach: They present a dataset with 2,437 dialogues and 10,917 QA pairs for FAQs . they use the Wizard of Oz method with crowdsourcing to create dialogues using the original post and the original reply.
Outcome: The proposed system can access domain-specific FAQ information without training data.
Towards Integration of Statistical Hypothesis Tests into Deep Neural Networks (P19-1)

Copied to clipboard

Challenge: Existing approaches for text classification are lexicallevel features with Naive Bayes or Support Vector Machines (SVM) .
Approach: They propose a deep-learning model that uses label descriptions to train texts and their labels for multi-label and multi-class classification tasks.
Outcome: The proposed model improves on one set with a high margin and on all other sets with competitive results.
Correction of Errors in Preference Ratings from Automated Metrics for Text Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods are over-confident in assigning significant differences between systems . Currently, the most reliable evaluation methods for text generation are human-based evaluations.
Approach: They propose to combine human ratings with automated ratings to reduce the amount of human ratings needed to arrive at robust results.
Outcome: The proposed evaluation protocol reduces the amount of human ratings by 50% while yielding the same evaluation outcome as the pure human evaluation in 95% of cases.
Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation (2024.acl-long)

Copied to clipboard

Challenge: Generative AI systems are becoming ubiquitous for all kinds of modalities . evaluation of generated outputs is increasingly difficult due to cost and complexity of human evaluations.
Approach: They propose to evaluate preference ratings on sign accuracy and favoritism . they propose to use automated metrics to assess generated outputs .
Outcome: The proposed evaluations of preference ratings rely on correlation to human judgments or sign accuracy scores, but this does not tell the whole story.
LEDGAR: A Large-Scale Multi-label Corpus for Text Classification of Legal Provisions in Contracts (2020.lrec-1)

Copied to clipboard

Challenge: Contractual provisions are a primary research target in law studies as they constitute the legal essence of a contract.
Approach: They propose to use LEDGAR to construct a multilabel corpus of legal provisions in contracts that is crawled and scraped from the public domain.
Outcome: The proposed corpus is the first freely available corpus of its kind.
Probing the Robustness of Trained Metrics for Conversational Dialogue Systems (2022.acl-short)

Copied to clipboard

Challenge: Existing methods for evaluating conversational dialogue systems have been shown to be inefficient and instabile.
Approach: They propose an adversarial method to stress-test trained metrics for evaluation of conversational dialogue systems using Reinforcement Learning.
Outcome: The proposed method outperforms existing methods and can be applied to stress-test trained metrics for conversational dialogue systems.
Do NOT Classify and Count: Hybrid Attribute Control Success Evaluation (2026.eacl-long)

Copied to clipboard

Challenge: evaluating attribute control success in controllable text generation relies on pretrained classifiers.
Approach: They propose a Bayesian method that combines classifier predictions with a small number of human labels for calibration.
Outcome: The proposed method produces robust estimates across both text and image generation tasks, offering an alternative to current evaluation practices.
Error-preserving Automatic Speech Recognition of Young English Learners’ Language (2024.acl-long)

Copied to clipboard

Challenge: State-of-the-art speech recognition models are often trained on adult read-aloud data by native speakers and do not transfer well to young language learners’ speech.
Approach: They propose to use an automated speech recognition module to train language learners' speaking skills on spontaneous speech by young language learners.
Outcome: The proposed model improves on 85 hours of English audio spoken by Swiss learners and preserves their mistakes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations