Papers by Christian Bentz
From characters to words: the turning point of BPE merges (2021.eacl-main)
Copied to clipboard
| Challenge: | morphological complexity is still a major challenge for NLP and the study of language. |
| Approach: | They perform a cross-linguistic comparison following incremental merges of BPE for 47 diverse languages. |
| Outcome: | The results show that language distributions are similar under specific levels of tokenization. |
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)
Copied to clipboard
| Challenge: | a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages . |
| Approach: | They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity. |
| Outcome: | The proposed measure can be used to identify the types of languages that are not represented in a data set. |
TeDDi Sample: Text Data Diversity Sample for Language Comparison and Multilingual NLP (2022.lrec-1)
Copied to clipboard
| Challenge: | a deep understanding of language is being achieved through increased access to data from minority and low-resource languages. |
| Approach: | They present a diversity sample of text data for language comparison and multilingual natural language processing. |
| Outcome: | The TeDDi sample features 89 languages based on the typological diversity sample in the World Atlas of Language Structures . |
Grammatical error detection in transcriptions of spoken English (2020.coling-main)
Copied to clipboard
| Challenge: | CrowdED corpus of spoken English monologues on business topics was crowdsourced from native speakers of English and learners of English with German as their first language. |
| Approach: | They propose to use the corpus recordings to correct existing speech transcriptions and edit them to make them more fluent. |
| Outcome: | The proposed transcription corrections and annotations can be used for automatic transcription post-editing and grammatical error correction for spoken English. |