Papers by Christian Bentz

4 papers
From characters to words: the turning point of BPE merges (2021.eacl-main)

Copied to clipboard

Challenge: morphological complexity is still a major challenge for NLP and the study of language.
Approach: They perform a cross-linguistic comparison following incremental merges of BPE for 47 diverse languages.
Outcome: The results show that language distributions are similar under specific levels of tokenization.
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)

Copied to clipboard

Challenge: a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages .
Approach: They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity.
Outcome: The proposed measure can be used to identify the types of languages that are not represented in a data set.
TeDDi Sample: Text Data Diversity Sample for Language Comparison and Multilingual NLP (2022.lrec-1)

Copied to clipboard

Challenge: a deep understanding of language is being achieved through increased access to data from minority and low-resource languages.
Approach: They present a diversity sample of text data for language comparison and multilingual natural language processing.
Outcome: The TeDDi sample features 89 languages based on the typological diversity sample in the World Atlas of Language Structures .
Grammatical error detection in transcriptions of spoken English (2020.coling-main)

Copied to clipboard

Challenge: CrowdED corpus of spoken English monologues on business topics was crowdsourced from native speakers of English and learners of English with German as their first language.
Approach: They propose to use the corpus recordings to correct existing speech transcriptions and edit them to make them more fluent.
Outcome: The proposed transcription corrections and annotations can be used for automatic transcription post-editing and grammatical error correction for spoken English.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations