Papers by Francis Tyers

13 papers
Common Voice: A Massively-Multilingual Speech Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development.
Approach: They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments.
Outcome: The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages.
Producing a Parallel Universal Dependencies Treebank of Ancient Hebrew and Ancient Greek via Cross-Lingual Projection (2024.lrec-main)

Copied to clipboard

Challenge: Using parallel treebanks, syntactic changes can be identified and evaluated in translations, redactions, and commentaries.
Approach: They propose to construct a treebank of Ancient Greek containing portions of the Septuagint by word-aligning and projecting from the parallel Ancient Hebrew text.
Outcome: The proposed treebank contains portions of the Hebrew Scriptures, which are translated into Ancient Greek, and is based on the results of a collaborative effort to create a crosslinguistically consistent treebank annotation scheme.
A Large-Scale Study of Machine Translation in Turkic Languages (2021.emnlp-main)

Copied to clipboard

Challenge: a large corpus covering 22 Turkic languages is included in this paper . low-resource MT evaluation has traditionally focused on European languages due to limitations of available technology and resources.
Approach: They present a case study of the practical application of MT in the Turkic language family . they propose to realize the gains of NMT for Turkic languages under high-resource to extremely low-resourced scenarios.
Outcome: The proposed study shows that the new methods can be used in the Turkic language family . the results highlight bottlenecks in building competitive systems .
A Universal Dependencies Treebank for Highland Puebla Nahuatl (2024.naacl-long)

Copied to clipboard

Challenge: Annotated linguistic corpora are essential component of natural language processing (NLP) Annotation frameworks are used for morphological and dependency-based syntactic phenomena in endangered, indigenous, and/or marginalized languages.
Approach: They propose a Universal Dependencies (UD) treebank for Highland Puebla Nahuatl . they describe the process of data collection, annotation decisions and interesting syntactic constructions .
Outcome: The proposed treebank is the second such UD treebank for a Mexican language . it is a significant addition to an existing treebank of another Nahuatl language based on the framework .
A Free/Open-Source Morphological Analyser and Generator for Sakha (2022.lrec-1)

Copied to clipboard

Challenge: a morphological transducer for Sakha is being developed for use in downstream tasks . the marginalised language is subject to increasing economic and cultural peril due to climate change .
Approach: They describe the development of a morphological analyser and generator for Sakha . the transducer has coverage of solidly above 90%, and high precision . it is already being used in downstream tasks such as linguistic maintenance .
Outcome: The proposed morphological analyser has coverage of 90% and high precision . it is already being used in computer assisted language learning applications .
Developing a Benchmark for Pronunciation Feedback: Creation of a Phonemically Annotated Speech Corpus of isiZulu Language Learner Speech (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora for computer-assisted pronunciation training (CAPT) do not apply well to research in pronunciation feedback.
Approach: They propose to create a corpus of isiZulu language learner speech that has been annotated for phoneme errors and suprasegmental errors in tone.
Outcome: The proposed corpus is comprised of gold standard recordings from isiZulu teachers and recordings from students that have been annotated for pronunciation errors.
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)

Copied to clipboard

Challenge: Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew.
Approach: They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax.
Outcome: The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax.
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
Universal Dependencies for Western Sierra Puebla Nahuatl (2022.lrec-1)

Copied to clipboard

Challenge: Annotated corpus of western Sierra Puebla Nahuatl conforms to universal dependency project annotation guidelines . morphological and syntactic phenomena can be analyzed quantitatively with a large enough corpus .
Approach: They present a morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl . it is the first indigenous language of Mexico to be added to the Universal Dependencies project . UD is a widely-used annotation framework whose aim is to provide a consistent schema for morphological and syntactic phenomena for all of the world's languages.
Outcome: The morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl conforms to the universal dependency project annotation guidelines.
A Finite-State Morphological Analyser for Evenki (2020.lrec-1)

Copied to clipboard

Challenge: Evenki is a language with rich morphology, therefore a morphological analyser is highly desirable for processing Evenki texts.
Approach: They propose to use a morphological analyser for Evenki to analyze half of the corpus . they evaluate the morphology of available corpora and estimate accuracy, recall and F-score .
Outcome: The proposed morphological analyser can analyse less than a half of the available corpora on Evenki . it is based on the Helsinki Finite-State Transducer toolkit (HFST).
Do RNN States Encode Abstract Phonological Alternations? (2021.naacl-main)

Copied to clipboard

Challenge: Sequence-to-sequence models have been successful in word formation tasks, but the opacity of the models makes it difficult to determine whether complex generalizations are learned or whether there is some level of generalization across related sound changes.
Approach: They propose to train character-based sequence-to-sequence models for inflection of Finnish nouns into the genitive case, an inflation type which is encoded in the hidden states of an LSTM encoderdecoder trained to perform word infference.
Outcome: The proposed models encode 17 different consonant gradation processes in a handful of dimensions in the RNN.
Finite-state morphological analysis for Gagauz (L18-1)

Copied to clipboard

Challenge: a finite-state approach to morphological analysis and generation of Gagauz is used . the model has a reasonable coverage over a range of freely-available corpora .
Approach: They propose a finite-state approach to morphological analysis and generation of Gagauz . they explicitly handle orthographic errors and variance, in addition to loan words .
Outcome: The proposed approach has a reasonable coverage over a range of freely-available corpora.
An Unsupervised Method for Weighting Finite-state Morphological Analyzers (2020.lrec-1)

Copied to clipboard

Challenge: Morphological analysis is one of the tasks that have been studied for years.
Approach: They propose a method for weighting a morphological analyzer built using finite state transducers in order to disambiguate its results.
Outcome: The proposed model weights a word2vec model using untagged corpora and captures the semantic meaning of the words.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations