FAB: The French Absolute Beginner Corpus for Pronunciation Training (2020.lrec-1)
Copied to clipboard
| Challenge: | French Absolute Beginner corpus is intended for the development and study of Computer-Assisted Pronunciation Training (CAPT) tools for absolute beginner learners. |
| Approach: | They introduce the French Absolute Beginner (FAB) speech corpus which is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners. |
| Outcome: | The proposed corpus is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners. |
Similar Papers
Automatic Pronunciation Assessment - A Review (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Pronunciation assessment and its application in computer-aided pronunciation training (CAPT) have seen impressive progress in recent years. |
| Approach: | They review methods employed in computer-aided pronunciation training for both phonemic and prosodic pronunciations. |
| Outcome: | The proposed system should be able to automatically score non-native speech segments and give meaningful feedback. |
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names. |
| Approach: | They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations . |
| Outcome: | The French TreeBank is the main source of morphosyntactic and syntactical annotations for French. |
FrSemCor: Annotating a French Corpus with Supersenses (2020.lrec-1)
Copied to clipboard
Lucie Barque, Pauline Haas, Richard Huyghe, Delphine Tribout, Marie Candito, Benoit Crabbé, Vincent Segonne
| Challenge: | a new project aims to provide a sense-annotated corpus of French for NLP and linguistics research . the project uses WordNet Unique Beginners as semantic tags to provide interoperability . |
| Approach: | They propose to use WordNet Unique Beginners as semantic tags to annotate French nouns . the project aims to provide a gold standard resource for linguistics and linguistic research . |
| Outcome: | The proposed resource is released online under a Creative Commons license. |
The French-Algerian Code-Switching Triggered audio corpus (FACST) (L18-1)
Copied to clipboard
| Challenge: | The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody . |
| Approach: | They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language . |
| Outcome: | The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies. |
CBFC: a parallel L2 speech corpus for Korean and French learners (L18-1)
Copied to clipboard
| Challenge: | Using corpora for second language acquisition has become more and more common . corporata are used to study morpho-syntactic phenomena in English as a foreign language . |
| Approach: | They propose to use a bilingual corpus of French learners of Korean and Korean learners of French to provide a translated and annotated corpus to the scientific community. |
| Outcome: | The proposed corpus can be used for a wide array of purposes in the field of theoretical but also applied linguistics. |
What Has LeBenchmark Learnt about French Syntax? (2024.lrec-main)
Copied to clipboard
| Challenge: | Pretrained acoustic models are increasingly used for downstream speech tasks such as automatic speech recognition, speech translation, spoken language understanding or speech parsing. |
| Approach: | They propose to probing a pretrained acoustic model for French for syntactic information using the Orféo treebank. |
| Outcome: | The proposed model is trained on 7k hours of spoken French and obtained reasonable results on tasks that require higher level linguistic knowledge. |
ODIL_Syntax: a Free Spontaneous Spoken French Treebank Annotated with Constituent Trees (2020.lrec-1)
Copied to clipboard
| Challenge: | ODIL Syntax is a French treebank built on spontaneous speech transcripts . the structure of every speech turn is represented by constituent trees . |
| Approach: | They propose a French treebank built on spontaneous speech transcripts with a constituency tree representation. |
| Outcome: | The proposed treebank is based on the French TreeBank, with some annotation guidelines . the proposed tree bank will be freely distributed by January 2020 under a Creative Commons licence . |
Developing a Benchmark for Pronunciation Feedback: Creation of a Phonemically Annotated Speech Corpus of isiZulu Language Learner Speech (2024.lrec-main)
Copied to clipboard
Alexandra O’Neil, Nils Hjortnaes, Francis Tyers, Zinhle Nkosi, Thulile Ndlovu, Zanele Mlondo, Ngami Phumzile Pewa
| Challenge: | Existing corpora for computer-assisted pronunciation training (CAPT) do not apply well to research in pronunciation feedback. |
| Approach: | They propose to create a corpus of isiZulu language learner speech that has been annotated for phoneme errors and suprasegmental errors in tone. |
| Outcome: | The proposed corpus is comprised of gold standard recordings from isiZulu teachers and recordings from students that have been annotated for pronunciation errors. |
What data should I include in my POS tagging training set? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | POS tagging is a crucial task for descriptive linguistics and language documentation . POS tags are not available in all languages, but are used for training sets for understudied languages . |
| Approach: | They compare POS tagging with in-context learning, active learning, and random sampling . they find that POS can deliver reasonable results for communities with limited resources . |
| Outcome: | The proposed training set for Indigenous and endangered languages performs better than random sampling. |
FABRA: French Aggregator-Based Readability Assessment toolkit (2022.lrec-1)
Copied to clipboard
Rodrigo Wilkens, David Alfter, Xiaoou Wang, Alice Pintard, Anaïs Tack, Kevin P. Yancey, Thomas François
| Challenge: | a large number of readability predictor variables are used to predict reading difficulty of texts . the most important predictors for native texts are lexical diversity, dependency counts and text coherence . |
| Approach: | They propose a readability toolkit based on aggregation of readability predictor variables . they show which features are most predictive on two different corpora . |
| Outcome: | The proposed toolkit improves performance over standard feature-based readability prediction. |