Challenge: French Absolute Beginner corpus is intended for the development and study of Computer-Assisted Pronunciation Training (CAPT) tools for absolute beginner learners.
Approach: They introduce the French Absolute Beginner (FAB) speech corpus which is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners.
Outcome: The proposed corpus is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners.

Similar Papers

Automatic Pronunciation Assessment - A Review (2023.findings-emnlp)

Copied to clipboard

Challenge: Pronunciation assessment and its application in computer-aided pronunciation training (CAPT) have seen impressive progress in recent years.
Approach: They review methods employed in computer-aided pronunciation training for both phonemic and prosodic pronunciations.
Outcome: The proposed system should be able to automatically score non-native speech segments and give meaningful feedback.
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names.
Approach: They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations .
Outcome: The French TreeBank is the main source of morphosyntactic and syntactical annotations for French.
FrSemCor: Annotating a French Corpus with Supersenses (2020.lrec-1)

Copied to clipboard

Challenge: a new project aims to provide a sense-annotated corpus of French for NLP and linguistics research . the project uses WordNet Unique Beginners as semantic tags to provide interoperability .
Approach: They propose to use WordNet Unique Beginners as semantic tags to annotate French nouns . the project aims to provide a gold standard resource for linguistics and linguistic research .
Outcome: The proposed resource is released online under a Creative Commons license.
The French-Algerian Code-Switching Triggered audio corpus (FACST) (L18-1)

Copied to clipboard

Challenge: The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody .
Approach: They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language .
Outcome: The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies.
CBFC: a parallel L2 speech corpus for Korean and French learners (L18-1)

Copied to clipboard

Challenge: Using corpora for second language acquisition has become more and more common . corporata are used to study morpho-syntactic phenomena in English as a foreign language .
Approach: They propose to use a bilingual corpus of French learners of Korean and Korean learners of French to provide a translated and annotated corpus to the scientific community.
Outcome: The proposed corpus can be used for a wide array of purposes in the field of theoretical but also applied linguistics.
What Has LeBenchmark Learnt about French Syntax? (2024.lrec-main)

Copied to clipboard

Challenge: Pretrained acoustic models are increasingly used for downstream speech tasks such as automatic speech recognition, speech translation, spoken language understanding or speech parsing.
Approach: They propose to probing a pretrained acoustic model for French for syntactic information using the Orféo treebank.
Outcome: The proposed model is trained on 7k hours of spoken French and obtained reasonable results on tasks that require higher level linguistic knowledge.
ODIL_Syntax: a Free Spontaneous Spoken French Treebank Annotated with Constituent Trees (2020.lrec-1)

Copied to clipboard

Challenge: ODIL Syntax is a French treebank built on spontaneous speech transcripts . the structure of every speech turn is represented by constituent trees .
Approach: They propose a French treebank built on spontaneous speech transcripts with a constituency tree representation.
Outcome: The proposed treebank is based on the French TreeBank, with some annotation guidelines . the proposed tree bank will be freely distributed by January 2020 under a Creative Commons licence .
Developing a Benchmark for Pronunciation Feedback: Creation of a Phonemically Annotated Speech Corpus of isiZulu Language Learner Speech (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora for computer-assisted pronunciation training (CAPT) do not apply well to research in pronunciation feedback.
Approach: They propose to create a corpus of isiZulu language learner speech that has been annotated for phoneme errors and suprasegmental errors in tone.
Outcome: The proposed corpus is comprised of gold standard recordings from isiZulu teachers and recordings from students that have been annotated for pronunciation errors.
What data should I include in my POS tagging training set? (2025.findings-emnlp)

Copied to clipboard

Challenge: POS tagging is a crucial task for descriptive linguistics and language documentation . POS tags are not available in all languages, but are used for training sets for understudied languages .
Approach: They compare POS tagging with in-context learning, active learning, and random sampling . they find that POS can deliver reasonable results for communities with limited resources .
Outcome: The proposed training set for Indigenous and endangered languages performs better than random sampling.
FABRA: French Aggregator-Based Readability Assessment toolkit (2022.lrec-1)

Copied to clipboard

Challenge: a large number of readability predictor variables are used to predict reading difficulty of texts . the most important predictors for native texts are lexical diversity, dependency counts and text coherence .
Approach: They propose a readability toolkit based on aggregation of readability predictor variables . they show which features are most predictive on two different corpora .
Outcome: The proposed toolkit improves performance over standard feature-based readability prediction.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations