Papers by Gilles Adda

4 papers
Establishing degrees of closeness between audio recordings along different dimensions using large-scale cross-lingual models (2024.findings-eacl)

Copied to clipboard

Challenge: Existing methods to analyze speech representations using pretraining data are difficult to achieve for endangered languages.
Approach: They propose an unsupervised method to examine the level of abstraction in vector representations of speech from a pretrained model to determine their level of abstractness.
Outcome: The proposed method is fully unsupervised and could be used in comparative studies on under-documented languages.
A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments (L18-1)

Copied to clipboard

Challenge: a new study aims to document endangered languages using a speech corpus . linguistic documentation is limited to the phonetic, lexical and syntactic levels .
Approach: They propose to use a speech corpus to document endangered languages in field . they propose to collect 5k speech utterances aligned to French text translations .
Outcome: The proposed language corpus is used to document endangered languages in field linguists . it is multilingual and contains 5k speech utterances aligned to french text translations - the authors show it can be used in a zero-resource task .
BULBasaa: A Bilingual Basaa-French Speech Corpus for the Evaluation of Language Documentation Tools (L18-1)

Copied to clipboard

Challenge: Approximately 50 hours of Bàsàá speech were collected and then carefully re-spoken and orally translated into French .
Approach: They propose to provide an automatic phonetic transcription using a set of derived phone-like units.
Outcome: The proposed method provides an automatic phonetic transcription using a set of derived phone-like units.
Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville) (L18-1)

Copied to clipboard

Challenge: BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages.
Approach: This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project.
Outcome: The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations