Xuân-Nga Cao, Cyrille Dakhlia, Patricia Del Carmen, Mohamed-Amine Jaouani, Malik Ould-Arbi, Emmanuel Dupoux
| Challenge: | a platform for capturing, storing and analyzing day-long audio recordings and photos of children's linguistic environments is proposed . the proposed platform connects families and academics, with strong innovation potential for each type of users. |
| Approach: | They propose a platform for capturing, storing and analyzing audio recordings and photos of children's linguistic environments. |
| Outcome: | The proposed platform connects families and academics with strong innovation potential for each type of users. |
Similar Papers
lingvis.io - A Linguistic Visual Analytics Framework (P19-3)
Copied to clipboard
Mennatallah El-Assady, Wolfgang Jentner, Fabian Sperrle, Rita Sevastjanova, Annette Hautli-Janisz, Miriam Butt, Daniel Keim
| Challenge: | Using a modular framework, linguistic visual analytics applications can be rapidly prototypized using a web-based framework. |
| Approach: | They propose a modular framework for rapid prototyping of linguistic, web-based, visual analytics applications. |
| Outcome: | The proposed framework supports rapid prototyping of linguistic, web-based, visual analytics applications. |
KidLM: Advancing Language Models for Children – Early Insights and Future Directions (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have been shown to be effective in creating educational tools for children, yet there are significant challenges in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards. |
| Approach: | They propose a user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children. |
| Outcome: | The proposed model excels in understanding lower grade-level text, maintains safety by avoiding stereotypes, and captures children’s unique preferences. |
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data (2026.eacl-long)
Copied to clipboard
Jaap Jumelet, Abdellah Fourtassi, Akari Haga, Bastian Bunzeck, Bhargav Shandilya, Diana Galvan-Sosa, Faiz Ghifari Haznitrama, Francesca Padovani, Francois Meyer, Hai Hu, Julen Etxaniz, Laurent Prevot, Linyang He, María Grandury, Mila Marcheva, Negar Foroutan, Nikitas Theodoropoulos, Pouya Sadeghi, Siyuan Song, Suchir Salhan, Susana Zhou, Yurii Paniv, Ziyin Zhang, Arianna Bisazza, Alex Warstadt, Leshem Choshen
| Challenge: | prevailing trend in language modeling research is to prioritize scaling, authors say . from infancy to maturity, English learners acquire language through exposure to less than 100M words . |
| Approach: | They propose a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. |
| Outcome: | The proposed models outperform models trained on a fixed, developmentally plausible English corpus on various benchmarks. |
Is Word Segmentation Child’s Play in All Languages? (P19-1)
Copied to clipboard
| Challenge: | Existing word learning strategies for infants are cross-linguistically robust . infants do not know which language(s) will be found in their environment at the beginning of development . |
| Approach: | They propose to use 11 conceptually diverse algorithms to learn word-like units in infants . they propose to employ cross-linguistically robust algorithms that can be used by all infants. |
| Outcome: | The proposed algorithms perform above chance on 8 different languages . the results show that some of the algorithms are cross-linguistically valid . |
Social Web Observatory: A Platform and Method for Gathering Knowledge on Entities from Different Textual Sources (2020.lrec-1)
Copied to clipboard
| Challenge: | a framework for gathering entity-centered information is needed in real-life scenarios . a social web observatory system allows users to define their own entities . |
| Approach: | They propose a framework for the collection and summarization of information from the Web in an entity-driven manner. |
| Outcome: | The proposed framework is based on a language analysis pipeline and a human user study. |
Is Child-Directed Speech Effective Training Data for Language Models? (2024.emnlp-main)
Copied to clipboard
| Challenge: | High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data. |
| Approach: | They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset. |
| Outcome: | The proposed models show that child language input is not valuable for training language models. |
ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5 (2025.acl-long)
Copied to clipboard
Jiaming Zhou, Shiyao Wang, Shiwan Zhao, Jiabei He, Haoqin Sun, Hui Wang, Cheng Liu, Aobo Kong, Yujie Guo, Xi Yang, Yequan Wang, Yonghua Lin, Yong Qin
| Challenge: | Automatic speech recognition systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0. |
| Approach: | They propose to use Mandarin speech datasets to analyze pronunciation and tone of children aged 3 to 5 and evaluate their models on speaker verification (SV) They find that the datasets are more robust than those used by adult speech recognition systems and are open-source and available for all academic purposes. |
| Outcome: | The proposed dataset includes 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. |
The ACQDIV Corpus Database and Aggregation Pipeline (2020.lrec-1)
Copied to clipboard
| Challenge: | ACQDIV corpus database and aggregation pipeline aims to identify universal cognitive processes that allow children to acquire any language. |
| Approach: | They present the ACQDIV corpus database and aggregation pipeline . the tool aims to identify universal cognitive processes that allow children to acquire any language . |
| Outcome: | The ACQDIV corpus database and aggregation pipeline is a tool developed by the European Research Council . the database represents 15 corpora from 14 typologically maximally diverse languages . |
CHICA: A Developmental Corpus of Child-Caregiver’s Face-to-face vs. Video Call Conversations in Middle Childhood (2024.lrec-main)
Copied to clipboard
Dhia Elhak Goumri, Abhishek Agrawal, Mitja Nikolaus, Hong Duc Thang Vu, Kübra Bodur, Elias Emmar, Cassandre Armand, Chiara Mazzocconi, Shreejata Gupta, Laurent Prévot, Benoit Favre, Leonor Becerra-Bonache, Abdellah Fourtassi
| Challenge: | Existing studies of language-in-interaction focus on the two ends of the developmental spectrum, i.e., early childhood and adulthood, leaving a gap in our knowledge about how development unfolds, especially across middle childhood. |
| Approach: | They propose to use CHICA to analyze child-caregiver conversations at home . they use mobile, lightweight eye-tracking and head motion detection to optimize the naturalness of the recordings. |
| Outcome: | The proposed corpus of child-caregiver conversations at home was compared with a previous corpus based on a set of conversations between children aged 7, 9, and 11 years old. |
The Road to Success: Assessing the Fate of Linguistic Innovations in Online Communities (C18-1)
Copied to clipboard
| Challenge: | a longitudinal study of online social networks investigates the birth and spread of lexical innovations. |
| Approach: | They investigate the birth and diffusion of lexical innovations in online communities . they build on sociolinguistic theories and focus on the relationship between the spread of a new term and the social role of the individuals who use it . |
| Outcome: | The proposed method predicts whether an innovation will succeed in a community. |