CHICA: A Developmental Corpus of Child-Caregiver’s Face-to-face vs. Video Call Conversations in Middle Childhood (2024.lrec-main)
Copied to clipboard
Dhia Elhak Goumri, Abhishek Agrawal, Mitja Nikolaus, Hong Duc Thang Vu, Kübra Bodur, Elias Emmar, Cassandre Armand, Chiara Mazzocconi, Shreejata Gupta, Laurent Prévot, Benoit Favre, Leonor Becerra-Bonache, Abdellah Fourtassi
| Challenge: | Existing studies of language-in-interaction focus on the two ends of the developmental spectrum, i.e., early childhood and adulthood, leaving a gap in our knowledge about how development unfolds, especially across middle childhood. |
| Approach: | They propose to use CHICA to analyze child-caregiver conversations at home . they use mobile, lightweight eye-tracking and head motion detection to optimize the naturalness of the recordings. |
| Outcome: | The proposed corpus of child-caregiver conversations at home was compared with a previous corpus based on a set of conversations between children aged 7, 9, and 11 years old. |
Similar Papers
Automatic Coding of Contingency in Child-Caregiver Conversations (2024.lrec-main)
Copied to clipboard
| Challenge: | Current research on children's language development relies on manual annotation of a small sample of children, which limits our ability to draw general conclusions about development. |
| Approach: | They propose to use automatic tools to assess contingency in children's natural interactions with caregivers by annotating a small set of data with a Transformer-based model. |
| Outcome: | The proposed model replicates existing results and generates new data-driven hypotheses. |
From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology (2022.acl-long)
Copied to clipboard
| Challenge: | Informal social interaction is the primordial home of human language. |
| Approach: | They show that linguistically diverse conversational corpora can provide empirical foundations for flexible, localizable language technologies of the future. |
| Outcome: | The results suggest that even relatively small corpora can support robust generalizations about key aspects of interactional infrastructure. |
Automatic Annotation of Grammaticality in Child-Caregiver Conversations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for analyzing child language acquisition have been tedious and inconsistent. |
| Approach: | They propose a coding scheme for context-dependent grammaticality in child-caregiver conversations and annotate 4,000 utterances from a large corpus of transcribed conversations. |
| Outcome: | The proposed method achieves human inter-annotation agreement levels and is faster and reproducible than manual methods. |
The Badalona Corpus - An Audio, Video and Neuro-Physiological Conversational Dataset (2022.lrec-1)
Copied to clipboard
Philippe Blache, Salomé Antoine, Dorina De Jong, Lena-Marie Huttner, Emilia Kerr, Thierry Legou, Eliot Maës, Clément François
| Challenge: | Using the same dyads at different periods, we can study the evolution of interlocutors’ alignment during the time. |
| Approach: | They propose to record 5 dyads with all modalities and neuro-physiological signals in a natural conversation corpus. |
| Outcome: | The proposed corpus is the first to capture all modalities and neuro-physiological signals in a natural conversation situation. |
Representing the Toddler Lexicon: Do the Corpus and Semantics Matter? (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on child language development have relied on adult-based measures to model their lexicons. |
| Approach: | They propose to use transcripts of child-directed conversations, picture books and dialog from G-rated movies to approximate the language input a North American preschooler might hear. |
| Outcome: | The proposed model outperforms models based on the existing corpus and the existing model. |
ChildTalk: A Multi-Dialect Chinese Child Speech Corpus with Full-Length Child–Caregiver Conversations for Speech Recognition (2026.findings-acl)
Copied to clipboard
| Challenge: | Automatic speech recognition (ASR) for children remains challenging due to developmental variability and the scarcity of high-quality corpora. |
| Approach: | They propose a large-scale Chinese child speech corpus that contains 112.5 hours of speech from 498 children and 500 caregivers. |
| Outcome: | The proposed model improves in-domain and cross-domain performance on children's speech. |
Is Child-Directed Speech Effective Training Data for Language Models? (2024.emnlp-main)
Copied to clipboard
| Challenge: | High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data. |
| Approach: | They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset. |
| Outcome: | The proposed models show that child language input is not valuable for training language models. |
Learning from Child-directed Speech in Two-language Scenarios: A French-English Case-Study (2026.findings-eacl)
Copied to clipboard
| Challenge: | a systematic study of compact language models with limited computational resources is challenging for many research contexts and real-world applications. |
| Approach: | They extend BabyBERTa to English-French scenarios under strictly sizematched data conditions. |
| Outcome: | The proposed model extends to English-French scenarios under sizematched data conditions . the results show context-dependent effects of multilingual training . |
Morphological Complexity of Children Narratives in Eight Languages (2022.lrec-1)
Copied to clipboard
Gordana Hržica, Chaya Liebeskind, Kristina Š. Despot, Olga Dontcheva-Navratilova, Laura Kamandulytė-Merfeldienė, Sara Košutar, Matea Kramarić, Giedrė Valūnaitė Oleškevičienė
| Challenge: | morphological complexity of a corpus representing the language production of younger and older children is compared across different languages. |
| Approach: | a study compares morphological complexity of a corpus representing language production of younger and older children across different languages. |
| Outcome: | The results show that younger children corpora have lower morphological complexity than older children corpus for Spanish and Russian. |
ChiSense-12: An English Sense-Annotated Child-Directed Speech Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent evidence suggests that the speech children hear early in development is rich in word sense ambiguity, and also that children's early vocabularies are populated by ambiguous words. |
| Approach: | They sense-tagged 53 corpora of American and English speech directed to 958 target children up to 59 months of age and selected target senses that they know young children understand. |
| Outcome: | The sense-tagged corpus ChiSense-12 was used to examine the role of verb-event structure in child word sense disambiguation. |