Challenge: Existing studies of language-in-interaction focus on the two ends of the developmental spectrum, i.e., early childhood and adulthood, leaving a gap in our knowledge about how development unfolds, especially across middle childhood.
Approach: They propose to use CHICA to analyze child-caregiver conversations at home . they use mobile, lightweight eye-tracking and head motion detection to optimize the naturalness of the recordings.
Outcome: The proposed corpus of child-caregiver conversations at home was compared with a previous corpus based on a set of conversations between children aged 7, 9, and 11 years old.

Similar Papers

Automatic Coding of Contingency in Child-Caregiver Conversations (2024.lrec-main)

Copied to clipboard

Challenge: Current research on children's language development relies on manual annotation of a small sample of children, which limits our ability to draw general conclusions about development.
Approach: They propose to use automatic tools to assess contingency in children's natural interactions with caregivers by annotating a small set of data with a Transformer-based model.
Outcome: The proposed model replicates existing results and generates new data-driven hypotheses.
From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology (2022.acl-long)

Copied to clipboard

Challenge: Informal social interaction is the primordial home of human language.
Approach: They show that linguistically diverse conversational corpora can provide empirical foundations for flexible, localizable language technologies of the future.
Outcome: The results suggest that even relatively small corpora can support robust generalizations about key aspects of interactional infrastructure.
Automatic Annotation of Grammaticality in Child-Caregiver Conversations (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for analyzing child language acquisition have been tedious and inconsistent.
Approach: They propose a coding scheme for context-dependent grammaticality in child-caregiver conversations and annotate 4,000 utterances from a large corpus of transcribed conversations.
Outcome: The proposed method achieves human inter-annotation agreement levels and is faster and reproducible than manual methods.
The Badalona Corpus - An Audio, Video and Neuro-Physiological Conversational Dataset (2022.lrec-1)

Copied to clipboard

Challenge: Using the same dyads at different periods, we can study the evolution of interlocutors’ alignment during the time.
Approach: They propose to record 5 dyads with all modalities and neuro-physiological signals in a natural conversation corpus.
Outcome: The proposed corpus is the first to capture all modalities and neuro-physiological signals in a natural conversation situation.
Representing the Toddler Lexicon: Do the Corpus and Semantics Matter? (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on child language development have relied on adult-based measures to model their lexicons.
Approach: They propose to use transcripts of child-directed conversations, picture books and dialog from G-rated movies to approximate the language input a North American preschooler might hear.
Outcome: The proposed model outperforms models based on the existing corpus and the existing model.
ChildTalk: A Multi-Dialect Chinese Child Speech Corpus with Full-Length Child–Caregiver Conversations for Speech Recognition (2026.findings-acl)

Copied to clipboard

Challenge: Automatic speech recognition (ASR) for children remains challenging due to developmental variability and the scarcity of high-quality corpora.
Approach: They propose a large-scale Chinese child speech corpus that contains 112.5 hours of speech from 498 children and 500 caregivers.
Outcome: The proposed model improves in-domain and cross-domain performance on children's speech.
Is Child-Directed Speech Effective Training Data for Language Models? (2024.emnlp-main)

Copied to clipboard

Challenge: High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data.
Approach: They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset.
Outcome: The proposed models show that child language input is not valuable for training language models.
Learning from Child-directed Speech in Two-language Scenarios: A French-English Case-Study (2026.findings-eacl)

Copied to clipboard

Challenge: a systematic study of compact language models with limited computational resources is challenging for many research contexts and real-world applications.
Approach: They extend BabyBERTa to English-French scenarios under strictly sizematched data conditions.
Outcome: The proposed model extends to English-French scenarios under sizematched data conditions . the results show context-dependent effects of multilingual training .
Morphological Complexity of Children Narratives in Eight Languages (2022.lrec-1)

Copied to clipboard

Challenge: morphological complexity of a corpus representing the language production of younger and older children is compared across different languages.
Approach: a study compares morphological complexity of a corpus representing language production of younger and older children across different languages.
Outcome: The results show that younger children corpora have lower morphological complexity than older children corpus for Spanish and Russian.
ChiSense-12: An English Sense-Annotated Child-Directed Speech Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Recent evidence suggests that the speech children hear early in development is rich in word sense ambiguity, and also that children's early vocabularies are populated by ambiguous words.
Approach: They sense-tagged 53 corpora of American and English speech directed to 958 target children up to 59 months of age and selected target senses that they know young children understand.
Outcome: The sense-tagged corpus ChiSense-12 was used to examine the role of verb-event structure in child word sense disambiguation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations