Challenge: In this paper we use network theory to model graphs of child-directed speech from caregivers of children from nine typologically and morphologically diverse languages.
Approach: They use network theory to model child-directed speech from caregivers of children from nine typologically and morphologically diverse languages.
Outcome: The proposed model adds to the repertoire of universal distributional patterns found in the input to children cross-linguistically.

Similar Papers

Learning from Child-directed Speech in Two-language Scenarios: A French-English Case-Study (2026.findings-eacl)

Copied to clipboard

Challenge: a systematic study of compact language models with limited computational resources is challenging for many research contexts and real-world applications.
Approach: They extend BabyBERTa to English-French scenarios under strictly sizematched data conditions.
Outcome: The proposed model extends to English-French scenarios under sizematched data conditions . the results show context-dependent effects of multilingual training .
On the Relation between Linguistic Typology and (Limitations of) Multilingual Language Modeling (D18-1)

Copied to clipboard

Challenge: a key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language.
Approach: They propose to use a full-vocabulary setup to test the performance of language modeling (LM) on 50 typologically diverse languages.
Outcome: The proposed language modeling task is based on a full vocabulary setup focused on word-level prediction on 50 typologically diverse languages.
Is Word Segmentation Child’s Play in All Languages? (P19-1)

Copied to clipboard

Challenge: Existing word learning strategies for infants are cross-linguistically robust . infants do not know which language(s) will be found in their environment at the beginning of development .
Approach: They propose to use 11 conceptually diverse algorithms to learn word-like units in infants . they propose to employ cross-linguistically robust algorithms that can be used by all infants.
Outcome: The proposed algorithms perform above chance on 8 different languages . the results show that some of the algorithms are cross-linguistically valid .
Measuring the perceptual availability of phonological features during language acquisition using unsupervised binary stochastic autoencoders (N19-1)

Copied to clipboard

Challenge: Xitsonga and English are typologically unrelated languages . phonological features are not directly observed by humans .
Approach: They deploy binary stochastic neural autoencoder networks as models of infant language learning in two typologically unrelated languages.
Outcome: The proposed model is well represented in both languages, while others are less so.
Morphological Complexity of Children Narratives in Eight Languages (2022.lrec-1)

Copied to clipboard

Challenge: morphological complexity of a corpus representing the language production of younger and older children is compared across different languages.
Approach: a study compares morphological complexity of a corpus representing language production of younger and older children across different languages.
Outcome: The results show that younger children corpora have lower morphological complexity than older children corpus for Spanish and Russian.
Multilingual Transfer Learning for Children Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in automatic speech recognition (ASR) systems have been criticized for high acoustic variability and limited amount of available training data.
Approach: They propose a two-step training strategy that uses multilingual learning followed by language-specific transfer learning to generalize children's speech.
Outcome: The proposed training strategy outperforms single language training and multilingual and transfer learning alone in English.
The Linguistic Connectivities Within Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have discovered notable disparities in their performance across different languages.
Approach: They conduct a systematic investigation into the behaviors of large language models across 27 different languages on 3 different scenarios and reveals a Linguistic Map correlates with the richness of available resources and linguistic family relations.
Outcome: The proposed model demonstrates that there are significant disparities in performance across languages across 27 different languages on 3 different scenarios.
Is Child-Directed Speech Effective Training Data for Language Models? (2024.emnlp-main)

Copied to clipboard

Challenge: High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data.
Approach: They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset.
Outcome: The proposed models show that child language input is not valuable for training language models.
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)

Copied to clipboard

Challenge: linguistic typology is the classification of languages according to their linguistic properties.
Approach: They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale.
Outcome: The proposed model can predict typological properties on a massively multilingual scale.
Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: prevailing view in language acquisition research has long held that child-directed language is more effective than adultdirected language (ADL)
Approach: They propose a frequency-controlled testing methodology to enable balanced comparisons across training corpora.
Outcome: The proposed method outperforms models trained on English Child-Directed Language (CDL) but it does not yield stronger generalizations for acquiring syntax.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations