Challenge: High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data.
Approach: They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset.
Outcome: The proposed models show that child language input is not valuable for training language models.

Similar Papers

Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: prevailing view in language acquisition research has long held that child-directed language is more effective than adultdirected language (ADL)
Approach: They propose a frequency-controlled testing methodology to enable balanced comparisons across training corpora.
Outcome: The proposed method outperforms models trained on English Child-Directed Language (CDL) but it does not yield stronger generalizations for acquiring syntax.
Language acquisition: do children and language models follow similar learning stages? (2023.findings-acl)

Copied to clipboard

Challenge: During language acquisition, children follow a typical sequence of learning stages, whereby they first learn to categorize phonemes before they develop their lexicon and eventually master complex syntactic structures.
Approach: They train 48 GPT-2 models from scratch and evaluate their syntactic and semantic abilities at each training step using 96 probes curated from the BLiMP, Zorro and BIG-Bench benchmarks.
Outcome: The proposed model exhibits similar learning trajectories to human children aged between 18 months and 6 years.
Learning to Understand Child-directed and Adult-directed Speech (2020.acl-main)

Copied to clipboard

Challenge: linguistic properties of child-directed speech differ from adult-directed in many ways . linguistic differences between CDS and ADS are retained, but the acoustic properties are similar.
Approach: They compare the task performance of models trained on adult-directed speech and child-directed language . they propose that CDS is optimized for learnability, but not for comprehension .
Outcome: The proposed model trains on adult-directed speech and child-directed language . the model generalizes better on the training register and on synthesized speech .
Evaluating Large Language Models on Controlled Generation Tasks (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies have looked into the ability of large language models in various benchmark tasks, including question generation, reading comprehension, multilingual and etc. However, few studies investigate the controllability of large languages.
Approach: They propose to compare large language models with state-of-the-start finetuned smaller models to find that large language model controls are comparable to smaller models.
Outcome: The proposed model can meet hard constraints and perform better than state-of-the-art models.
KidLM: Advancing Language Models for Children – Early Insights and Future Directions (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models have been shown to be effective in creating educational tools for children, yet there are significant challenges in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards.
Approach: They propose a user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children.
Outcome: The proposed model excels in understanding lower grade-level text, maintains safety by avoiding stereotypes, and captures children’s unique preferences.
Multilingual Transfer Learning for Children Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in automatic speech recognition (ASR) systems have been criticized for high acoustic variability and limited amount of available training data.
Approach: They propose a two-step training strategy that uses multilingual learning followed by language-specific transfer learning to generalize children's speech.
Outcome: The proposed training strategy outperforms single language training and multilingual and transfer learning alone in English.
Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck? (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit a puzzling disparity in their formal linguistic competence, even after training on trillions of tokens.
Approach: They pre-train Large Language Models on 100M-token corpora and inject a minimal amount of synthetic data targeting specific linguistic phenomena into the model.
Outcome: The proposed intervention significantly improves model performance in 8 out of the 9 worst-performing BLiMP paradigms.
On the Automatic Generation and Simplification of Children’s Stories (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have made it possible to generate children's educational texts with appropriate lexical and readability levels.
Approach: They first examine the ability of several popular LLMs to generate stories with properly adjusted lexical and readability levels.
Outcome: The proposed models can generalize to the domain of children's stories and create an efficient pipeline for their automatic generation.
Evaluating and Improving Child-Directed Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that adult speech recognition systems are lagging behind child models due to the fact that children's vocal tracts are smaller than adults .
Approach: They evaluate a model that trains on adult data and apply additional tuning to varied amounts of child speech data to improve child-directed speech recognition.
Outcome: The proposed model improves over baseline models using child data and small amounts of child audio data.
Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies have explored using large language models to generate synthetic datasets . however, the effectiveness of the LLM-generated synthetic data is inconsistent across different classification tasks.
Approach: They propose to use large language models to generate synthetic datasets to better understand factors that moderate the effectiveness of LLM-generated synthetic data.
Outcome: The results show that subjectivity is negatively associated with the performance of the model trained on synthetic data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations