Challenge: a vocabulary commonality index is used to investigate to what extent each child acquires common words during the early stages of lexical development.
Approach: They propose a vocabulary commonality index to investigate to what extent each child acquires common words during the early stages of lexical development.
Outcome: The proposed index can be used to understand to what extent each child acquires common words during the early stages of lexical development.

Similar Papers

Representing the Toddler Lexicon: Do the Corpus and Semantics Matter? (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on child language development have relied on adult-based measures to model their lexicons.
Approach: They propose to use transcripts of child-directed conversations, picture books and dialog from G-rated movies to approximate the language input a North American preschooler might hear.
Outcome: The proposed model outperforms models based on the existing corpus and the existing model.
Word Acquisition in Neural Language Models (2022.tacl-1)

Copied to clipboard

Challenge: Language models acquire individual words during training, based on unigram token frequencies, before transitioning loosely to bigram probabilities, eventually converging on more nuanced predictions.
Approach: They examine how neural language models acquire individual words during training, extracting learning curves and ages of acquisition for over 600 words on the MacArthur-Bates Communicative Development Inventory.
Outcome: The models follow consistent patterns during training for both unidirectional and bidirectional models, and for both LSTM and Transformer architectures.
Building an English Vocabulary Knowledge Dataset of Japanese English-as-a-Second-Language Learners Using Crowdsourcing (L18-1)

Copied to clipboard

Challenge: a dataset for analyzing the English vocabulary of English-as-a-second language learners is available . a vocabulary size test was performed by 100 test takers hired via crowdsourcing .
Approach: They propose a dataset for analyzing the English vocabulary of English-as-a-second language learners.
Outcome: a dataset for analyzing the English vocabulary of English-as-a-second language learners is available online . the results show that the test is reliable and can be predicted with high accuracy .
A Large-Scale Leveled Readability Lexicon for Standard Arabic (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale leveled readability lexicon for Modern Standard Arabic is not available in many other languages.
Approach: They propose a large-scale leveled readability lexicon for Arabic with 26,000 lemmas . they manually annotate a lexico from three different regions in the arab world .
Outcome: The proposed lexicon is publicly available for Arabic readability tasks.
PoKi: A Large Dataset of Poems by Children (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of child-written texts is available for study of child language . authors use non-parametric regressions to model developmental differences from early childhood to late-adolescence .
Approach: They propose to analyze 62 thousand child-written poems written by children from grades 1 to 12 . they use non-parametric regressions to model developmental differences from early childhood to late-adolescence .
Outcome: The proposed corpus includes about 62 thousand poems written by children from grades 1 to 12 . results show decreases in valence that are especially pronounced during mid-adolescence .
Reading Time and Vocabulary Rating in the Japanese Language: Large-Scale Japanese Reading Time Data Collection Using Crowdsourcing (2022.lrec-1)

Copied to clipboard

Challenge: a study examines how differences in human vocabulary affect reading time . vocabulary size is inversely correlated to reading time due to the COVID-19 pandemic .
Approach: They assume that vocabulary is random effect of research participants . they then asked participants to take part in a self-paced reading task to collect reading times .
Outcome: The proposed method clarifies the tendency that vocabulary differences give to reading time.
Is Word Segmentation Child’s Play in All Languages? (P19-1)

Copied to clipboard

Challenge: Existing word learning strategies for infants are cross-linguistically robust . infants do not know which language(s) will be found in their environment at the beginning of development .
Approach: They propose to use 11 conceptually diverse algorithms to learn word-like units in infants . they propose to employ cross-linguistically robust algorithms that can be used by all infants.
Outcome: The proposed algorithms perform above chance on 8 different languages . the results show that some of the algorithms are cross-linguistically valid .
KidLM: Advancing Language Models for Children – Early Insights and Future Directions (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models have been shown to be effective in creating educational tools for children, yet there are significant challenges in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards.
Approach: They propose a user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children.
Outcome: The proposed model excels in understanding lower grade-level text, maintains safety by avoiding stereotypes, and captures children’s unique preferences.
Estimating Lexical Complexity from Document-Level Distributions (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for complexity estimation are limited to entire documents . health assessment tools are too short for existing methods to apply .
Approach: They propose a two-step approach for estimating lexical complexity that does not rely on pre-annotated data.
Outcome: The proposed method is tested on the Norwegian language and compares with other assessment tools.
Is Child-Directed Speech Effective Training Data for Language Models? (2024.emnlp-main)

Copied to clipboard

Challenge: High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data.
Approach: They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset.
Outcome: The proposed models show that child language input is not valuable for training language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations