Analyzing Vocabulary Commonality Index Using Large-scaled Database of Child Language Development (L18-1)
Copied to clipboard
| Challenge: | a vocabulary commonality index is used to investigate to what extent each child acquires common words during the early stages of lexical development. |
| Approach: | They propose a vocabulary commonality index to investigate to what extent each child acquires common words during the early stages of lexical development. |
| Outcome: | The proposed index can be used to understand to what extent each child acquires common words during the early stages of lexical development. |
Similar Papers
Representing the Toddler Lexicon: Do the Corpus and Semantics Matter? (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on child language development have relied on adult-based measures to model their lexicons. |
| Approach: | They propose to use transcripts of child-directed conversations, picture books and dialog from G-rated movies to approximate the language input a North American preschooler might hear. |
| Outcome: | The proposed model outperforms models based on the existing corpus and the existing model. |
Word Acquisition in Neural Language Models (2022.tacl-1)
Copied to clipboard
| Challenge: | Language models acquire individual words during training, based on unigram token frequencies, before transitioning loosely to bigram probabilities, eventually converging on more nuanced predictions. |
| Approach: | They examine how neural language models acquire individual words during training, extracting learning curves and ages of acquisition for over 600 words on the MacArthur-Bates Communicative Development Inventory. |
| Outcome: | The models follow consistent patterns during training for both unidirectional and bidirectional models, and for both LSTM and Transformer architectures. |
Building an English Vocabulary Knowledge Dataset of Japanese English-as-a-Second-Language Learners Using Crowdsourcing (L18-1)
Copied to clipboard
| Challenge: | a dataset for analyzing the English vocabulary of English-as-a-second language learners is available . a vocabulary size test was performed by 100 test takers hired via crowdsourcing . |
| Approach: | They propose a dataset for analyzing the English vocabulary of English-as-a-second language learners. |
| Outcome: | a dataset for analyzing the English vocabulary of English-as-a-second language learners is available online . the results show that the test is reliable and can be predicted with high accuracy . |
A Large-Scale Leveled Readability Lexicon for Standard Arabic (2020.lrec-1)
Copied to clipboard
| Challenge: | a large-scale leveled readability lexicon for Modern Standard Arabic is not available in many other languages. |
| Approach: | They propose a large-scale leveled readability lexicon for Arabic with 26,000 lemmas . they manually annotate a lexico from three different regions in the arab world . |
| Outcome: | The proposed lexicon is publicly available for Arabic readability tasks. |
PoKi: A Large Dataset of Poems by Children (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of child-written texts is available for study of child language . authors use non-parametric regressions to model developmental differences from early childhood to late-adolescence . |
| Approach: | They propose to analyze 62 thousand child-written poems written by children from grades 1 to 12 . they use non-parametric regressions to model developmental differences from early childhood to late-adolescence . |
| Outcome: | The proposed corpus includes about 62 thousand poems written by children from grades 1 to 12 . results show decreases in valence that are especially pronounced during mid-adolescence . |
Reading Time and Vocabulary Rating in the Japanese Language: Large-Scale Japanese Reading Time Data Collection Using Crowdsourcing (2022.lrec-1)
Copied to clipboard
| Challenge: | a study examines how differences in human vocabulary affect reading time . vocabulary size is inversely correlated to reading time due to the COVID-19 pandemic . |
| Approach: | They assume that vocabulary is random effect of research participants . they then asked participants to take part in a self-paced reading task to collect reading times . |
| Outcome: | The proposed method clarifies the tendency that vocabulary differences give to reading time. |
Is Word Segmentation Child’s Play in All Languages? (P19-1)
Copied to clipboard
| Challenge: | Existing word learning strategies for infants are cross-linguistically robust . infants do not know which language(s) will be found in their environment at the beginning of development . |
| Approach: | They propose to use 11 conceptually diverse algorithms to learn word-like units in infants . they propose to employ cross-linguistically robust algorithms that can be used by all infants. |
| Outcome: | The proposed algorithms perform above chance on 8 different languages . the results show that some of the algorithms are cross-linguistically valid . |
KidLM: Advancing Language Models for Children – Early Insights and Future Directions (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have been shown to be effective in creating educational tools for children, yet there are significant challenges in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards. |
| Approach: | They propose a user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children. |
| Outcome: | The proposed model excels in understanding lower grade-level text, maintains safety by avoiding stereotypes, and captures children’s unique preferences. |
Estimating Lexical Complexity from Document-Level Distributions (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for complexity estimation are limited to entire documents . health assessment tools are too short for existing methods to apply . |
| Approach: | They propose a two-step approach for estimating lexical complexity that does not rely on pre-annotated data. |
| Outcome: | The proposed method is tested on the Norwegian language and compares with other assessment tools. |
Is Child-Directed Speech Effective Training Data for Language Models? (2024.emnlp-main)
Copied to clipboard
| Challenge: | High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data. |
| Approach: | They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset. |
| Outcome: | The proposed models show that child language input is not valuable for training language models. |