Reading Time and Vocabulary Rating in the Japanese Language: Large-Scale Japanese Reading Time Data Collection Using Crowdsourcing (2022.lrec-1)
Copied to clipboard
| Challenge: | a study examines how differences in human vocabulary affect reading time . vocabulary size is inversely correlated to reading time due to the COVID-19 pandemic . |
| Approach: | They assume that vocabulary is random effect of research participants . they then asked participants to take part in a self-paced reading task to collect reading times . |
| Outcome: | The proposed method clarifies the tendency that vocabulary differences give to reading time. |
Similar Papers
You Don’t Have Time to Read This: An Exploration of Document Reading Time Prediction (2020.acl-main)
Copied to clipboard
Orion Weller, Jordan Hildebrandt, Ilya Reznik, Christopher Challis, E. Shannon Tass, Quinn Snell, Kevin Seppi
| Challenge: | Existing work on reading time prediction has focused on word level only predictions . however, previous work has focused only on word levels . |
| Approach: | They perform an experiment to examine how different features of text contribute to the time it takes to read, distributing and collecting data from over a thousand participants. |
| Outcome: | The proposed method combines a large number of machine learning methods with textual and stylistic factors to predict the time it takes to read. |
Building an English Vocabulary Knowledge Dataset of Japanese English-as-a-Second-Language Learners Using Crowdsourcing (L18-1)
Copied to clipboard
| Challenge: | a dataset for analyzing the English vocabulary of English-as-a-second language learners is available . a vocabulary size test was performed by 100 test takers hired via crowdsourcing . |
| Approach: | They propose a dataset for analyzing the English vocabulary of English-as-a-second language learners. |
| Outcome: | a dataset for analyzing the English vocabulary of English-as-a-second language learners is available online . the results show that the test is reliable and can be predicted with high accuracy . |
Exploring the Relationship Between Algorithm Performance, Vocabulary, and Run-Time in Text Classification (2021.naacl-main)
Copied to clipboard
| Challenge: | Many text classification algorithms depend on the size of the corpus’ vocabulary due to their bag-of-words representation. |
| Approach: | They propose to evaluate how preprocessing techniques affect the run-time of models by evaluating ten techniques over four models and two datasets. |
| Outcome: | The proposed methods can reduce run-time with no loss of accuracy while sacrificing up to 65%. |
Exploring the Effect of Nominal Compound Structure in Scientific Texts on Reading Times of Experts and Novices (2025.acl-srw)
Copied to clipboard
| Challenge: | Using a corpus of eye-tracking data of German native speakers, we find that some compound types are associated with longer reading times. |
| Approach: | They use a corpus containing eye-tracking data of german native speakers reading scientific texts. |
| Outcome: | The authors show that some compound types are associated with longer reading times and that experts may have an advantage while reading in-domain texts, but also while reading out-of-domain. |
The Linearity of the Effect of Surprisal on Reading Times across Languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a large amount of insight into human language processing can be gleaned by studying word-by-word processing difficulty. |
| Approach: | They extend the study by examining eyetracking corpora of seven languages . they find evidence for superlinearity in some languages, but highly sensitive to language models . |
| Outcome: | The study extends existing studies on english to Danish, Dutch, English, German, Japanese, Mandarin, and Russian. |
Construction of a Japanese Word Similarity Dataset (L18-1)
Copied to clipboard
| Challenge: | evaluating distributed word representations in languages that do not have such resources is difficult . et al., 2015: distributed word represent a sparse vector indicating the word itself or the context of the word. |
| Approach: | They constructed a Japanese word similarity dataset to evaluate distributed representations in Japanese. |
| Outcome: | a Japanese word similarity dataset is the first resource that can be used to evaluate distributed representations in Japanese . the dataset contains various parts of speech and includes rare words in addition to common words . |
Measuring and Modeling Language Change (N19-5)
Copied to clipboard
| Challenge: | This tutorial will help researchers answer questions fundamental to the social sciences and humanities . |
| Approach: | This tutorial is designed to help researchers answer questions in the social sciences and humanities . it synthesizes recent computational techniques for handling and modeling temporal data . |
| Outcome: | The tutorial will synthesize recent techniques for handling and modeling temporal data, such as dynamic word embeddings, and identify useful tools for social scientists and digital humanities scholars. |
Analyzing Vocabulary Commonality Index Using Large-scaled Database of Child Language Development (L18-1)
Copied to clipboard
| Challenge: | a vocabulary commonality index is used to investigate to what extent each child acquires common words during the early stages of lexical development. |
| Approach: | They propose a vocabulary commonality index to investigate to what extent each child acquires common words during the early stages of lexical development. |
| Outcome: | The proposed index can be used to understand to what extent each child acquires common words during the early stages of lexical development. |
Timesteps of Mamba Align with Human Reading Times (2026.findings-acl)
Copied to clipboard
| Challenge: | In Mamba, the recurrent state transition at each layer conceptually takes some duration of time, the discretization timestep t, determined dynamically in response to the input. |
| Approach: | They propose to align per-word processing time in a popular state-space language model Mamba with human reading time using a naturalistic reading dataset. |
| Outcome: | The proposed model can predict reading times comparable to baselines such as word frequency and GPT-2 surprisal and significant even when they are controlled for. |
Word Complexity Estimation for Japanese Lexical Simplification (2020.lrec-1)
Copied to clipboard
| Challenge: | Experimental results show that the proposed method achieves the highest performance of Japanese lexical simplification. |
| Approach: | They propose a large-scale word complexity lexicon, a synonym lexicone and a toolkit for developing and benchmarking Japanese lexical simplification systems. |
| Outcome: | The proposed method achieves the highest performance of Japanese lexical simplification. |