Modeling the Impact of Syntactic Distance and Surprisal on Cross-Slavic Text Comprehension (2022.lrec-1)
Copied to clipboard
| Challenge: | Using symmetric measures of insertion, deletion and movement of syntactic units, we investigate phonetic and orthographic asymmetries between selected languages. |
| Approach: | They focus on the syntactic variation and measure syntaktic distances between nine Slavic languages using symmetric measures of insertion, deletion and movement of syntak units in parallel sentences of the fable “The North Wind and the Sun”. |
| Outcome: | The proposed measures are validated on spoken and written cloze tests for Slavic native speakers to determine whether variations in syntax lead to slower or impeded intercomprehension of Slav texts. |
Similar Papers
The Effects of Surprisal across Languages: Results from Native and Non-native Reading (2022.findings-aacl)
Copied to clipboard
| Challenge: | Context-dependent predictive processes have been proposed as a core component of the human cognitive system. |
| Approach: | They extract surprisal estimates from mBERT and assess their predictive power on the MECO corpus, a cross-linguistic dataset of eye movement behavior in reading. |
| Outcome: | The proposed model is based on a cross-linguistic dataset of eye movement behavior in reading. |
The Linearity of the Effect of Surprisal on Reading Times across Languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a large amount of insight into human language processing can be gleaned by studying word-by-word processing difficulty. |
| Approach: | They extend the study by examining eyetracking corpora of seven languages . they find evidence for superlinearity in some languages, but highly sensitive to language models . |
| Outcome: | The study extends existing studies on english to Danish, Dutch, English, German, Japanese, Mandarin, and Russian. |
DistaLs: a Comprehensive Collection of Language Distance Measures (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Existing work on how to measure distances between languages has focused on intuition and typological distance. |
| Approach: | They propose a toolkit that provides users with easy access to language distance measures. |
| Outcome: | The proposed toolkit provides easy access to a wide variety of language distance measures. |
Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets (2026.eacl-long)
Copied to clipboard
| Challenge: | Novel metaphor comprehension involves complex semantic processes and linguistic creativity. |
| Approach: | They propose a cloze-style surprisal method that conditions on full-sentence context. |
| Outcome: | The proposed method shows that LM surprisal yields moderate correlations with scores/labels of metaphor novelty. |
A surprisal–duration trade-off across and within the world’s languages (2021.emnlp-main)
Copied to clipboard
| Challenge: | Throughout human evolution, countless languages have evolved, each with unique features. |
| Approach: | They analysed a corpus of 600 languages to find strong evidence for a surprisal–duration trade-off between languages and languages. |
| Outcome: | The proposed model shows that phones are produced faster in languages where they are less surprising and vice versa. |
Word Surprisal Correlates with Sentential Contradiction in LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing models are primarily optimized for task-specific performance, lacking well-defined objectives or linguistic grounding. |
| Approach: | They propose a token-to-word decoding algorithm that extends theoretically grounded probability estimation to open-vocabulary settings. |
| Outcome: | The proposed algorithm can localize sentence-level inconsistency at the word level, establishing a quantitative link between lexical uncertainty and sentential semantics. |
Finding Concept-specific Biases in Form–Meaning Associations (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to detect cross-linguistic associations are not effective, but their effects are minor. |
| Approach: | They propose a method to measure cross-linguistic associations by controlling for the influence of language family and geographic proximity within a large concept-aligned, cross-lingual lexicon. |
| Outcome: | The proposed method shows that it is small, but it is unsurprisingly small (less than 0.5% on average). |
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2024.acl-demos)
Copied to clipboard
| Challenge: | ACL 2024 System Demonstration Track invites submissions describing system demonstrations . submissions will undergo a single-blind review process . |
| Approach: | the ACL 2024 System Demonstration Track invites submissions . papers will be published in a companion volume of the conference proceedings . submissions will undergo a single-blind review process . |
| Outcome: | the Demonstration Track at ACL 2024 is a venue for papers describing system demonstrations . publicly available open-source or open-access systems are of special interest . submissions will undergo a single-blind review process . |
Dialect Clustering with Character-Based Metrics: in Search of the Boundary of Language and Dialect (2020.lrec-1)
Copied to clipboard
| Challenge: | 'A language is a dialect with an army and navy' is attributed to sociologist Max Weinrich. |
| Approach: | They propose a universal character-based method for representing sentences so that one can calculate the distance between any two sentence pairs. |
| Outcome: | The proposed method can be used to calculate distance between two sentences by clustering a dialect/sub-language mixed corpus into sub-groups and to partially answer the question of what separates languages from dialects. |
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2023.acl-demo)
Copied to clipboard
| Challenge: | 58 papers were selected for inclusion in the program, while a small number received only two reviews. |
| Approach: | the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023) will be held in london from July 9-14, 2023 . 58 submissions were selected for inclusion in the program, with an acceptance rate of 37%) |
| Outcome: | the system demonstration track received a record number of submissions . 58 papers were selected for inclusion in the program . |