Learning to Write Rationally: How Information Is Distributed in Non-native Speakers’ Essays (2024.emnlp-main)
Copied to clipboard
| Challenge: | a study of second language learners with different native language backgrounds shows that people distribute information evenly in language production. |
| Approach: | They compare essays written by second language learners with different native language backgrounds to examine how they distribute information in non-native L2 production. |
| Outcome: | The authors found that writers with higher L2 proficiency can reduce uncertainty of language production while still conveying informative content. |
Similar Papers
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models. |
| Approach: | They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key . |
| Outcome: | The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key . |
Addressing Linguistic Bias through a Contrastive Analysis of Academic Writing in the NLP Domain (2023.emnlp-main)
Copied to clipboard
| Challenge: | a reviewer’s opinion of the nativeness of expression in an academic paper affects the likelihood of it being accepted for publication. |
| Approach: | They conduct a statistical analysis of paper abstracts from the natural language processing domain to identify how authors from different linguistic backgrounds differ in the lexical, morphological, syntactic and cohesive aspects of their writing. |
| Outcome: | The results suggest that there is potential for linguistic bias in the domain of natural language processing. |
Facilitating Cross-lingual Transfer of Empathy through Language-independent Latent Diffusion: A Case Study in Chinese (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing human empathy data are limited to English . a new study examines the pragmatic transferability of empathy across languages . |
| Approach: | a team of researchers integrate language-independent diffusion processes to facilitate the cross-lingual transfer of empathy. |
| Outcome: | The proposed method demonstrates that empathy can be transferred across languages without compromising linguistic naturalness. |
Revisiting Entropy Rate Constancy in Text (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evidence supports the uniform information density hypothesis . however, we re-evaluate the hypothesis with neural language models . |
| Approach: | They propose to use n-gram language models to argue that English documents exhibit entropy rate constancy . they re-evaluate the claims of Genzel and Charniak with neural language models . |
| Outcome: | The proposed hypothesis fails to support the proposed hypothesis with language models. |
Native Language Prediction from Gaze: a Reproducibility Study (2023.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have shown that the linguistic properties of a speaker’s native language affect the cognitive processing of other languages. |
| Approach: | They found that the correlation between eye movements and native language similarity may be more complex than the original study found. |
| Outcome: | The proposed model shows that the correlation between eye movements and native language similarity may be more complex than the original study. |
Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases (2025.acl-long)
Copied to clipboard
Rena Wei Gao, Xuetong Wu, Tatsuki Kuribayashi, Mingrui Ye, Siya Qi, Carsten Roever, Yuanxing Liu, Zheng Yuan, Jey Han Lau
| Challenge: | Large Language Models (LLMs) can simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (N1) knowledge. |
| Approach: | They use large language models to simulate non-native-like English use observed in human second language (L2) learners, and then compare their outputs to real L2 learner data. |
| Outcome: | The proposed models replicate L1-dependent patterns observed in human second language (L2) learners, with distinct influences from various languages. |
How Distributed are Distributed Representations? An Observation on the Locality of Syntactic Information in Verb Agreement Tasks (2022.acl-short)
Copied to clipboard
| Challenge: | Using probing, causal analysis and feature selection, we find that syntactic information is encoded locally in the transformers representations consistent with the French grammar. |
| Approach: | They address the question of the localization of syntactic information encoded in transformers representations by probing, causal analysis and feature selection methods. |
| Outcome: | The proposed representations are consistent with the object-past participle agreement in French and are consistent in both languages. |
Efficient Methods for Natural Language Processing: A Survey (2023.tacl-1)
Copied to clipboard
Marcos Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Qingqing Cao, Manuel R. Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, Pedro H. Martins, André F. T. Martins, Jessica Zosa Forde, Peter Milder, Edwin Simpson, Noam Slonim, Jesse Dodge, Emma Strubell, Niranjan Balasubramanian, Leon Derczynski, Iryna Gurevych, Roy Schwartz
| Challenge: | Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data, but using only scale to improve performance means resource consumption also grows. |
| Approach: | They propose to use data, time, storage, or energy to improve model performance. |
| Outcome: | The proposed methods and findings provide guidance for conducting NLP under limited resources and point towards promising research directions for developing more efficient methods. |
Beyond Facts- Benchmarking Distributional Reading Comprehension in Large Language Models (2026.findings-acl)
Copied to clipboard
Pei-Fu Guo, Ya An Tsai, Chun-Chia Hsu, Kai-Xin Chen, Yun-Da Tsai, Kai-Wei Chang, Nanyun Peng, Mi-Yen Yeh, Shou-De Lin
| Challenge: | Existing reading comprehension benchmarks focus on factual information, but many real-world tasks require distributional knowledge expressed across text. |
| Approach: | They propose a reading comprehension benchmark for LLMs to evaluate their ability to infer distributional knowledge from natural language. |
| Outcome: | Experiments with multiple LLMs show that the model outperforms baselines, but performance varies widely across distribution types and characteristics. |
How Much Knowledge Can You Pack Into the Parameters of a Language Model? (2020.emnlp-main)
Copied to clipboard
| Challenge: | In this paper, we show that large neural language models trained on unstructured text can attain competitive results on open-domain question answering benchmarks without access to external knowledge. |
| Approach: | They propose to fine-tune pre-trained neural language models to answer questions without external knowledge . they show that this approach scales with model size and performs competitively . |
| Outcome: | The proposed approach scales with model size and performs competitively with open-domain systems that explicitly retrieve answers from an external knowledge source when answering questions. |