Do LLMs learn a true syntactic universal? (2024.emnlp-main)

Copied to clipboard

Challenge: linguistics literature has debated whether large multilingual language models learn language universals . Typological generalizations are a key battleground in such debates - e.g. van der Hulst, 2023, chapter 7).
Approach: They consider a candidate universal for language universals, the Final-over-Final Condition . they suggest that modern language models may need additional sources of bias to become truly human-like .
Outcome: The proposed model only seems to recognize the Final-over-Final Condition in German, Russian, Hungarian and Serbian .

Similar Papers

Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on Understanding (2023.emnlp-main)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text.
Approach: They argue that LLMs only parrot statistical patterns in training data and that language learning in LLM cannot inform human language learning.
Outcome: The proposed model can generate grammatically correct, fluent text without requiring human intervention.
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment.
Approach: They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge.
Outcome: The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses.
Can Language Models Learn Typologically Implausible Languages? (2026.tacl-1)

Copied to clipboard

Challenge: Language models provide a naturalistic framework for studying artificial language learning . authors: typological universals and tendencies are thought to be caused by a learning bias .
Approach: They propose to train LMs on highly naturalistic counterfactual versions of English and Japanese . they show that LM learn subtly implausible languages more slowly .
Outcome: The proposed language models learn subtly implausible languages more slowly compared to human models . the findings suggest that LMs exhibit typologically aligned learning preferences .
Language Models Learn Universal Representations of Numbers and Here’s Why You Should Care (2026.acl-long)

Copied to clipboard

Challenge: Prior work has shown that large language models (LLMs) often converge to accurate input embedding for numbers, based on sinusoidal representations.
Approach: They show that large language models often converge to accurate input embedding for numbers, based on sinusoidal representations.
Outcome: The proposed representations are strikingly systematic, and are interchangeable in a large swathe of experimental setups.
On the Continued Value of Universal Dependencies in the Era of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: a growing belief that explicit linguistic representations are no longer necessary is questioned in large language models . a recent study examines whether and in what ways this cross-lingual syntactic framework can still benefit LLMs .
Approach: They use Universal Dependencies (UD) to examine whether and in what ways it can still benefit LLMs.
Outcome: The proposed model outperforms its syntax-agnostic counterparts in a cross-lingual evaluation task.
Towards Practical and Knowledgeable LLMs for a Multilingual World: A Thesis Proposal (2025.naacl-srw)

Copied to clipboard

Challenge: a proposed thesis examines the role that multilinguality occupies in the development of practical and knowledgeable LLMs.
Approach: They propose to use multilingual knowledge to improve LLM performance on NLP tasks . they extend the territorial disputes benchmark to retrieval-augmented generation setting .
Outcome: The proposed methods improve LLMs' performance on standard natural language processing tasks by leveraging their existing multilingual knowledge.
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs (2025.acl-long)

Copied to clipboard

Challenge: LMs are highly flexible learners, capable of acquiring linguistic patterns beyond those learnable by humans.
Approach: They train LMs to model impossible and typologically unattested languages . they find that the model does not achieve perfect separation between attested and unattest languages - suggesting some human-like inductive biases .
Outcome: The proposed model can largely distinguish attested from impossible languages, but does not achieve perfect separation between them and their impossible counterparts.
Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs (2025.acl-long)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are predominantly designed with English as the primary language, but many are still English-dominated.
Approach: They propose to use automatic corpus-level metrics to assess lexical and syntactic naturalness of LLMs in a multilingual context.
Outcome: The proposed method improves naturalness of LLMs in target languages without compromising performance on general-purpose benchmarks.
Domain Regeneration: How well do LLMs match syntactic properties of text domains? (2025.findings-acl)

Copied to clipboard

Challenge: Recent improvements in large language models have improved their ability to approximate distributions . authors find that LLMs can suffer from model collapse due to domain considerations based on pretraining .
Approach: They use open source LLMs to regenerate permissively licensed English text from Wikipedia and news text.
Outcome: The proposed model can faithfully match the human-generated distributions in a semantically-controlled setting.
Biasless Language Models Learn Unnaturally: How LLMs Fail to Distinguish the Possible from the Impossible (2026.eacl-long)

Copied to clipboard

Challenge: linguists have discovered patterns which hold across virtually all known natural languages . lingulists are able to learn languages by comparing their learning curves to those of humans .
Approach: They compare LLM learning curves on existing and "impossible" datasets . they find that GPT-2 learns each language and its impossible counterpart equally easily .
Outcome: The proposed model learns each language and its impossible counterpart equally easily, the study shows . the study also shows that the proposed model does not provide any kind of separation between the possible and the impossible .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations