Targeted Syntactic Evaluation on the Chomsky Hierarchy (2024.lrec-main)

Copied to clipboard

Challenge: a novel evaluation paradigm for targeted syntactic evaluations is proposed . we create formal languages that abstract four syntaktic phenomena in natural languages .
Approach: They propose a new evaluation paradigm for Targeted Syntactic Evaluations . they create formal languages that abstract syntactical phenomena in natural languages .
Outcome: The proposed evaluation paradigm evaluates language models on language modeling tasks . it shows that they can capture the structural patterns of the (Adj)n NP type formal language .

Similar Papers

Targeted Syntactic Evaluation of Language Models (D18-1)

Copied to clipboard

Challenge: Recent advances have led to an explosion of neural network-based LM architectures.
Approach: They propose to supplement perplexity with a metric that assesses whether a language model can predict the grammatical sentence more accurately than an ungrammatically-based model.
Outcome: The proposed model performed poorly on many of the constructions.
SyntaxGym: An Online Platform for Targeted Evaluation of Language Models (2020.acl-demos)

Copied to clipboard

Challenge: SyntaxGym is an online platform and open-source framework for targeted syntactic evaluation of neural network language models.
Approach: They propose to make targeted syntactic evaluations accessible to both experts in NLP and linguistics and reproducible across computing environments.
Outcome: The proposed framework is reproducible across computing environments and standardized following the norms of psycholinguistic experimental design.
Overestimation of Syntactic Representation in Neural Language Models (2020.acl-main)

Copied to clipboard

Challenge: Several testing methodologies have been developed to probe models’ syntactic representations.
Approach: They propose a method to determine syntactic structure by training a model on strings generated according to a template and testing its ability to distinguish between similar ones with different syntax.
Outcome: The proposed method reproduces positive results with two non-syntactic baseline language models: an n-gram model and an LSTM model trained on scrambled inputs.
On Efficiently Representing Regular Languages as RNNs (2024.findings-acl)

Copied to clipboard

Challenge: Recent work by Hewitt et al. (2020) provides an interpretation of the empirical success of recurrent neural networks (RNNs) as language models (LMs).
Approach: They generalize their construction and show that RNNs can efficiently represent a larger class of LMs than previously claimed.
Outcome: The results suggest that RNNs can represent a larger class of LMs than previously claimed .
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges (2024.emnlp-main)

Copied to clipboard

Challenge: introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance.
Approach: They propose a taxonomy for organizing existing LLM-based evaluation metrics and a structured framework to understand and compare them.
Outcome: The proposed taxonomy offers a framework to understand and compare LLM-based evaluation methods.
Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have not evaluated whether probing accuracy predicts syntactic outcomes.
Approach: They evaluate 32 open-weight transformer models and find that probing fails to predict outcomes of targeted syntax evaluations across English linguistic phenomena.
Outcome: The proposed model does not predict syntactic outcomes on English linguistic phenomena.
Do Neural Language Models Show Preferences for Syntactic Formalisms? (2020.acl-main)

Copied to clipboard

Challenge: Recent work on interpretability of deep neural language models concludes that many properties of natural language syntax are encoded in their representational spaces.
Approach: They propose to examine whether syntactic structure adheres to a surface-syntactical or deep syntaktic style of analysis.
Outcome: The proposed model prefers Universal Dependencies (UD) over Surface-Syntactic Universal Dependency (SUD) with interesting variations across languages and layers.
Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in language models have demonstrated their ability to process linguistic expressions with complex syntactic structures.
Approach: They investigate whether language models employ shared neural mechanisms across different constructions by applying causal interpretability methods at a granular level.
Outcome: The proposed model performance improves on acceptability judgment benchmarks.
Discovering Language Model Behaviors with Model-Written Evaluations (2023.findings-acl)

Copied to clipboard

Challenge: Prior work creates evaluations with crowdwork or existing data sources, which are not always available.
Approach: They generate evaluations automatically with language models (LMs) using crowdwork or existing data sources to find out how they behave .
Outcome: The results show that large LMs repeat back a dialog user’s preferred answer and express greater desire to pursue concerning goals like resource acquisition and goal preservation.
Probing LLMs for Joint Encoding of Linguistic Categories (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing research suggests that a linguistic hierarchy emerges in large language models . little is known about how encodings of different linguistic phenomena interact within the models - and to what extent processing of linguistically-related categories relies on the same, shared model representations.
Approach: They propose a framework for testing the joint encoding of linguistic categories in large language models.
Outcome: The proposed framework shows that the same patterns hold across languages in multilingual LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations