Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics (2024.lrec-main)
Copied to clipboard
| Challenge: | Scalar adjectives describe different domain scales and vary in intensity . they can be triggered by scalar adjective and require listeners to reason pragmatically about them. |
| Approach: | They probe different families of Large Language Models for their knowledge of the lexical semantics of scalar adjectives and one specific aspect of their pragmatics. |
| Outcome: | The proposed models encode rich lexical-semantic information about scalar adjectives but lack a good understanding of skalar diversity. |
Similar Papers
Pragmatic inference of scalar implicature by LLMs (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing Large Language Models (LLMs) engage in pragmatic inference of scalar implicature, such as some. |
| Approach: | They investigate how Large Language Models (LLMs) engage in pragmatic inference of scalar implicature, such as some. |
| Outcome: | The proposed models interpret some as pragmatic implicature not all in the absence of context, aligning with human language processing. |
Scalar Adjective Identification and Multilingual Ranking (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing studies on scalar adjective ranking have focused on English due to the availability of datasets for evaluation. |
| Approach: | They propose a binary classification task to examine the models’ ability to distinguish scalar from relational adjectives in English. |
| Outcome: | The proposed task compares the models' ability to distinguish scalar from relational adjectives in English using monolingual and multilingual models. |
Continuous Interpretive Steering for Scalar Diversity (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies on pragmatic inference in large language models rely on prompt-based manipulations to elicit a pragmatic interpretation. |
| Approach: | They propose a method that probes graded pragmatic interpretation by treating activation-level steering strength as a continuous experimental variable. |
| Outcome: | The proposed method increases pragmatic interpretations globally but collapses item-level variation whereas graded activation steering yields differentiated interpretive shifts aligned with scalar diversity grades. |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |
Annotating the French Wiktionary with supersenses for large scale lexical analysis: a use case to assess form-meaning relationships within the nominal lexicon (2025.coling-main)
Copied to clipboard
| Challenge: | Conducting large-scale empirical studies in lexical semantics remains an elusive goal for many languages lacking comprehensive semantic resources. |
| Approach: | They propose to use the Princeton WordNet to enrich the French Wiktionary with general semantic classes, known as supersenses, using a limited amount of manually annotated data. |
| Outcome: | The proposed method can be extended to other languages provided an electronic lexicon and manually annotated senses are available. |
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing LLM leaderboards often reference scores reported in other papers without consistent settings and prompts, which may encourage cherry-picking favored settings and for better results. |
| Approach: | They propose an open-source and reproducible LLM evaluation suite built on top of OpenAI Evals that systematically evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories. |
| Outcome: | The evaluation suite is built on top of OpenAI Evals and evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories. |
BERT Knows Punta Cana is not just beautiful, it’s gorgeous: Ranking Scalar Adjectives with Contextualised Representations (2020.emnlp-main)
Copied to clipboard
| Challenge: | Adjectives describe positive properties of nouns but with different intensity. |
| Approach: | They propose a BERT-based approach to intensity detection for scalar adjectives by generating vectors directly from contextualised representations. |
| Outcome: | The proposed model outperforms static embeddings and previous models with dedicated resources on an Indirect Question Answering task. |
Harnessing the linguistic signal to predict scalar inferences (2020.acl-main)
Copied to clipboard
| Challenge: | Recent Bayesian game-theoretic models of pragmatic reasoning can predict the strength of scalar inferences by using linguistic features. |
| Approach: | They propose to use a sentence encoder to predict the strength of scalar inferences by using a corpus of linguistic data. |
| Outcome: | The proposed model infers previously established associations between linguistic features and inference strength, suggesting that it learns to use linguistic feature to predict pragmatic inferences. |
Quantifying Generalizations: Exploring the Divide Between Human and LLMs’ Sensitivity to Quantification (2024.acl-long)
Copied to clipboard
| Challenge: | Generics are expressions used to communicate abstractions about categories . they allow for exceptions, and they are a powerful way to express knowledge about the world . |
| Approach: | They examine how large language models interpret generics to understand their meanings . they find that the presence of a generic sentence as context influences quantifiers based on the generalization . |
| Outcome: | The proposed models do not exhibit a strong sensitivity to quantification, the study finds . the results suggest that the presence of a generic sentence as context influences quantifiers . |
How Abstract Is Linguistic Generalization in Large Language Models? Experiments with Argument Structure (2023.tacl-1)
Copied to clipboard
| Challenge: | Competent speakers of a language know how likely a word w is to appear in a specific context . |
| Approach: | They use transformer-based large language models to generalize a novel noun argument . they show a bias to generalise based on linear order, instead of a linear order . |
| Outcome: | The proposed models perform well in generalizing the distribution of a novel noun argument between related contexts that were seen during pre-training. |