Papers by Pierluigi Cassotti

9 papers
Analyzing Semantic Change through Lexical Replacements (2024.acl-long)

Copied to clipboard

Challenge: Modern language models can contextualize words based on their surrounding contexts, but semantic change can compromise this capability.
Approach: They propose a replacement schema where a target word is replaced with lexical replacements of varying relatedness . they leverage the replacement schema as a basis for a novel interpretable model for semantic change .
Outcome: The proposed model is the first to evaluate LLaMa for semantic change detection . it shows that lexical replacements can detect unexpected contexts .
Towards Language-Agnostic STIPA: Universal Phonetic Transcription to Support Language Documentation at Scale (2025.emnlp-main)

Copied to clipboard

Challenge: Existing ASR systems focus on orthographic output for high-resource languages, but STIPA can be used as a language-agnostic interface for documenting under-resourced and unwritten languages.
Approach: They propose to use the International Phonetic Alphabet (STIPA) to generate phonetic transcriptions using a language-agnostic interface.
Outcome: The proposed model reduces phonetic error rates even in low-resource settings and can be used for documenting under-resourced and unwritten languages.
Computational modeling of semantic change (2024.eacl-tutorials)

Copied to clipboard

Challenge: Languages change constantly over time, influenced by social, technological, cultural and political factors that affect how people express themselves.
Approach: They propose to categorise the types of change, the causes and the mechanisms underlying the different types of changes using large diachronic corpora and evaluation benchmarks.
Outcome: In historical linguistics, tools and methods have been developed to analyse the process . they include categorisations of types of change, causes and mechanisms . but traditional methods, while informative, are often based on small, carefully curated samples.
More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Word Usage Graphs (WUGs) represent word sense clusters from simple pairwise word use judgments.
Approach: They propose to use a weighted graph to represent human semantic proximity judgments for pairs of word uses to infer word sense clusters from simple pairwise word use judgments.
Outcome: The proposed approach can be applied in a Word Sense Induction (WSI) setting or for Word sense disambiguation (WSD) it is the first and to date largest manually annotated, diachronic WUG dataset.
TRoTR: A Framework for Evaluating the Re-contextualization of Text Reuse (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for detecting text reuse focus on recontextualization . current approaches focus on text reuse across a diachronic corpus .
Approach: They propose a framework that relies on topic relatedness for evaluating the diachronic change of context in which text is reused.
Outcome: The proposed framework evaluates biblical text reuse human-annotated with topic relatedness . it exhibits greater sensitivity to textual similarity than topic relatedity, the authors show .
Using Synchronic Definitions and Semantic Relations to Classify Semantic Change Types (2024.acl-long)

Copied to clipboard

Challenge: Existing models for detecting semantic change in corpora have been disregarded due to lack of knowledge of the nature of semantic change and the way it takes place.
Approach: They propose a model that leverages synchronic lexical relations and definitions of word meanings to detect these types of change.
Outcome: The proposed model can detect changes in a digitized version of Blank's dataset and improve human judgments of semantic relatedness and binary Lexical Semantic Change Detection.
SenseRel: A Sense-Level Benchmark for Denotational and Connotational Meaning Relations (2026.acl-long)

Copied to clipboard

Challenge: Polysemy enables a single word to convey multiple related meanings . a word's sense is extended to new contexts and concepts, a process called semantic change is gradual .
Approach: They propose a benchmark for modeling semantic relations between word senses . they use a model that distinguishes denotational and connotationally related aspects of meaning .
Outcome: The proposed model is able to distinguish between denotational and connotationalist aspects of meaning . it is compared with models with GPT-4o, Llama 3.1, and DeepSeek .
XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE (2023.acl-short)

Copied to clipboard

Challenge: Existing approaches to the Word in Context task use cross-encoders, which prevent the possibility of deriving comparable word embeddings.
Approach: They propose a Lexical Semantic Change Detection model that extends SBERT, highlighting the target word in the sentence.
Outcome: The proposed model outperforms the state-of-the-art on the multilingual benchmarks for SemEval-2020 Task 1 - Lexical Semantic Change (LSC) Detection and the RuShiftEval shared task.
Elections go bananas: A First Large-scale Multilingual Study of Pluralia Tantum using LLMs (2026.eacl-long)

Copied to clipboard

Challenge: a large amount of annotated sentences for each feature can be used for in-depth analysis.
Approach: They propose an annotation framework for lexicalization of pluralia tantum . they use an LLM to annotate each instance from the reference corpus .
Outcome: The proposed framework provides useful annotators for semantic, syntactic and sense categories with accuracy ranging from 51% to 89% on a hand-annotated testset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations