Papers by Giovanni Semeraro

2 papers
Mind Your Special Tokens! On the Importance of Dedicated Sequence-End Tokens in Vision-Language Embedding Models (2026.eacl-short)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) are highly sensitive to end-of-input artifacts in fine-tuning and inference data, e.g., whether input sequences end with punctuation or newline characters.
Approach: They propose to convert generative LVLMs into vision-language encoders via contrastive learning objectives and use supervised contrastive objectives to train them.
Outcome: The proposed approach improves visual and text representations and improves retrieval and (semantic) similarity tasks.
XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE (2023.acl-short)

Copied to clipboard

Challenge: Existing approaches to the Word in Context task use cross-encoders, which prevent the possibility of deriving comparable word embeddings.
Approach: They propose a Lexical Semantic Change Detection model that extends SBERT, highlighting the target word in the sentence.
Outcome: The proposed model outperforms the state-of-the-art on the multilingual benchmarks for SemEval-2020 Task 1 - Lexical Semantic Change (LSC) Detection and the RuShiftEval shared task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations