Papers by Israfel Salazar

3 papers
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: N-gram-based evaluation metrics are unreliable due to low correlation to human judgments.
Approach: They propose a metric that rewards correct details and penalizes incorrect ones.
Outcome: The proposed metric matches the performance of open-source LLM-based metrics in correlation to human judgments while being far more efficient.
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for vision-language models treat compositionality and long-caption understanding in isolation.
Approach: They analyze when compositional reasoning and long-caption understanding transfer across tasks and when this relationship fails.
Outcome: The proposed model can generalize on poorly grounded captions and with strong visual grounding, while architectural choices can limit compositional learning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations