Papers by Panagiotis Kaliosis

4 papers
Learning to Align: Addressing Character Frequency Distribution Shifts in Handwritten Text Recognition (2025.findings-emnlp)

Copied to clipboard

Challenge: Character sets change over time and character frequency distributions shift across historical periods or regions . character distribution alignment can improve existing models at inference time without requiring retraining .
Approach: They propose a loss function that incorporates the Wasserstein distance between predicted and target distributions.
Outcome: The proposed method improves accuracy and robustness under temporal and contextual shifts.
A Data-Driven Guided Decoding Mechanism for Diagnostic Captioning (2024.findings-acl)

Copied to clipboard

Challenge: Diagnostic Captioning (DC) systems receive one or more medical images of a patient, such as X-Rays or Magnetic Resonance Images (MRIs).
Approach: They propose a data-driven guided decoding method that incorporates medical information into the beam search of the diagnostic text generation process.
Outcome: The proposed method improves on two medical datasets and can be used in few- and zero-shot learning scenarios.
LVLMs and Humans Ground Differently in Referential Communication (2026.acl-long)

Copied to clipboard

Challenge: generative AI agents cannot model common ground in a way that enables smooth communication . a recent study examined whether large language models and large vision language models engage in grounding as human discourse partners do .
Approach: They propose to use referential communication to model common ground between a pair of directors and a picture matching system.
Outcome: The proposed experiment shows that generative AI agents cannot model common ground . human conversation relies on common ground accrued and updated by interacting partners .
LVLMs are Bad at Overhearing Human Referential Communication (2025.emnlp-main)

Copied to clipboard

Challenge: a crucial skill for embodied AI agents working with humans is grounding in referential communication.
Approach: They use large vision language models to overhear spontaneous conversations between humans . they find that current LVLMs fail to show consistent performance improvement .
Outcome: The proposed models fail to show consistent performance improvement over previous models . the authors release the results to facilitate future research .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations