Papers by Anton Lavrouk

2 papers
What are Foundation Models Cooking in the Post-Soviet World? (2025.emnlp-main)

Copied to clipboard

Challenge: During the Soviet era, these identities were pressured through forced assimilation under the Russian language and culture.
Approach: They construct a multi-modal dataset encompassing 1147 and 823 dishes in the Russian and Ukrainian languages, centered around the Post-Soviet region.
Outcome: The results show that leading models struggle to correctly identify the origins of dishes from Post-Soviet nations in both text-only and multi-modal Question Answering (QA) the weak correlation between this task and QA suggests that QA alone may be insufficient as an evaluation of cultural understanding.
ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment (2024.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses.
Approach: They propose to use a multilingual multi-domain dataset to benchmark multilingual and monolingual models for multilingual readability assessment.
Outcome: The proposed model trains better in supervised, unsupervised, and few-shot prompting settings and identifies shortcomings in state-of-the-art unsupervised methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations