Papers by Itai Mondshine
HeGeL: A Novel Dataset for Geo-Location from Hebrew Text (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets in English for textual geolocation are limited because of the location of the place is implicit. |
| Approach: | They propose to use a Hebrew place description corpus to analyze lingual geospatial reasoning. |
| Outcome: | The Hebrew Geo-Location corpus collects literal Hebrew place descriptions and analyzes lingual geospatial reasoning. |
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization (2025.acl-long)
Copied to clipboard
| Challenge: | n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear. |
| Approach: | They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics. |
| Outcome: | The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand. |
Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs (2025.findings-naacl)
Copied to clipboard
| Challenge: | Current LLMs are primarily trained on English data but also include data from other languages. |
| Approach: | They propose to use a pre-translation strategy to translate a task prompt into English before inference . they use 'a modular entity' that could be translated into four different languages . |
| Outcome: | The proposed strategies are based on a set of pre-trained data across 35 languages covering both low and high-resource languages. |
HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models excel in various natural language tasks in English, but their performance in low-resource languages like Hebrew remains unclear. |
| Approach: | They propose a benchmark dataset specifically designed for Hebrew abstractive text summarization that combines 10,000 article-summary pairs from Hebrew news websites. |
| Outcome: | The proposed dataset shows that it presents distinct difficulties even for state-of-the-art LLMs. |