Papers by Tzuf Paz-Argaman

8 papers
RUN through the Streets: A New Dataset and Baseline Models for Realistic Urban Navigation (D19-1)

Copied to clipboard

Challenge: Existing work on map-based NL navigation relies on small artificial worlds with a fixed set of entities known in advance.
Approach: They propose a task to interpret navigation instructions in natural language (NL) they use a dataset aligned with real, dense, urban maps to study neural architectures .
Outcome: The proposed task is based on a dataset of 2515 navigation instructions aligned with real routes over three regions of Manhattan.
HeGeL: A Novel Dataset for Geo-Location from Hebrew Text (2023.findings-acl)

Copied to clipboard

Challenge: Existing datasets in English for textual geolocation are limited because of the location of the place is implicit.
Approach: They propose to use a Hebrew place description corpus to analyze lingual geospatial reasoning.
Outcome: The Hebrew Geo-Location corpus collects literal Hebrew place descriptions and analyzes lingual geospatial reasoning.
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization (2025.acl-long)

Copied to clipboard

Challenge: n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear.
Approach: They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics.
Outcome: The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand.
Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs (2025.findings-naacl)

Copied to clipboard

Challenge: Current LLMs are primarily trained on English data but also include data from other languages.
Approach: They propose to use a pre-translation strategy to translate a task prompt into English before inference . they use 'a modular entity' that could be translated into four different languages .
Outcome: The proposed strategies are based on a set of pre-trained data across 35 languages covering both low and high-resource languages.
Where Do We Go From Here? Multi-scale Allocentric Relational Inferencefrom Natural Spatial Descriptions (2024.eacl-long)

Copied to clipboard

Challenge: Current NLP navigation studies focus on egocentric local descriptions that require reasoning over the agent’s local perception.
Approach: They propose to use a dataset to analyse English geospatial instructions to find locations and paths from natural language descriptions.
Outcome: The proposed task and dataset includes 10,404 examples of English geospatial instructions for reaching a target location using map-knowledge.
Into the Unknown: Generating Geospatial Descriptions for New Environments (2024.findings-acl)

Copied to clipboard

Challenge: Similar to vision-and-language navigation tasks, the Rendezvous (RVS) task requires reasoning over allocentric spatial relationships using non-sequential navigation instructions and maps.
Approach: They propose a large-scale augmentation method for generating high-quality synthetic data for new environments using readily available geospatial data.
Outcome: The proposed method improves accuracy on unseen and seen environments by 45.83% on the Rendezvous (RVS) task.
HeSum: a Novel Dataset for Abstractive Text Summarization in Hebrew (2024.findings-acl)

Copied to clipboard

Challenge: Large language models excel in various natural language tasks in English, but their performance in low-resource languages like Hebrew remains unclear.
Approach: They propose a benchmark dataset specifically designed for Hebrew abstractive text summarization that combines 10,000 article-summary pairs from Hebrew news websites.
Outcome: The proposed dataset shows that it presents distinct difficulties even for state-of-the-art LLMs.
ZEST: Zero-shot Learning from Text Descriptions using Textual Similarity and Visual Summarization (2020.findings-emnlp)

Copied to clipboard

Challenge: Specifically, given birds’ images with free-text descriptions of their species, we learn to classify images of previously-unseen species based on specie descriptions.
Approach: They propose to leverage the similarity between species and extract visual summaries from the texts to match visual features to the parts of the text that discuss them.
Outcome: The proposed model outperforms the state-of-the-art on the largest benchmarks for text-based zero-shot learning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations