Papers by Omri Abend

43 papers
Comprehensive Supersense Disambiguation of English Prepositions and Possessives (P18-1)

Copied to clipboard

Challenge: Frequent prepositions like for are maddeningly polysemous, their interpretation depends especially on the object of the preposition.
Approach: They propose a new annotation scheme, corpus, and task for the disambiguation of prepositions and possessives in English.
Outcome: The proposed annotations are comprehensive with respect to types and tokens of these markers and use broadly applicable supersense classes rather than fine-grained dictionary definitions.
Simple and Effective Text Simplification Using Semantic and Neural Methods (P18-1)

Copied to clipboard

Challenge: Sentence splitting is a major simplification operation.
Approach: They propose a simple and efficient splitting algorithm based on an automatic semantic parser.
Outcome: The proposed method compares favorably to the state-of-the-art in combined lexical and structural simplification.
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for Theory of Mind (ToM) focus on whether agents have correct beliefs about others.
Approach: They propose to evaluate Theory of Mind (ToM) capabilities in Large Language Models (LLMs) they propose to use the theory of mind to determine whether and how to invoke ToM .
Outcome: The proposed frameworks can be used to evaluate the performance of large language models (LLMs) in biological agents.
DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Question answering models have access to two sources of knowledge during inference time: parametric knowledge and contextual knowledge.
Approach: They propose a new paradigm in which QA models are trained to disentangle the two sources of knowledge.
Outcome: The proposed model generates two answers for a given question based on parametric and contextual knowledge.
Inherent Biases in Reference-based Evaluation for Grammatical Error Correction (P18-1)

Copied to clipboard

Challenge: Existing evaluation systems obtain comparable or superior performance compared to humans by making few but targeted changes to the input.
Approach: They propose to re-scale M 2 by the inter-annotator agreement and increase the number of references in any feasible range to overcome low coverage bias in GEC evaluation.
Outcome: The proposed measure overcomes low coverage bias in GEC evaluation by re-scaling or increasing the number of references in any feasible range.
Exploring the Learning Capabilities of Language Models using LEVERWORLDS (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models of stochastic learning involve learning general structure rules and specific properties of the instance.
Approach: They propose a framework that allows the generation of physics-inspired worlds that follow a similar generative process with different distributions and their instances can be expressed in natural language.
Outcome: The proposed framework allows the generation of physics-inspired worlds that follow a similar generative process with different distributions and their instances can be expressed in natural language.
Fine-Grained Analysis of Cross-Linguistic Syntactic Divergences (2020.acl-main)

Copied to clipboard

Challenge: Existing work on quantifying the prevalence of syntactic divergences across languages has not been done.
Approach: They propose a framework for extracting divergence patterns for any language pair from a parallel corpus building on Universal Dependencies.
Outcome: The proposed framework provides a detailed picture of cross-language divergences, generalizes previous approaches, and lends itself to full automation.
Reference-less Measure of Faithfulness for Grammatical Error Correction (N18-2)

Copied to clipboard

Challenge: Existing reference-less measures (RLMs) for measuring grammaticality are expensive to collect and limited by the large number of valid outputs.
Approach: They propose a semantic measure for Grammatical Error Correction that compares the semantic symbolic structure of the source and correction without relying on manually-curated references.
Outcome: The proposed measure shows that it can be applied consistently to ungrammatical text, and that valid corrections obtain a high USim similarity score to the source, and invalid corrections get lower scores.
Cross-lingual Semantic Representation for NLP with UCCA (2020.coling-tutorials)

Copied to clipboard

Challenge: introductory tutorial to UCCA, a symbolic meaning representation for semantic representations.
Approach: This tutorial introduces UCCA, a cross-linguistically applicable framework for semantic representation . it will provide a detailed introduction to the UCca annotation guidelines, design philosophy and available resources .
Outcome: The tutorial will provide a detailed introduction to the UCCA framework and compare it to other meaning representations.
The Language of Legal and Illegal Activity on the Darknet (P19-1)

Copied to clipboard

Challenge: a study of the characteristics of text in the Darknet shows that it has legal and illegal content.
Approach: They compare texts for selling legal and illegal drugs to a control condition . they find several distinguishing features between legal and illicit texts .
Outcome: The authors compare legal and illegal texts to a clear net website with similar content as a control condition.
The Grammar-Learning Trajectories of Neural Language Models (2022.acl-long)

Copied to clipboard

Challenge: In this paper, we show that neural language models with different initialization, architecture, and training data acquire linguistic phenomena in a similar order, despite their different end performance.
Approach: They propose to use mutual inductive bias to study linguistic representations implicit in NLMs.
Outcome: The proposed approach shows that NLMs with different initialization, architecture, and training data acquire linguistic phenomena in a similar order, despite their different end performance.
CONTESTS: a Framework for Consistency Testing of Span Probabilities in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Language model scores are often treated as probabilities, but their reliability as probability estimators has mainly been studied through calibration, overlooking other aspects.
Approach: They propose a framework to assess model reliability across interchangeable completion and conditioning orders by performing statistical tests on real and synthetic data to eliminate training effects.
Outcome: The proposed framework assesses the consistency of model predictions across interchangeable completion and conditioning orders on real and synthetic data to eliminate training effects.
Q2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation methods for factual consistency in knowledge-grounded dialogues are unreliable and limit their applicability.
Approach: They propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering.
Outcome: The proposed evaluation metric consistently shows higher correlation with human judgements.
Mediators in Determining what Processing BERT Performs First (2021.naacl-main)

Copied to clipboard

Challenge: Probing neural models for the ability to perform downstream tasks using their activation patterns is often used to localize what parts of the network specialize in performing which tasks.
Approach: They propose to consider the prediction’s context length as a potential mediating factor and consider the length of the span whose processing is minimally required to perform the prediction.
Outcome: The proposed model can get 196 different rankings when probing with seven tasks, the authors show .
Semantics-aware Attention Improves Neural Machine Translation (2022.starsem-1)

Copied to clipboard

Challenge: Existing attempts to integrate semantic structures into NMT Transformers have failed .
Approach: They propose two parameter-free methods for injecting semantic information into Transformers, using a Scene-Aware Self-Attention (SASA) head and a Scenario-Award Cross-Action (SACrA) head.
Outcome: The proposed methods improve on the vanilla Transformer and syntax-aware models for four language pairs and show an additional gain when using both semantic and syntactic structures in some language pairs.
T5Score: A Methodology for Automatically Assessing the Quality of LLM Generated Multi-Document Topic Sets (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods for Multi-Document Topic Extraction are not designed for LLMs and result in low inter-annotator agreement scores.
Approach: They propose an evaluation methodology that decomposes the quality of a topic set into quantifiable aspects, measurable through easy-to-perform annotation tasks.
Outcome: The proposed evaluation methodology decomposes the quality of a topic set into quantifiable aspects, measurable through easy-to-perform annotation tasks.
Multitask Parsing Across Semantic Representations (P18-1)

Copied to clipboard

Challenge: UCCA parsing is a test case for multitask learning, with auxiliary tasks AMR, SDP and Universal Dependencies (UD) . Semantic parsers have arguably yet to reach their full potential due to the limited amount of semantically annotated training data.
Approach: They propose a general transition-based parser that can parse UCCA, AMR, SDP and Universal Dependencies (UD) they use a transition-driven learning architecture and a uniform transition-basic learning architecture to train the parsers.
Outcome: The proposed parser improves UCCA, AMR, SDP and Universal Dependencies (UD) parsing over training in English, German and French.
Computational Analysis of Character Development in Holocaust Testimonies (2025.emnlp-main)

Copied to clipboard

Challenge: This work examines character development along the narrative timeline by analyzing changes in the protagonist’s views and behavior and the interplay between them.
Approach: They propose to analyze character development along the narrative timeline using a transcript of Holocaust survivor testimonies as a test case.
Outcome: The proposed approach characterizes changes in the protagonist’s views and behavior and the interplay between them.
Comparison by Conversion: Reverse-Engineering UCCA from Syntax and Lexical Semantics (2020.coling-main)

Copied to clipboard

Challenge: a systematic comparative analysis of linguistic meaning representations from different frameworks is needed.
Approach: They compare a rule-based converter and a supervised delexicalized parser to map meaning representations from different frameworks.
Outcome: The proposed method yields surprisingly accurate representations close to fully supervised UCCA parser quality.
Semantic Structural Evaluation for Text Simplification (N18-1)

Copied to clipboard

Challenge: Current measures for evaluating text simplification systems focus on lexical aspects, neglecting its structural aspects.
Approach: They propose to use a reference-less automatic evaluation procedure to assess simplification quality by decomposing the input based on its semantic structure and comparing it to the output.
Outcome: The proposed measure has a significant correlation with human judgments and is highly comparable with existing measures.
Improving Cross-lingual Transfer through Subtree-aware Word Reordering (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that multilingual language models are not effective when dealing with less-represented languages.
Approach: They propose a powerful reordering method that learns word-order patterns conditioned on the syntactic context from a small amount of annotated data.
Outcome: The proposed method outperforms baselines on a variety of tasks and is effective in both zero-shot and few-shot scenarios.
Machine Reading of Historical Events (2020.acl-main)

Copied to clipboard

Challenge: Using a short text description of an event, we can extract relevant sentences from Wikipedia and apply a combination of task-specific and general-purpose feature embeddings for the classification.
Approach: They propose to use Wikipedia sentences to extract relevant sentences and apply feature embeddings to the task.
Outcome: The proposed model outperforms the historical event ordering task and the event focus time task in the literature.
Mediocrity is the key for LLM as a Judge Anchor Selection (2026.acl-long)

Copied to clipboard

Challenge: a poor selection of an anchor can dramatically reduce correlation with human rankings . traditional reference-based metrics are often ill-suited for open-ended generation .
Approach: They evaluate 22 different anchors on a Arena-Hard-v2.0 dataset and quantify the effect size of anchor selection.
Outcome: The proposed model is better or worse than all other models, but it is rarely indicative of the relative ranking of the models.
Locally Measuring Cross-lingual Lexical Alignment: A Domain and Word Level Perspective (2024.findings-emnlp)

Copied to clipboard

Challenge: a cognitive science research focus on aligning language spaces in their entirety . but, cognitive science has long focused on a local perspective . a new method for cross-lingual lexical alignment requires some methodology .
Approach: They propose a method for analyzing kinship domain kinematics and a new method for contextualization . they propose kin-level validations and contextualizations to validate the results .
Outcome: The proposed method analyzes synthetic validations and naturalistic validations using lexical gaps in the kinship domain.
Generating Benchmarks for Factuality Evaluation of Language Models (2024.eacl-long)

Copied to clipboard

Challenge: Existing methods for factuality evaluation of LLM generation focus on facts sampled from the LM itself and might under-represent domain specific or rare facts.
Approach: They propose a method that transforms a factual corpus into a benchmark evaluating an LM's propensity to generate true facts from the corpus .
Outcome: The proposed framework transforms a factual corpus of interest into a benchmark evaluating an LM's propensity to generate true facts from the corpus vs. similar but incorrect statements.
Putting Words in BERT’s Mouth: Navigating Contextualized Vector Spaces with Pseudowords (2021.emnlp-main)

Copied to clipboard

Challenge: a new technique for exploring contextualized vector space is proposed . masked prediction of a word in a sentence allows controlled exploration of the space .
Approach: They propose a method for exploring regions around individual points in a contextualized vector space . they use a static embedding to induce a "pseudoword" vector and masked prediction of a word .
Outcome: The proposed method investigates the geometry of the contextualized space around individual instances of a word . it uses a static embedding to induce a contextualized "pseudoword" vector .
A Computational Acquisition Model for Multimodal Word Categorization (2022.naacl-main)

Copied to clipboard

Challenge: Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition.
Approach: They propose a multimodal language acquisition model trained from image-caption pairs on naturalistic data using cross-modal self-supervision.
Outcome: The proposed model learns word categories and object recognition abilities, the authors show . their model is trained from image-caption pairs on naturalistic data using cross-modal self-supervision .
Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney (2023.emnlp-main)

Copied to clipboard

Challenge: Generating images with Text-to-Image models often requires multiple trials, where human users iteratively update their prompt based on feedback, namely the output image.
Approach: They compile a dataset of iterative interactions of human users with Midjourney and analyze the dynamics of the user prompts along these iterations.
Outcome: The proposed model produces better images for a specific language style than other models.
A Large-Scale Multilingual Study of Visual Constraints on Linguistic Selection of Descriptions (2023.findings-eacl)

Copied to clipboard

Challenge: a multilingual study examines how vision constrains linguistic choice . we use existing annotations to investigate the effect of different visual conditions on numeral expressions in captions .
Approach: They propose a method that leverages existing corpora of images with captions written by native speakers to constrain linguistic choice.
Outcome: The proposed method covers four languages and five linguistic properties, including verb transitivity and use of numerals.
BLEU is Not Suitable for the Evaluation of Text Simplification (D18-1)

Copied to clipboard

Challenge: BLEU is widely considered to be an informative metric for text-to-text generation . Xu et al. (2016) found that BLUE is not suitable for evaluation of sentence splitting .
Approach: They propose to use BLEU to evaluate sentence splitting as a metric for machine translation . they propose to compare BLUE with a corpus containing multiple structural paraphrases .
Outcome: The proposed BLEU is not suitable for evaluation of sentence splitting . a correlation analysis with human judgments shows low correlation with BLUE .
Reinforcement Learning with Large Action Spaces for Neural Machine Translation (2022.coling-1)

Copied to clipboard

Challenge: Recent work has argued that the gains produced by Reinforcement learning are mostly due to promoting tokens that have already received a fairly high probability in pre-training.
Approach: They hypothesize that the large action space is a main obstacle to RL’s effectiveness in MT by reducing the size of the vocabulary without changing the vocabulary.
Outcome: The proposed method improves by 1.5 BLEU points on average.
Automatic Metric Validation for Grammatical Error Correction (P18-1)

Copied to clipboard

Challenge: Existing methods for metric validation in GEC suffer from low inter-rater agreement.
Approach: They propose an automatic method for GEC metric validation that overcomes many of the difficulties in the existing method.
Outcome: The proposed method sheds new light on metric quality and shows valid edits are penalized by existing metrics.
Event-Location Tracking in Narratives: A Case Study on Holocaust Testimonies (2023.emnlp-main)

Copied to clipboard

Challenge: a primary goal of narrative analysis is to represent essential dimensions of stories in a schematic manner.
Approach: They propose a task to extract the sequence of locations where the narrative is set through its progression.
Outcome: The proposed task is based on the test case of Holocaust survivor testimonies . it shows that models that are aware of the larger context can generate more accurate locations chains.
On the Relation between Syntactic Divergence and Zero-Shot Performance (2021.emnlp-main)

Copied to clipboard

Challenge: Recent advances in cross-lingual transfer methods have enabled significant advances in grammatical processing tasks.
Approach: They examine the extent to which syntactic relations are preserved in translation and parsability in a zero-shot setting.
Outcome: The proposed model is based on a translation task in English and a subset of a standard English RE benchmark translated to Russian and Korean.
Evaluating and Improving the Coreference Capabilities of Machine Translation Models (2023.eacl-main)

Copied to clipboard

Challenge: Currently, end-to-end models learn coreference resolution implicitly by observing aligned sentences in bilingual corpora.
Approach: They develop a method that derives coreference clusters from MT output and evaluates them without requiring annotations in the target language.
Outcome: The proposed model outperforms existing models on three challenging benchmarks.
Paths to Relation Extraction through Semantic Structure (2021.findings-acl)

Copied to clipboard

Challenge: Syntactic and semantic structure directly reflect relations expressed by the text at hand and are therefore very useful for relation extraction (RE)
Approach: They propose two methods for integrating broad-coverage semantic structure into supervised RE models by encoding semantic DAGs.
Outcome: The proposed methods overshadow the use of syntactic integrations in RE . they reduce UCCA into a bilexical structure and encode semantic DAG structures .
Language (Re)modelling: Towards Embodied Language Understanding (2020.acl-main)

Copied to clipboard

Challenge: Despite the rapid progress in NLU, current systems lack the rich mental representations that people use for language understanding.
Approach: They propose an approach to representation and learning based on the tenets of embodied cognitive linguistics (ECL) they propose a system architecture along with a roadmap towards realizing this vision.
Outcome: The proposed approach will improve the performance of existing systems and provide a roadmap towards realizing this vision.
Semantic Structural Decomposition for Neural Machine Translation (2020.starsem-1)

Copied to clipboard

Challenge: Existing methods for translation of long sentences are limited by the translation of single sentences to single sentences.
Approach: They propose to use semantic splitting of the source sentence as preprocessing for machine translation.
Outcome: The proposed approach tackles two main limitations of state-of-the-art machine translation.
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community (2025.acl-demo)

Copied to clipboard

Challenge: Using human-model conversations is a valuable resource for model development and research, but the open source and research community lags behind.
Approach: They propose a unified set of human conversations with large language models and a plugin for voluntarily contributing user-model conversations.
Outcome: The ShareLM collection and its plugin allow users to share conversations from most platforms.
Content Differences in Syntactic and Semantic Representation (N19-1)

Copied to clipboard

Challenge: Syntactic analysis plays an important role in semantic parsing, but the nature of this role remains a topic of ongoing debate.
Approach: They propose to use Universal Dependencies and UCCA as test cases to compare syntactic and semantic schemes.
Outcome: The proposed comparison methodology can be used for fine-grained evaluation of UCCA parsing, highlighting both challenges and potential sources for improvement.
Topical Segmentation of Spoken Narratives: A Test Case on Holocaust Survivor Testimonies (2022.emnlp-main)

Copied to clipboard

Challenge: Topical segmentation is a task that has been neglected in recent work . a drawback of this approach is the lack of interpretability, which is crucial in some contexts.
Approach: They propose to model running (spoken) narratives using topic segmentation . they hypothesize that boundary points between segments correspond to low mutual information .
Outcome: The proposed approaches show significant improvements over manual approaches.
PreQuEL: Quality Estimation of Machine Translation Outputs in Advance (2022.emnlp-main)

Copied to clipboard

Challenge: A PreQuEL system predicts how well a given sentence will be translated without recourse to the actual translation.
Approach: They propose a task that uses a model to predict how well a given sentence will be translated . they show that the model is sensitive to syntactic and semantic distinctions .
Outcome: The proposed model improves on the Quality-Estimation task and on challenge sets and languages.
Parallel Context Windows for Large Language Models (2023.acl-long)

Copied to clipboard

Challenge: Existing efforts to address context window limitation for off-the-shelf LLMs involve training specialized architectures.
Approach: They propose a method that carves a long context into chunks and restricts attention to apply only within each window.
Outcome: The proposed method shows significant improvements on in-context learning tasks with diverse input and output spaces.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations