Papers by Amir Zeldes

20 papers
GUM-SAGE: A Novel Dataset and Approach for Graded Entity Salience Prediction (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for graded entity salience are subjective but lack consistency.
Approach: They propose a method for graded entity salience that combines subjective judgments and summarization-based methods that define saliency as mention-worthiness in a summary.
Outcome: The proposed approach outperforms existing methods and shows stronger correlation with human summaries and alignments.
To Ask LLMs about English Grammaticality, Prompt Them in a Different Language (2024.findings-emnlp)

Copied to clipboard

Challenge: a study focuses on questions about grammar and fluency in multilingual LLMs . english is the dominant training language for all three models, but prompting in a different language often yields better results.
Approach: They ask three multilingual language models in multiple languages to test their model's grammatical accuracy.
Outcome: The language of the prompt can significantly affect model performance, the study finds . english is the dominant training language for all three models, the researchers show .
A Deeper Look into Dependency-Based Word Embeddings (N18-4)

Copied to clipboard

Challenge: Word embeddings trained with dependency contexts excel at different tasks, and enhanced dependencies often improve performance.
Approach: They propose to use dependency-based word embeddings to capture semantic similarity rather than relatedness.
Outcome: The results show that word embeddings trained with Universal and Stanford dependencies excel at different tasks and that enhanced dependencies often improve performance.
SPLICE: A Singleton-Enhanced PipeLIne for Coreference REsolution (2024.lrec-main)

Copied to clipboard

Challenge: Existing attempts to integrate singleton mention detection into end-to-end coreference resolution for English have been hampered by the lack of singletont mention spans in the OntoNotes benchmark.
Approach: They propose a two-step neural mention and coreference resolution system that integrates singleton mentions with OntoNotes syntax trees to achieve a near approximation of the Ontonotes dataset with all singletont mentions.
Outcome: The proposed system achieves 94% recall on a sample of gold singletons.
UCxn: Typologically-Informed Annotation of Constructions Atop Universal Dependencies (2024.lrec-main)

Copied to clipboard

Challenge: Grammatical constructions that convey meaning through a particular combination of several morphosyntactic elements are not labeled holistically.
Approach: They propose to augment UD annotations with a ‘UCxn’ annotation layer for such meaning-bearing grammatical constructions and to approach this in a typologically informed way so that morphosyntactic strategies can be compared across languages.
Outcome: The proposed annotation layer could be used to annotate meaning-bearing constructions across languages and to compare them across languages.
GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains (2024.emnlp-main)

Copied to clipboard

Challenge: Existing shallow discourse parsing systems focus on the Wall Street Journal corpus, but the data is limited to the news domain and is 35 years old.
Approach: They propose to use the Wall Street Journal corpus as a benchmark for PDTB-style shallow discourse parsing.
Outcome: The proposed dataset is compatible with PDTB, but suffers from degradation out-of-domain.
ELQA: A Corpus of Metalinguistic Questions and Answers about English (2023.acl-long)

Copied to clipboard

Challenge: ELQA corpus is metalinguistic—it consists of language about language.
Approach: They present a corpus of questions and answers in and about the English language . they use a free-form question answering task and multiple LLMs to analyze their capacity .
Outcome: The ELQA corpus covers grammar, meaning, fluency, and etymology . the results can be used to investigate metalinguistic capabilities of NLU models .
Universal Anaphora: The First Three Years (2024.lrec-main)

Copied to clipboard

Challenge: Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by expanding the aspects of anaphonic interpretation which are or can be reliably annotated in an anagraphic corpora.
Approach: They propose to develop a standard for anaphoric annotations and a method for evaluating models that can carry out this type of interpretation.
Outcome: The Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by producing unified standards to annotate and encode annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation.
OntoGUM: Evaluating Contextualized SOTA Coreference Resolution on 12 More Genres (2021.acl-short)

Copied to clipboard

Challenge: Existing methods for coreference resolution are unable to evaluate generalizability to open domain data.
Approach: They propose to make an OntoNotes-like coreference dataset publicly available and convert it into an English corpus.
Outcome: The proposed dataset is the largest human-annotated coreference corpus following the OntoNotes guidelines and the first to be evaluated for consistency with the OnToNote's scheme.
GCDT: A Chinese RST Treebank for Multigenre and Multilingual Discourse Parsing (2022.aacl-short)

Copied to clipboard

Challenge: GCDT is the largest hierarchical discourse treebank for Mandarin Chinese in the framework of Rhetorical Structure Theory (RST).
Approach: They propose to use a Chinese hierarchical discourse treebank to parse Mandarin Chinese using relation inventory and a multilingual training program.
Outcome: The proposed dataset includes state-of-the-art scores for Chinese RST parsing and RST Parsing on the English GUM dataset, using cross-lingual training in Chinese and English with multilingual embeddings.
Syntax as a Rosetta Stone: Universal Dependencies for In-Context Coptic Translation (2026.findings-acl)

Copied to clipboard

Challenge: Existing work using bilingual dictionaries to support inference for vocabulary items is lacking for low-resource languages.
Approach: They propose to use universal dependency parses of input sentences to augment in-context learning prompts for low resource machine translation for the Coptic language.
Outcome: The proposed approach achieves state-of-the-art results for the Coptic language.
GUMSum: Multi-Genre Data and Evaluation for English Abstractive Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing datasets are limited to newswire text, which is a fraction of extant genres in general and on the Web.
Approach: They present a small but carefully crafted dataset of English summaries in 12 written and spoken genres for evaluation of abstractive summarization.
Outcome: The proposed dataset of English summaries in 12 written and spoken genres is compared with human outputs and compared to untuned and prompt-based approaches.
Why Can’t Discourse Parsing Generalize? A Thorough Investigation of the Impact of Data Diversity (2023.eacl-main)

Copied to clipboard

Challenge: Discourse parsing performance is not reliable for high-resource languages such as English . a heterogeneous training regime is critical for stable and generalizable models .
Approach: They investigate the impact of genre diversity on RST parsing stability . they use two largest RST corpora of English with text from multiple genres .
Outcome: The proposed model can generalize to text types unseen during training, but it is not reliable for high-resource languages.
CorefUD 1.0: Coreference Meets Universal Dependencies (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in standardization for annotated language resources have led to successful large scale efforts, such as the Universal Dependencies (UD) project for multilingual syntactically annotized data.
Approach: They propose a multilingual collection of corpora and a standardized format for coreference resolution compatible with morphosyntactic annotations in the UD framework.
Outcome: The proposed framework is compatible with morphosyntactic annotations and includes facilities for related tasks such as named entity recognition.
Treebanking User-Generated Content: A Proposal for a Unified Representation in Universal Dependencies (2020.lrec-1)

Copied to clipboard

Challenge: Despite the increasing number of contributions on Part-of-Speech tagging and parsing, automatic processing of user-generated content (UGC) still represents a challenging task.
Approach: They propose a set of guidelines for the annotation of user-generated texts within the Universal Dependencies framework.
Outcome: The proposed annotation guidelines promote cross-linguistic consistency, which has always been in the spirit of UD.
A Second Wave of UD Hebrew Treebanking and Cross-Domain Parsing (2022.emnlp-main)

Copied to clipboard

Challenge: Foundational Hebrew NLP tasks have relied on various versions of the Hebrew Treebank . however, the data in the HTB is now over 30 years old and does not cover many aspects of contemporary Hebrew on the web.
Approach: They propose to use Hebrew Wikipedia to stratify the text from a UD treebank.
Outcome: The proposed treebank is based on a single-source newswire corpus selected from Hebrew Wikipedia.
Expect the Unexpected? Testing the Surprisal of Salient Entities (2026.acl-long)

Copied to clipboard

Challenge: Existing work on the Uniform Information Density hypothesis has neglected the relative salience of discourse participants.
Approach: They propose to use an annotated text to examine how overall salience of entities in discourse relates to surprisal.
Outcome: The proposed method shows that global salience is a mechanism shaping information distribution in discourse.
GUMsley: Evaluating Entity Salience in Summarization for 12 English Genres (2024.eacl-long)

Copied to clipboard

Challenge: Existing work on salient entity extraction relies on crowdsourcing or user statistics to derive labels for entities.
Approach: They propose a dataset that defines salience using human summaries and shows high agreement between annotations based on whether a source entity is mentioned in the summary.
Outcome: The proposed dataset shows that pre-trained models and zero-shot LLM prompting fail to capture salient entities in generated summaries.
DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse Processing (2024.lrec-main)

Copied to clipboard

Challenge: DISRPT is a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing.
Approach: They present a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing that includes 13 languages and 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks.
Outcome: The DISRPT dataset includes data from 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks.
AMALGUM – A Free, Balanced, Multilayer English Web Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 4M tokens is available online with a large number of high-quality annotation layers.
Approach: They propose to use a genre-balanced English web corpus with multiple annotation layers . they harness knowledge from multiple annotation layer to achieve a "better than NLP" benchmark .
Outcome: The proposed corpus is genre-balanced and features high-quality automatic annotation layers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations