Papers by Ori Shapira

17 papers
McPhraSy: Multi-Context Phrase Similarity and Clustering (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for estimating phrase similarity use the phrase context only during training, instead relying on the phrase itself.
Approach: They propose a novel algorithm that leverages multiple contexts during inference to estimate the similarity of phrases based on multiple context.
Outcome: The proposed method outperforms existing models on two phrase similarity datasets by 13.3% and a new task that relies on phrase similarities in the product reviews domain.
Re-Examining Summarization Evaluation across Multiple Quality Criteria (2023.findings-emnlp)

Copied to clipboard

Challenge: a number of automated evaluation metrics are evaluated by multiple quality criteria, such as relevance, consistency, fluency and coherence.
Approach: They propose a method that removes the confounding variable and detects unreliable correlations.
Outcome: The proposed method detects unreliable correlations between QCs and human scores . it is based on a multi-QC setup, but it fails to detect summary corruptions .
Better Rewards Yield Better Summaries: Learning to Summarise Without References (D19-1)

Copied to clipboard

Challenge: Reinforcement Learning (RL)-based document summarisation systems produce state-of-the-art performance in terms of ROUGE scores, but high summaries receive low human judgement.
Approach: They propose to learn a reward function from human ratings on 2,500 summaries to generate human-appealing summary.
Outcome: The proposed reward function can generate human-appealing summaries without reference summary input.
SummHelper: Collaborative Human-Computer Summarization (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing approaches for text summarization are mostly automated, with limited space for human intervention and control.
Approach: They propose a 2-phase summarization assistant that facilitates human-machine collaboration . it suggests possible content and generates a coherent summary from these selections . authors hope to improve the efficiency of the computer and human-involved approach .
Outcome: The proposed summarization assistant is a 2-phase summarizing assistant . it suggests potential content and consolidates the output with visual mappings . the proposed system is available for free on youtube .
Crowdsourcing Lightweight Pyramids for Manual Summary Evaluation (N19-1)

Copied to clipboard

Challenge: Manual evaluation methods are perceived as insufficient due to the high cost of the Pyramid method and the required expertise.
Approach: They propose a crowdsourced method that compares system summaries to references and uses crowdsourced scripts to analyze the results.
Outcome: The proposed method shows higher correlation relative to the original Pyramid method.
Extending Multi-Document Summarization Evaluation to the Interactive Setting (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches to interactive summarization are incomparable and divergent . a key gap in the development and adoption of interactive summaries is the lack of evaluation methodologies and benchmarks for meaningful comparison of systems.
Approach: They propose an end-to-end evaluation framework for interactive summarization based on expansion-based interaction . framework includes procedure of collecting real user sessions, evaluation measures relying on summarizing standards, but adapted to reflect interaction.
Outcome: The proposed evaluation framework is based on evaluations of baseline implementations and is available publicly as a benchmark.
Multi Document Summarization Evaluation in the Presence of Damaging Content (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing metrics evaluate a summary based on relevance and consistency with the source documents.
Approach: They propose to measure the ability of MDS systems to handle damaging documents in their input set by lexical similarity and language model likelihood.
Outcome: The proposed metrics show that they can summarize a set of documents without damaging content.
Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks (2025.acl-long)

Copied to clipboard

Challenge: a growing number of recorded human speech is recorded for automated processing, resulting in errors in the transcripts . a configurable framework is proposed to analyze transcript noise impact across noise levels and transcript-cleaning techniques.
Approach: They propose a configurable framework for assessing task models in diverse noisy settings . framework facilitates investigation of task model behavior, which can support effective SLU solutions.
Outcome: The proposed framework can analyze model behavior in various noise levels and transcript-cleaning techniques.
Proposition-Level Clustering for Multi-Document Summarization (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods focused on clustering sentences to indicate information saliency and avoid redundancy.
Approach: They propose to group together sub-sentential propositions to generate a representative sentence for each cluster via text fusion.
Outcome: The proposed method improves over the previous state-of-the-art method in the DUC 2004 and TAC 2011 datasets, both in automatic ROUGE scores and human preference.
The Overlooked Role of Graded Relevance Thresholds in Multilingual Dense Retrieval (2026.findings-acl)

Copied to clipboard

Challenge: Dense retrieval models are fine-tuned with contrastive learning objectives that require binary relevance judgments.
Approach: They examine how graded relevance scores affect multilingual dense retrieval . they argue that a well-chosen threshold can improve effectiveness and mitigate annotation noise .
Outcome: The optimal threshold varies systematically across languages and tasks, the authors show . a well-chosen threshold can improve effectiveness and mitigate annotation noise .
Evaluating Multiple System Summary Lengths: A Case Study (D18-1)

Copied to clipboard

Challenge: Practical summarization systems are expected to produce summaries of varying lengths, per user needs.
Approach: They propose to use ROUGE metric to evaluate system summaries of multiple lengths.
Outcome: The evaluation protocol in question is competitive, the authors show . they found that the evaluation protocol is competitive with existing benchmarks.
Multi-Review Fusion-in-Context (2024.findings-naacl)

Copied to clipboard

Challenge: Current methods for generating text are opaque and difficult to control and interpret due to their opaque nature.
Approach: They propose a modular approach with separate components for each step . they formalize Fusion-in-Context as a standalone task, whose input consists of source texts with highlighted spans of targeted content.
Outcome: The proposed approach is based on a curated dataset of 1000 instances in the reviews domain and a novel evaluation framework for assessing the faithfulness and coverage of highlights.
Interactive Query-Assisted Summarization via Deep Reinforcement Learning (2022.naacl-main)

Copied to clipboard

Challenge: Existing systems that can perform interactive summarization cannot ingest the full document set or operate at sufficient speed for interactivity.
Approach: They propose two deep reinforcement learning models for interactive summarization task . they use interactive session state and history to refrain from redundancy .
Outcome: The proposed model improves informativeness while preserving positive user experience.
The Power of Summary-Source Alignments (2024.findings-acl)

Copied to clipboard

Challenge: Multi-document summarization (MDS) is a challenging task, often decomposed to subtasks of salience and redundancy detection, followed by text generation.
Approach: They propose to extend the summary-source alignment framework by applying it at the more fine-grained proposition span level and annotating alignment manually in a multi-document setup.
Outcome: The proposed framework can yield several datasets for at least six different tasks.
iFacetSum: Coreference-based Interactive Faceted Summarization for Multi-Document Exploration (2021.emnlp-demo)

Copied to clipboard

Challenge: iFS provides a faceted navigation scheme that provides abstractive summaries for the user’s selections.
Approach: They propose a web application that integrates interactive summarization and faceted search to provide a faceted navigation scheme that yields abstractive summaries for the user's selections.
Outcome: The proposed system provides a comprehensive overview as well as particular details regard-ing subtopics of interest.
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to assess data quality for training and testing large language models are lacking.
Approach: They propose two approaches to assess the reliability of data for training large language models for external tool usage.
Outcome: The proposed approaches outperform models trained on high-quality data on two popular benchmarks and an extrinsic evaluation that showcases the impact of data quality on model performance.
OpenAsp: A Benchmark for Multi-document Open Aspect-based Summarization (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models focus on a limited set of predefined aspects, resulting in a lack of realistic open aspect setting.
Approach: They propose a benchmark for multi-document open aspect-based summarization using an annotation protocol.
Outcome: The proposed benchmark satisfies the needs of users in real-world scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations