Papers by Esin Durmus

17 papers
Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive Summarization (2022.acl-long)

Copied to clipboard

Challenge: Abstractive summarization systems still suffer from faithfulness errors, authors say . prior work has proposed models that improve faithfulness, but it is unclear whether this improvement comes from an increased level of extractiveness of the outputs.
Approach: They propose a faithfulness-abstractiveness trade-off curve that serves as a control . they also learn a selector to identify the most faithful and abstractive summary for a given document .
Outcome: The proposed model achieves higher faithfulness scores while being abstractive than the baseline system on two datasets.
A Corpus for Modeling User and Language Effects in Argumentation on Online Debating (P19-1)

Copied to clipboard

Challenge: Existing argumentation datasets have allowed only limited assessment of "user" traits because information on background of users is generally unavailable.
Approach: They present a dataset of 78,376 debates generated over a 10-year period along with surprisingly comprehensive participant profiles.
Outcome: The proposed dataset includes 78,376 debates generated over a 10-year period along with comprehensive participant profiles.
Exploring the Role of Argument Structure in Online Debate Persuasion (2020.emnlp-main)

Copied to clipboard

Challenge: Existing work in NLP has shown that linguistic features extracted from debate text and features encoding the characteristics of the audience are both critical in persuasion studies.
Approach: They propose to incorporate argument structure features into an LSTM-based model to assess the persuasiveness of debates.
Outcome: The proposed model incorporates argument structure features to predict debaters that make the most convincing arguments on online debate forums.
When Do Pre-Training Biases Propagate to Downstream Tasks? A Case Study in Text Summarization (2023.eacl-main)

Copied to clipboard

Challenge: Existing studies have shown that large language models contain linguistic and societal biases, but it is unclear how these biase amplify to downstream tasks.
Approach: They investigate how name-nationality bias propagates from pre-training to downstream tasks . they show that these biases manifest themselves as hallucinations in summarization .
Outcome: The proposed model can reduce the rate of hallucinations, but does not change the types of biases that do appear.
Exploring the Role of Prior Beliefs for Argument Persuasion (N18-1)

Copied to clipboard

Challenge: Recent studies in natural language processing (NLP) have shown that the language of opinion holders and their patterns of interaction play a key role in changing the mind of a reader.
Approach: They propose to use a dataset to study the effect of language use vs. prior beliefs on persuasion in a controlled setting that takes into account political and religious ideology.
Outcome: The proposed controlled setting takes into account political and religious ideology and shows that prior beliefs play a more important role than language use effects.
Improving Faithfulness by Augmenting Negative Summaries from Fake Documents (2022.emnlp-main)

Copied to clipboard

Challenge: Current abstractive summarization systems tend to hallucinate unfaithful content . however, the most common method does not disentangle factual errors from other errors.
Approach: They propose a back-translation-style approach to augment negative samples that mimic factual errors made by the model.
Outcome: The proposed method improves faithfulness without sacrificing informativeness . it incorporates negative samples into training, and produces faithful/unfaithful summaries .
Towards Reference-free Text Simplification Evaluation with a BERT Siamese Network Architecture (2023.findings-acl)

Copied to clipboard

Challenge: Text simplification (TS) aims to modify sentences to make their content and structure easier to understand.
Approach: They propose a neural-network-based TS metric that uses a human reference to evaluate simplification and meaning preservation.
Outcome: The proposed metric correlates better with human judgments for simplicity and meaning preservation than existing metrics.
Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to measure stereotypes in large language models rely on manual templates or natural sentences that contain stereotypes.
Approach: They propose a prompt-based method to measure stereotypes in large language models . they use natural language descriptions of the target demographic group alongside unmarked defaults .
Outcome: The proposed method detects that portrayals contain higher rates of racial stereotypes than human-written portrayals.
Spurious Correlations in Reference-Free Evaluation of Text Generation (2022.acl-long)

Copied to clipboard

Challenge: Recent work suggests that reference-free evaluation metrics may rely on spurious correlations with human judgments.
Approach: They propose to use model-based, reference-free evaluation metrics to evaluate natural language generation systems.
Outcome: The proposed metrics achieve high correlations with human judgments, but they may not be robust enough to evaluate their efficacy and robustness.
Determining Relative Argument Specificity and Stance for Complex Argumentative Structures (P19-1)

Copied to clipboard

Challenge: Existing work on claim specificity and stance has been limited to shallow arguments . a system that can determine the stance of claims employed in argumentation is not sufficient .
Approach: They propose to use a dataset of manually curated argument trees to study claim specificity and stance in argumentation.
Outcome: The proposed dataset consists of manually curated argument trees for 741 controversial topics covering 95,312 unique claims.
Contrastive Error Attribution for Finetuned Language Models (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for error tracing do not detect faithfulness errors in NLG datasets.
Approach: They propose a framework to identify and remove low-quality training instances that lead to undesirable outputs.
Outcome: The proposed method outperforms existing methods for detecting faithfulness errors in NLG datasets.
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)

Copied to clipboard

Challenge: Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work.
Approach: They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations.
Outcome: The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work.
Leveraging Topic Relatedness for Argument Persuasion (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies of argumentation focus on the effects of factors such as source, audience, and language style, but the impact of exploiting the relationships among controversial topics is under-explored.
Approach: They propose to model topic relatedness among controversial topics using topic embedding features and topic semantics features extracted from the arguments.
Outcome: The proposed method improves predicting persuasiveness and generalizes to rare topics in a few-shot setting.
The Role of Pragmatic and Discourse Context in Determining Argument Impact (D19-1)

Copied to clipboard

Challenge: Recent work shows that attributes of both the audience and communicator constitute important cues for determining argument strength.
Approach: They propose to use a dataset to study the pragmatic and discourse context of argumentative claims to build predictive models that incorporate the pragmatic context of the argument.
Outcome: The proposed models outperform models that rely on claim-specific linguistic features for predicting the perceived impact of individual claims within a particular line of argument.
WikiLingua: A New Benchmark Dataset for Cross-Lingual Abstractive Summarization (2020.findings-emnlp)

Copied to clipboard

Challenge: a lack of high quality multilingual data for cross-lingual summarization is a costly endeavor since it requires humans to read, comprehend, condense, and paraphrase entire articles.
Approach: They propose to use a large-scale, multilingual dataset to evaluate cross-lingual abstractive summarization systems.
Outcome: The proposed method significantly outperforms baseline approaches while being more cost efficient during inference.
FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization (2020.acl-main)

Copied to clipboard

Challenge: Existing automatic metrics do not capture errors in abstractive summarization models.
Approach: They propose an automatic question answering metric for faithfulness that leverages recent advances in reading comprehension.
Outcome: The proposed metric has significantly higher correlation with human faithfulness scores on highly abstracted summaries.
NLP Systems That Can’t Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps (2024.naacl-long)

Copied to clipboard

Challenge: Existing language models fail to distinguish use from mention, leading to misinformation and hate speech detection, resulting in censorship of counterspeech.
Approach: They propose prompting mitigations that teach the use-mention distinction and show they reduce these errors.
Outcome: The proposed model reduces misinformation and hate speech detection errors by reducing misinformation, and reducing hate speech.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations