Challenge: State-of-the-art models in NLP are opaque in terms of how they come to make predictions.
Approach: They propose to release a benchmark to measure the quality of rationales extracted by models and how faithful these rationale are to human annotators.
Outcome: The proposed benchmark will enable researchers to compare models and track progress on interpretable models for NLP.

Similar Papers

Plausible Extractive Rationalization through Semi-Supervised Entailment Signal (2024.findings-acl)

Copied to clipboard

Challenge: Abstract: Large language models are gaining widespread adoption in natural language processing tasks.
Approach: They propose a semi-supervised approach to optimize for plausibility of extracted rationales by using a pre-trained natural language inference model and a supervised NLI predictor.
Outcome: The proposed model outperforms unsupervised models by > 100% on a ERASER dataset.
Goodhart’s Law Applies to NLP’s Explanation Benchmarks (2024.findings-eacl)

Copied to clipboard

Challenge: Popular methods for "explaining" the outputs of natural language processing (NLP) models operate by highlighting a subset of input tokens that ought, in some sense, to be salient.
Approach: They propose to inflate a model’s comprehensiveness and sufficiency scores dramatically without altering its predictions or explanations on in-distribution inputs.
Outcome: The proposed metrics exploit the tendency for extracted explanations and complements to be “out-of-support” relative to each other and in-distribution inputs.
Model Interpretability and Rationale Extraction by Input Mask Optimization (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for creating explanations for black-box models struggle with deriving easily interpretable explanations.
Approach: They propose a model-agnostic method to generate extractive explanations for neural network predictions using masking parts of the input that the model does not consider indicative of the respective class.
Outcome: The proposed method achieves state-of-the-art results in a paragraph-level rationale extraction task, showing that this task can be performed without training a specialized model.
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Approach: This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement .
Outcome: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Did the Models Understand Documents? Benchmarking Models for Language Understanding in Document-Level Relation Extraction (2023.acl-long)

Copied to clipboard

Challenge: Document-level relation extraction (DocRE) models achieve consistent performance gains in DocRE, but their underlying decision rules are still understudied.
Approach: They propose to use annotations to provide rationales for document-level relation extraction (DocRE) they then propose to apply a method to evaluate models' reasoning capabilities .
Outcome: The proposed models exhibit different reasoning processes in contrast to humans . the proposed models render models more trustworthy and robust to be deployed in real-world scenarios.
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
Interpretation of NLP models through input marginalization (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods to interpret NLP predictions replace each token with a predefined value, resulting in misleading interpretations.
Approach: They propose to marginalize each token out of the training data distribution to demystify the "black box" property of deep neural networks for natural language processing.
Outcome: The proposed method marginalizes each token out of the training data distribution.
Navigating the Modern Evaluation Landscape: Considerations in Benchmarks and Frameworks for Large Language Models (LLMs) (2024.lrec-tutorials)

Copied to clipboard

Challenge: General-purpose Language Models have changed the world of Natural Language Processing, if not the world itself.
Approach: This tutorial will lay the foundations and explain the basics of evaluation and compare traditional methods to newly developed methods.
Outcome: The tutorial assumes little familiarity with metrics, datasets, prompts and benchmarks . it will compare traditional methods to newly developed methods .
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks (2023.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape . established automatic evaluation metrics are poor surrogates, correlating weakly with human judgement.
Approach: They propose to use both automatic and human evaluation to evaluate generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction.
Outcome: The proposed model outperforms many popular models according to human reviewers on the majority of metrics, while scoring much worse when using classic automatic evaluation metrics.
DHP Benchmark: Are LLMs Good NLG Evaluators? (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly serving as evaluators in Natural Language Generation (NLG) tasks.
Approach: They propose a framework that measures the discernment of Large Language Models (LLMs) across diverse NLG tasks.
Outcome: The proposed framework provides quantitative discernment scores for LLMs across four NLG tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations