Challenge: Existing Ordinal Classification metrics ignore the ordering between items or assume additional information.
Approach: They propose a Closeness Evaluation Measure for Ordinal Classification based on Measurement Theory and Information Theory.
Outcome: The proposed metric captures quality aspects from different traditional tasks simultaneously.

Similar Papers

Evaluating Evaluation Measures for Ordinal Classification and Ordinal Quantification (2021.acl-long)

Copied to clipboard

Challenge: Ordinal Classification (OC) tasks require ordinal classes, not nominal ones, to be evaluated.
Approach: They use data from the SemEval and NTCIR communities to clarify evaluation measures for Ordinal Classification and Ordinal Quantification tasks.
Outcome: The evaluation measures for Ordinal Classification (OC) and Ordinal Quantification (OQ) tasks are ordinal, not nominal.
Exploring Ordinality in Text Classification: A Comparative Study of Explicit and Implicit Techniques (2024.findings-acl)

Copied to clipboard

Challenge: Ordinal classification (OC) is a key task in natural language processing with applications in various domains such as sentiment analysis, rating prediction, and more.
Approach: They propose to tackle ordinal classification (OC) through the implicit semantics of the labels . they propose to use a classical explicit approach and an implicit approach that organically engages the semantics.
Outcome: The proposed methods are based on pre-trained language models and offer strategic recommendations based upon specific settings.
What happens if you treat ordinal ratings as interval data? Human evaluations in NLP are even more under-powered than you think (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that human evaluations in NLP are under-powered because of two common factors: they treat ordinal data as interval data and operate under high variance settings.
Approach: They propose to use ordinal mixed effects models to detect small differences between models, especially in high variance settings common in NLP evaluations of generated texts.
Outcome: The proposed models detect small differences in high variance settings, especially in high-variance evaluations of generated texts.
SMATCH++: Standardized and Extended Evaluation of Semantic Graphs (2023.findings-eacl)

Copied to clipboard

Challenge: Existing graph-alignment metrics that measure graph distances are not reliable, we show . metric is spread out and does not provide upper bounds for extended tasks.
Approach: They propose a metric to measure a distance between graphs by aligning nodes and counting matching graph triples.
Outcome: The proposed method reduces search space and improves scoring by reducing the number of errors.
Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics (2021.tacl-1)

Copied to clipboard

Challenge: Existing fairness metrics quantify the differences in a model’s behaviour across a range of demographic groups.
Approach: They propose to unify existing fairness metrics and compare them to three generalized fairness measures to reveal the connections between them.
Outcome: The proposed measures can be explained by differences in parameter choices, and the results are consistent with previous studies.
On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations (2022.acl-short)

Copied to clipboard

Challenge: Recent natural language processing systems use large language models as the backbone . however, societal biases are encoded in these models and transferred to downstream applications .
Approach: They propose to use two categories to measure fairness in natural language processing tasks . they find intrinsic and extrinsic metrics do not correlate in their original setting .
Outcome: The proposed metrics do not correlate in their original setting, the authors show . they find that they are not accurate when correcting for metric misalignments and noise .
Evaluating Extreme Hierarchical Multi-label Classification (2022.acl-long)

Copied to clipboard

Challenge: Several natural language processing tasks are defined as a classification problem in its most complex form: Multi-label Hierarchical Extreme classification.
Approach: They propose a classification metric inspired by the Information Contrast Model (ICM) they use a set of formal properties to analyze the evaluation metrics.
Outcome: The proposed evaluation metrics are suitable for multi-label hierarchical extreme classification scenarios.
Studying Summarization Evaluation Metrics in the Appropriate Scoring Range (P19-1)

Copied to clipboard

Challenge: Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate.
Approach: They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
Outcome: The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)

Copied to clipboard

Challenge: a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes .
Approach: They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones .
Outcome: The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions.
Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics are conflated and can mislead models, resulting in downstream harms.
Approach: They propose a framework for conceptualizing and evaluating the reliability and validity of evaluation metrics based on empirical data.
Outcome: The proposed framework formalizes the source of measurement error and offers statistical tools for evaluating evaluation metrics based on empirical data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations