An Effectiveness Metric for Ordinal Classification: Formal Properties and Experimental Results (2020.acl-main)
Copied to clipboard
| Challenge: | Existing Ordinal Classification metrics ignore the ordering between items or assume additional information. |
| Approach: | They propose a Closeness Evaluation Measure for Ordinal Classification based on Measurement Theory and Information Theory. |
| Outcome: | The proposed metric captures quality aspects from different traditional tasks simultaneously. |
Similar Papers
Evaluating Evaluation Measures for Ordinal Classification and Ordinal Quantification (2021.acl-long)
Copied to clipboard
| Challenge: | Ordinal Classification (OC) tasks require ordinal classes, not nominal ones, to be evaluated. |
| Approach: | They use data from the SemEval and NTCIR communities to clarify evaluation measures for Ordinal Classification and Ordinal Quantification tasks. |
| Outcome: | The evaluation measures for Ordinal Classification (OC) and Ordinal Quantification (OQ) tasks are ordinal, not nominal. |
Exploring Ordinality in Text Classification: A Comparative Study of Explicit and Implicit Techniques (2024.findings-acl)
Copied to clipboard
Siva Rajesh Kasa, Aniket Goel, Karan Gupta, Sumegh Roychowdhury, Pattisapu Priyatam, Anish Bhanushali, Prasanna Srinivasa Murthy
| Challenge: | Ordinal classification (OC) is a key task in natural language processing with applications in various domains such as sentiment analysis, rating prediction, and more. |
| Approach: | They propose to tackle ordinal classification (OC) through the implicit semantics of the labels . they propose to use a classical explicit approach and an implicit approach that organically engages the semantics. |
| Outcome: | The proposed methods are based on pre-trained language models and offer strategic recommendations based upon specific settings. |
What happens if you treat ordinal ratings as interval data? Human evaluations in NLP are even more under-powered than you think (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that human evaluations in NLP are under-powered because of two common factors: they treat ordinal data as interval data and operate under high variance settings. |
| Approach: | They propose to use ordinal mixed effects models to detect small differences between models, especially in high variance settings common in NLP evaluations of generated texts. |
| Outcome: | The proposed models detect small differences in high variance settings, especially in high-variance evaluations of generated texts. |
SMATCH++: Standardized and Extended Evaluation of Semantic Graphs (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing graph-alignment metrics that measure graph distances are not reliable, we show . metric is spread out and does not provide upper bounds for extended tasks. |
| Approach: | They propose a metric to measure a distance between graphs by aligning nodes and counting matching graph triples. |
| Outcome: | The proposed method reduces search space and improves scoring by reducing the number of errors. |
Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing fairness metrics quantify the differences in a model’s behaviour across a range of demographic groups. |
| Approach: | They propose to unify existing fairness metrics and compare them to three generalized fairness measures to reveal the connections between them. |
| Outcome: | The proposed measures can be explained by differences in parameter choices, and the results are consistent with previous studies. |
On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations (2022.acl-short)
Copied to clipboard
Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, Aram Galstyan
| Challenge: | Recent natural language processing systems use large language models as the backbone . however, societal biases are encoded in these models and transferred to downstream applications . |
| Approach: | They propose to use two categories to measure fairness in natural language processing tasks . they find intrinsic and extrinsic metrics do not correlate in their original setting . |
| Outcome: | The proposed metrics do not correlate in their original setting, the authors show . they find that they are not accurate when correcting for metric misalignments and noise . |
Evaluating Extreme Hierarchical Multi-label Classification (2022.acl-long)
Copied to clipboard
| Challenge: | Several natural language processing tasks are defined as a classification problem in its most complex form: Multi-label Hierarchical Extreme classification. |
| Approach: | They propose a classification metric inspired by the Information Contrast Model (ICM) they use a set of formal properties to analyze the evaluation metrics. |
| Outcome: | The proposed evaluation metrics are suitable for multi-label hierarchical extreme classification scenarios. |
Studying Summarization Evaluation Metrics in the Appropriate Scoring Range (P19-1)
Copied to clipboard
| Challenge: | Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate. |
| Approach: | They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate. |
| Outcome: | The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate. |
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)
Copied to clipboard
| Challenge: | a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes . |
| Approach: | They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones . |
| Outcome: | The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions. |
Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics are conflated and can mislead models, resulting in downstream harms. |
| Approach: | They propose a framework for conceptualizing and evaluating the reliability and validity of evaluation metrics based on empirical data. |
| Outcome: | The proposed framework formalizes the source of measurement error and offers statistical tools for evaluating evaluation metrics based on empirical data. |