“Seeing the Big through the Small”: Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Human label variation arises when multiple human annotators provide different labels for valid reasons. |
| Approach: | They propose to use crowd workers to represent human judgment distributions or expert linguists to provide detailed explanations for their chosen labels. |
| Outcome: | The proposed model can approximate human judgment distributions using a small number of expert labels and explanations. |
Similar Papers
A Rose by Any Other Name: LLM-Generated Explanations Are Good Proxies for Human Explanations to Collect Label Distributions on NLI (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent research has shown that explanations provide valuable information for understanding human label variation (HLV) Large language models (LLMs) can approximate HJD from a few human-provided label-explanation pairs, but collecting explanations for every label is still time-consuming. |
| Approach: | They propose to use Large Language Models (LLMs) as annotators to generate model explanations for a few given human labels. |
| Outcome: | The proposed models can generate human-provided explanations from human labels, but they are still time-consuming. |
Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown considerable promise for annotation purposes, but questions remain about their ability to capture human label variation (HLV) label variation is genuine disagreement between annotators observed across NLP tasks. |
| Approach: | They investigate how label and explanation variation manifests within and across LLMs with respect to the Natural Language Inference task. |
| Outcome: | The proposed models generate label distributions similar to humans but exhibit distinct, idiosyncratic judgments and disagreement patterns. |
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI (2026.findings-acl)
Copied to clipboard
| Challenge: | Human label variation (HLV) arises when multiple labels are valid for the same instance. |
| Approach: | They propose a framework for generating and validating explanations to detect errors using large language models (LLMs) EVADE framework provides broader explanation coverage and requires less human intervention . |
| Outcome: | The proposed framework provides broader explanation coverage, requires less human intervention, and delivers better downstream performance in predicting label distributions. |
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets. |
| Approach: | They propose to use an ensemble of large language models to flag mislabeled examples by using an LLM-as-a-judge approach to detect label errors in existing datasets. |
| Outcome: | The proposed method improves label accuracy and consistency in large language models. |
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evidence of human label variation in Natural Language Inference (NLI) however, within-label variation is an additional challenge. |
| Approach: | They propose a linguistically-informed taxonomy for categorizing free-text explanations in English that captures different reasoning strategies behind NLI explanations with a particular focus on within-label variation. |
| Outcome: | The proposed taxonomy can be used to classify explanations in English using a linguistically-informed taxonomies. |
Can Large Language Models Capture Dissenting Human Voices? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown impressive achievements in solving a broad range of tasks. |
| Approach: | They evaluate the performance and alignment of large language models with humans using Monte Carlo Estimation and Log Probability Estimationic methods to estimate the multinomial distribution. |
| Outcome: | The proposed models fail to capture human disagreement distribution and inference and human alignment performance plunge even further on data samples with high disagreement levels raising concerns about their natural language understanding ability and representativeness to a larger human population. |
Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large language models have shown the power of chain-of-thought reasoning in improving complex decision-making tasks. |
| Approach: | They propose a pipeline that generates chain-of-thought (CoT) explanations from CoTs with improved accuracy. |
| Outcome: | The proposed pipeline outperforms a direct generation method and baselines on three datasets. |
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)
Copied to clipboard
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andre Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, Alberto Testoni
| Challenge: | Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models . |
| Approach: | They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets. |
| Outcome: | The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets. |
Style Over Substance: Evaluation Biases for Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Ranking the relative performance of large language models based on Elo ratings is gaining popularity . however, the extent to which humans and LLMs are capable evaluators remains uncertain . |
| Approach: | They propose to evaluate machine-generated text across multiple dimensions using the Elo rating system . they propose to use crowd-sourced and expert annotators to rank models based on Elo ratings . |
| Outcome: | The proposed method improves the quality of LLM-based evaluations, but there is no improvement in crowd-sourced evaluations. |
Ecologically Valid Explanations for Label Variation in NLI (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Human label variation exists in many natural language processing tasks, including NLI . |
| Approach: | They build an English dataset of 1,415 ecologically valid explanations for 122 MNLI items . they find that people can systematically vary on their interpretation . |
| Outcome: | The proposed dataset contains 1,415 ecologically valid explanations for 122 items . the results show that people can vary on interpretation and highlight differences . |