Challenge: a lack of human-centered considerations about people’s needs for language technologies is causing an “evaluation crisis” in NLP.
Approach: This tutorial introduces perspectives and methodologies from human-computer interaction (HCI) it will introduce what to evaluate for, how generalizable the results are to the real-world contexts, and pragmatic costs to conduct the evaluation.
Outcome: This tutorial introduces perspectives and methodologies from human-computer interaction (HCI) the tutorial will also encourage reflection on how these HCI perspectives and methods can complement NLP evaluation through Q&A discussions and a hands-on exercise.

Similar Papers

Designing, Evaluating, and Learning from Humans Interacting with NLP Models (2023.emnlp-tutorial)

Copied to clipboard

Challenge: This tutorial will cover how to conduct human-in-the-loop usability evaluations to ensure that models are capable of interacting with humans.
Approach: They will provide a systematic overview of key considerations and effective approaches for studying human-NLP model interactions.
Outcome: This tutorial will cover how to conduct human-in-the-loop usability evaluations to ensure that models are capable of interacting with humans.
Human-Centered Evaluation of Explanations (2022.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial will provide an overview of human-centered evaluations of explanations .
Approach: This tutorial will provide an overview of human-centered evaluations of explanations . it will introduce the psychological foundation of explanation and types of NLP explanations.
Outcome: This tutorial will provide an overview of human-centered evaluations of explanations . it will cover the two categories of evaluation: evaluation based on human-annotated explanations and evaluation with human-subjects studies.
Lessons from a User Experience Evaluation of NLP Interfaces (2025.findings-naacl)

Copied to clipboard

Challenge: Increasingly, questions are being asked on whether evaluations are reproducible and repeatable.
Approach: They propose to design user interfaces that are more consistent and reproducible . only a minority of published experiments can be reproduced due to non-working code or resource limits .
Outcome: The proposed UIs are based on standardized human-centered interaction principles and are evaluated by four experts.
Navigating the Modern Evaluation Landscape: Considerations in Benchmarks and Frameworks for Large Language Models (LLMs) (2024.lrec-tutorials)

Copied to clipboard

Challenge: General-purpose Language Models have changed the world of Natural Language Processing, if not the world itself.
Approach: This tutorial will lay the foundations and explain the basics of evaluation and compare traditional methods to newly developed methods.
Outcome: The tutorial assumes little familiarity with metrics, datasets, prompts and benchmarks . it will compare traditional methods to newly developed methods .
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon the insights from disciplines such as user experience research and human behavioral psychology to ensure that the results are reliable.
Approach: They propose a framework for human evaluation of generative large language models that takes into account usability, aesthetics and cognitive biases.
Outcome: The proposed framework is based on the framework proposed by Deutsch and alnajjar . it is aimed at ensuring that human evaluation is accurate in the age of generative AI .
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications (2022.naacl-main)

Copied to clipboard

Challenge: Evaluating natural language generation systems is difficult, as there are many ways to express similar things in text.
Approach: They combine interviews with NLG practitioners to examine ethical considerations and their implications for NLG evaluation.
Outcome: The findings of the study surface goals, community practices, assumptions, and constraints that shape NLG evaluations, and examine their implications and how they embody ethical considerations.
Opportunities for Human-centered Evaluation of Machine Translation Systems (2022.findings-naacl)

Copied to clipboard

Challenge: a new study examines the role of machine translation in larger user-facing systems . a sysadmin and a human factors researcher are developing evaluation tools .
Approach: They argue that machine translation models are embedded in larger user-facing systems . they argue that evaluation at the systems level is still lacking .
Outcome: The proposed model evaluations are based on human-computer interaction models . the authors argue that evaluations should be based more on the entire system .
Human-AI Interaction in the Age of LLMs (2024.naacl-tutorials)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized the capabilities of AI systems.
Approach: This tutorial will provide an overview of the interaction between humans and Large Language Models (LLMs) it will start with a review of the types of AI models we interact with and walkthrough of the core concepts in Human-AI Interaction.
Outcome: This tutorial will provide an overview of the interaction between humans and LLMs, exploring the challenges, opportunities, and ethical considerations that arise in this dynamic landscape.
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges (2024.emnlp-main)

Copied to clipboard

Challenge: introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance.
Approach: They propose a taxonomy for organizing existing LLM-based evaluation metrics and a structured framework to understand and compare them.
Outcome: The proposed taxonomy offers a framework to understand and compare LLM-based evaluation methods.
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Approach: This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement .
Outcome: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations