Challenge: 'information value' quantifies the predictability of an utterance relative to a set of plausible alternatives.
Approach: They propose a method to obtain interpretable estimates of information value using neural text generators and exploit their psychometric predictive power to investigate the dimensions of predictability that drive human comprehension behaviour.
Outcome: The proposed method is able to obtain interpretable estimates of information value using neural text generators and exploits their psychometric predictive power to investigate the dimensions of predictability that drive human comprehension behaviour.

Similar Papers

Towards a Similarity-adjusted Surprisal Theory (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that surprisal theory ignores the possibility of similarity between words and treats them as distinct entities.
Approach: They propose a new measure of comprehension effort called information value that accounts for communicative equivalences between possible continuations.
Outcome: The proposed measure of comprehension effort is based on the diversity index of the diversity of communicative units.
Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers (2021.findings-emnlp)

Copied to clipboard

Challenge: Large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, but statistical bias in benchmark data and probing studies has recently called into question their true capabilities.
Approach: They propose to evaluate systems through a measure of prediction coherence by using two existing language understanding benchmarks with different properties to demonstrate its versatility.
Outcome: The proposed evaluation framework is quick, effective, and versatile to provide insight into the coherence of machines’ predictions.
Not Every Metric is Equal: Cognitive Models for Predicting N400 and P600 Components During Reading Comprehension (2025.coling-main)

Copied to clipboard

Challenge: Several studies have focused on predicting the surprisal of a word and its reading time, but only recently, attention has been given to other components, such as P600.
Approach: They propose to model reading times and ERP amplitudes using surprisal and entropy . they also propose a metric based on semantic similarity for N400 and P600 .
Outcome: The proposed metric predicts reading times and ERP amplitudes in Mandarin Chinese.
Information Parity: Measuring and Predicting the Multilingual Capabilities of Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in user-facing applications worldwide, necessitating handling multiple languages across various tasks.
Approach: They propose a metric called Information Parity (IP) that can predict an LLM’s capabilities across multiple languages in a task-agnostic manner.
Outcome: The proposed metric can predict LLM’s capabilities across multiple languages in a task-agnostic manner.
Language models emulate certain cognitive profiles: An investigation of how predictability measures interact with individual differences (2024.findings-acl)

Copied to clipboard

Challenge: incorporating cognitive capacities increases predictive power of surprisal and entropy measures on reading data, whereas high performance in the psychometric tests is associated with lower sensitivity to predictability effects.
Approach: They examine the predictive power (PP) of surprisal and entropy estimated from generative language models (LMs) on reading data from individuals who also completed a wide range of psychometric tests.
Outcome: The LMs' predictive power is based on cognitive capacities and high performance in psychometric tests is associated with lower sensitivity to predictability effects.
Large Language Models for Predictive Analysis: How Far Are They? (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on LLMs do not evaluate their capability in predictive analysis.
Approach: They propose a benchmark to evaluate Large Language Models (LLMs) they integrate 1130 queries from 44 real-world datasets of 8 different fields to evaluate their capability .
Outcome: The proposed benchmark evaluates 12 renowned LLMs from 44 real-world datasets . results offer insights into their practical use in predictive analysis .
Prediction or Comparison: Toward Interpretable Qualitative Reasoning (2021.findings-acl)

Copied to clipboard

Challenge: Qualitative relationships are a significant portion of textual knowledge . current approaches use semantic parsers to transform natural language inputs into logical expressions or a "black-box" model to solve them in one step.
Approach: They propose to use neural network modules to simulate qualitative reasoning tasks . they use two qualitative reasoning question answering datasets to test their methods .
Outcome: Experiments on two qualitative reasoning question answering datasets show the proposed methods are general and general and interpretable.
Comparing and Developing Tools to Measure the Readability of Domain-Specific Texts (D19-1)

Copied to clipboard

Challenge: Despite this, we lack a thorough understanding of how to validly measure readability at scale, especially for domain-specific texts.
Approach: They present a comparison of the validity of well-known readability measures and introduce a novel approach to measure readability at scale.
Outcome: The proposed approach addresses shortcomings of existing measures.
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies.
Approach: They propose a psychometric evaluation pipeline grounded in realistic human-AI interactions to probe value orientations and novel tasks for evaluating value understanding in an open-ended value space.
Outcome: The proposed evaluation pipeline is grounded in realistic human-AI interactions and performs tasks that approximate expert conclusions in value-related extraction and generation tasks.
Beyond Facts- Benchmarking Distributional Reading Comprehension in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing reading comprehension benchmarks focus on factual information, but many real-world tasks require distributional knowledge expressed across text.
Approach: They propose a reading comprehension benchmark for LLMs to evaluate their ability to infer distributional knowledge from natural language.
Outcome: Experiments with multiple LLMs show that the model outperforms baselines, but performance varies widely across distribution types and characteristics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations