ChronoBias: A Benchmark for Evaluating Temporal Group Bias in the Time-sensitive Knowledge of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Using a template-based semi-automated generation method, we evaluate time-conditional group bias in time-sensitive knowledge of large language models (LLMs). |
| Approach: | They propose a template-based semi-automated generation method to construct a time-conditional group bias benchmark. |
| Outcome: | The proposed method balancing quality-quantity trade-off in existing benchmark curation approaches. |
Similar Papers
Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for testing time scales treat reasoning traces or tokens equally, ignoring substantial variations in trajectory quality and localized logical failures. |
| Approach: | They propose a chronological reasoning scorer that models each trajectory as a time series. |
| Outcome: | The proposed method achieves relative improvements of 34.21% over Pass@128 and 22.70% over Maj@135 on HMMT25, highlighting its effectiveness. |
DateLogicQA: Benchmarking Temporal Biases in Large Language Models (2025.naacl-srw)
Copied to clipboard
| Challenge: | DateLogicQA examines temporal biases in Large Language Models (LLMs) 190 questions are curated by humans to examine temporal reasoning across date formats and contexts . |
| Approach: | They propose a human-curated benchmark of 190 questions specifically designed to understand temporal bias in Large Language Models. |
| Outcome: | The proposed dataset covers seven date formats across past, present, and future contexts . it examines four reasoning types: commonsense, factual, conceptual, and numerical . |
ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events (2025.acl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) still face significant challenges in reasoning and arithmetic. |
| Approach: | They propose a new benchmark to evaluate LLMs' temporal understanding that includes 16 tasks identifying the Allen relation between two temporal events and temporal arithmetic. |
| Outcome: | The proposed model handles Allen relations, even symmetrical ones, quite differently. |
ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) struggle with ex-ante reasoning—making inferences or predictions without access to future information. |
| Approach: | They propose a benchmark that assesses LLMs’ ex-ante inference ability across four tasks: stock prediction, question answering, Wikipedia event generation, and scientific publication generation. |
| Outcome: | The proposed benchmark assesses LLMs’ ex-ante inference ability across four tasks. |
MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown nearly saturated performance on many NLP tasks. |
| Approach: | They construct multiple sensitive factors time QA which encompasses three temporal factors . they test current mainstream LLMs with different parameter sizes . |
| Outcome: | The proposed model incorporates three temporal factors with 2,853 samples . the results show that LLMs fall behind smaller models on these factors . |
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing efforts to ensure temporal consistency in large language models are lacking in time-sensitive fields . temporal reasoning is essential for time- sensitive fields such as finance and healthcare . a new benchmark aims to improve temporal referent consistency of LLMs . |
| Approach: | They propose a temporal referential consistency benchmark with a resource TEMP-ReCon to assess LLMs across temporal references. |
| Outcome: | The proposed model improves LLMs' temporal consistency by comparing them to baseline models. |
Chronocept: Instilling a Sense of Time in Machines (2026.eacl-srw)
Copied to clipboard
| Challenge: | Human cognition is deeply intertwined with a sense of time, known as Chronoception, which allows us to judge how long facts remain valid and when knowledge becomes outdated. |
| Approach: | They propose a model that captures nuanced patterns of emergence, decay, and peak relevance using skew-normal curves fitted along semantically decomposed temporal axes. |
| Outcome: | The proposed model captures nuanced patterns of emergence, decay, and peak relevance in two datasets. |
TWBias: A Benchmark for Assessing Social Bias in Traditional Chinese Large Language Models through a Taiwan Cultural Lens (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models have shown remarkable capabilities in natural language processing, but concerns about social bias amplification remain. |
| Approach: | They propose a social bias evaluation benchmark for Traditional Chinese LLMs that integrates chat templates and diverse prompts for comprehensive bias assessment. |
| Outcome: | The proposed model incorporates chat templates and diverse prompts for comprehensive bias assessment focusing on Taiwan's cultural context and prioritizing gender and ethnicity bias evaluation. |
Is Your LLM Outdated? A Deep Look at Temporal Generalization (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing methods to evaluate large language models are limited due to their inherent dynamic nature and the inherent dynamicity of language and information. |
| Approach: | They introduce a new evaluation framework that employs fresh text and event prediction for assessing LLMs’ temporal adaptability. |
| Outcome: | The proposed framework shows significant temporal biases and a decline in performance over time. |
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Grasping the concept of time is a fundamental facet of human cognition. |
| Approach: | They propose a hierarchical temporal reasoning benchmark that covers a broad spectrum of temporal phenomena. |
| Outcome: | The proposed benchmark shows that state-of-the-art LLMs are still far behind humans in temporal reasoning . |