| Challenge: | Human cognition is deeply intertwined with a sense of time, known as Chronoception, which allows us to judge how long facts remain valid and when knowledge becomes outdated. |
| Approach: | They propose a model that captures nuanced patterns of emergence, decay, and peak relevance using skew-normal curves fitted along semantically decomposed temporal axes. |
| Outcome: | The proposed model captures nuanced patterns of emergence, decay, and peak relevance in two datasets. |
Similar Papers
Mitigating Temporal Misalignment by Discarding Outdated Facts (2023.emnlp-main)
Copied to clipboard
| Challenge: | Temporal misalignment is a problem for knowledge-intensive tasks where models must rely on data from the past to make predictions. |
| Approach: | They propose a temporal misalignment task to predict how long a given fact will remain true. |
| Outcome: | The proposed task improves calibration for knowledge-intensive tasks under temporal misalignment by discarding volatile facts. |
Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for testing time scales treat reasoning traces or tokens equally, ignoring substantial variations in trajectory quality and localized logical failures. |
| Approach: | They propose a chronological reasoning scorer that models each trajectory as a time series. |
| Outcome: | The proposed method achieves relative improvements of 34.21% over Pass@128 and 22.70% over Maj@135 on HMMT25, highlighting its effectiveness. |
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Grasping the concept of time is a fundamental facet of human cognition. |
| Approach: | They propose a hierarchical temporal reasoning benchmark that covers a broad spectrum of temporal phenomena. |
| Outcome: | The proposed benchmark shows that state-of-the-art LLMs are still far behind humans in temporal reasoning . |
ChronoBias: A Benchmark for Evaluating Temporal Group Bias in the Time-sensitive Knowledge of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Using a template-based semi-automated generation method, we evaluate time-conditional group bias in time-sensitive knowledge of large language models (LLMs). |
| Approach: | They propose a template-based semi-automated generation method to construct a time-conditional group bias benchmark. |
| Outcome: | The proposed method balancing quality-quantity trade-off in existing benchmark curation approaches. |
TSVer: A Benchmark for Fact Verification Against Time-Series Evidence (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing systems for fact-checking lack structured evidence, provide insufficient justifications for verdicts, or rely on synthetic claims. |
| Approach: | They propose a temporal and numerical reasoning dataset based on time-series evidence that is annotated with time frames and a verdict and justifications reflecting how the evidence is used to reach the verdict. |
| Outcome: | The proposed dataset improves the quality of the annotations and achieves an inter-annotator agreement of = 0.745 on verdicts. |
ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events (2025.acl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) still face significant challenges in reasoning and arithmetic. |
| Approach: | They propose a new benchmark to evaluate LLMs' temporal understanding that includes 16 tasks identifying the Allen relation between two temporal events and temporal arithmetic. |
| Outcome: | The proposed model handles Allen relations, even symmetrical ones, quite differently. |
Confidence is not Timeless: Modeling Temporal Validity for Rule-based Temporal Knowledge Graph Forecasting (2024.acl-long)
Copied to clipboard
| Challenge: | Existing literature on temporal knowledge Graph Forecasting lacks in-depth investigation into how confidence evolves with time. |
| Approach: | They propose a framework to model the temporal validity of rules for Temporal Knowledge Graph Forecasting (TKGF) they propose rule-adversarial negative sampling and time-aware negative sampling strategies to facilitate TempValid learning. |
| Outcome: | The proposed framework outperforms state-of-the-art (SOTA) rule-based methods on six TKGF datasets. |
TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Reasoning-oriented language models expose explicit reasoning as a long, front-loaded chain of “thinking” tokens before the main output, either always enabled or externally toggled at inference time. |
| Approach: | They introduce a behavioral alignment framework that learns explicit reasoning as a context-triggered control policy rather than a fixed response mode. |
| Outcome: | The proposed framework improves TIMEBench scores over the base model in thinking and no-thinking modes while keeping output compact. |
A Diachronic Perspective on User Trust in AI under Uncertainty (2023.emnlp-main)
Copied to clipboard
| Challenge: | Modern NLP systems are rarely calibrated and are often confidently incorrect about their predictions, which violates users’ mental model and erodes their trust. |
| Approach: | They propose to use a mental model to bet on the correctness of an NLP system and to study how trust is rebuilt as a function of time after these events. |
| Outcome: | The proposed model shows that even a few highly inaccurate confidence estimation instances damage users’ trust in the system and performance, which does not easily recover over time. |
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties (2026.eacl-short)
Copied to clipboard
| Challenge: | Existing temporal reasoning benchmarks rely on rule-based construction and lack contextual depth . a recent study found existing LLMs struggle with nuanced temporal understanding . |
| Approach: | a benchmark is designed to evaluate LLMs on temporal reasoning in Chinese dynasties. |
| Outcome: | a new benchmark evaluates LLMs on temporal reasoning across Chinese dynasties . it emphasizes cross-entity relationships, pairwise temporal alignment, contextualized and culturally-grounded reasoning . results show existing LLM benchmarks struggle with nuanced temporal understanding . |