Challenge: Prior work on reasoning about time series in conjunction with natural language has largely overlooked event descriptions and focused on tasks involving just numeric data like trend analysis or anomaly detection.
Approach: They propose a method for generating tasks that test a model’s ability to reason about events associated with time series data based on sports data and develop a benchmarking method.
Outcome: The proposed method can infer unobserved events from time series data, even when providing minimal context.

Similar Papers

Evaluating Large Language Models on Time Series Feature Understanding: A Comprehensive Taxonomy and Benchmark (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a critical tool for time series analysis and reporting in many fields, including healthcare, finance, climate, and many more.
Approach: They propose a framework for rigorously evaluating the capabilities of Large Language Models (LLMs) on time series understanding, encompassing both univariate and multivariate forms.
Outcome: The proposed framework delineates various characteristics inherent in time series data.
EvEntS ReaLM: Event Reasoning of Entity States via Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to model event implications fail to reason about the world, despite their knowledge of physical attributes.
Approach: They propose to use a model prompting technique to prompt models of event implications by targeting their understanding of physical attributes.
Outcome: The proposed model prompting technique is especially useful for unseen attributes or when only limited data is available.
Language Models Still Struggle to Zero-shot Reason about Time Series (2024.findings-emnlp)

Copied to clipboard

Challenge: Time series are critical for decision-making in fields like finance and healthcare.
Approach: They propose a framework for time series reasoning that includes formal tasks and a dataset of multi-scale time series paired with text captions across ten domains.
Outcome: The proposed framework combines formal tasks and a dataset of multi-scale time series paired with text captions across ten domains to examine whether language models achieve three forms of reasoning.
Harnessing LLMs for Temporal Data - A Study on Explainable Financial Time Series Forecasting (2023.emnlp-industry)

Copied to clipboard

Challenge: Recent advances in machine learning and artificial intelligence have opened up numerous opportunities and challenges in financial time series forecasting.
Approach: They propose to use Large Language Models for explainable financial time series forecasting to leverage cross-sequence information and extract insights from text and price time series.
Outcome: The proposed model outperforms ARMA-GARCH and gradient-boosting tree models while underperforming on other models.
Temporal Reasoning in Natural Language Inference (2020.findings-emnlp)

Copied to clipboard

Challenge: We use five new natural language inference (NLI) datasets focused on temporal reasoning.
Approach: They introduce five new natural language inference datasets focused on temporal reasoning.
Outcome: The proposed models capture the temporal reasoning of four existing datasets.
Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Temporal reasoning is a vital component of human communication and understanding, yet remains an underexplored area within the context of Large Language Models (LLMs).
Approach: They propose to use 3 prompting strategies to evaluate 8 different LLMs across 6 datasets and 2 Code Generation LMs to perform the analysis.
Outcome: The proposed models perform better on NLP tasks than the standard models on the same dataset.
Can Large Language Models Adequately Perform Symbolic Reasoning Over Time Series? (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) and Multimodal LLMs (MLLMs) show strong performance in complex reasoning tasks, but their ability to extract symbolic laws from time series data remains underexplored.
Approach: They propose a benchmark to assess symbolic reasoning over real-world time series across three tasks: multivariate symbolic regression, Boolean network inference, and causal discovery.
Outcome: The proposed framework integrates LLMs with genetic programming to form a closed-loop symbolic reasoning system.
Causal Inference with Large Language Model: A Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Existing causal inference frameworks do not match human judgment in several key areas, such as domain knowledge, logical inference, and cultural context.
Approach: They propose to apply large language models to causal inference tasks . they summarize the main causal problems and approaches and compare their results .
Outcome: The proposed methods are compared with traditional methods in healthcare, finance, and economics.
Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and Natural Language Data (2026.eacl-long)

Copied to clipboard

Challenge: Large language models are being rapidly applied across many fields such as healthcare, finance, transportation, and energy.
Approach: They propose a large language model framework that integrates time-series tokens into LLMs’ vocabulary, enhancing its reasoning ability over time- and textual data.
Outcome: The proposed framework enhances reasoning ability over time-series and textual data without compromising core natural language capabilities.
Forecasting Future International Events: A Reliable Dataset for Text-Based Event Modeling (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for text-based event prediction are limited in quality due to dynamic nature of international relations and conflicting economic dynamics.
Approach: They propose a novel dataset that leverages the advanced reasoning capabilities of large-language models to address these limitations.
Outcome: The proposed dataset features high-quality scoring labels generated through advanced prompt modeling and rigorously validated by domain experts in political science.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations