Challenge: Large language models are being rapidly applied across many fields such as healthcare, finance, transportation, and energy.
Approach: They propose a large language model framework that integrates time-series tokens into LLMs’ vocabulary, enhancing its reasoning ability over time- and textual data.
Outcome: The proposed framework enhances reasoning ability over time-series and textual data without compromising core natural language capabilities.

Similar Papers

Language Models Still Struggle to Zero-shot Reason about Time Series (2024.findings-emnlp)

Copied to clipboard

Challenge: Time series are critical for decision-making in fields like finance and healthcare.
Approach: They propose a framework for time series reasoning that includes formal tasks and a dataset of multi-scale time series paired with text captions across ten domains.
Outcome: The proposed framework combines formal tasks and a dataset of multi-scale time series paired with text captions across ten domains to examine whether language models achieve three forms of reasoning.
Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement (2025.acl-long)

Copied to clipboard

Challenge: Existing time series models focus on a narrow spectrum of tasks, such as forecasting or anomaly detection.
Approach: They propose a framework that enables natural language queries across multiple time series tasks such as numerical analytical tasks and open-ended question answering with reasoning.
Outcome: The proposed framework enables natural language queries across multiple time series tasks and allows for more advanced and intuitive interactions with temporal data.
LLaTiSA: Towards Difficulty-Stratified Time Series Reasoning from Visual Perception to Semantics (2026.findings-acl)

Copied to clipboard

Challenge: Current research hinders the development of unified Time Series Reasoning Models (TSRMs) time series data are a fundamental modality for capturing the temporal dynamics of complex systems.
Approach: They propose a time series reasoning model that integrates visualized patterns with precision-calibrated numerical tables to enhance the temporal perception of Vision-Language Models.
Outcome: The proposed model outperforms existing models and exhibits robust out-of-distribution generalization across diverse tasks and real-world scenarios.
A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated powerful reasoning abilities across multiple domains, but have been underexplored for time-series reasoning (TsR)
Approach: They propose a prompt-based solution for evaluating large language models’ TsR performance.
Outcome: The proposed solution improves performance and costs by 140% and reduces costs by 99%.
Inferring Events from Time Series using Language Models (2026.acl-long)

Copied to clipboard

Challenge: Prior work on reasoning about time series in conjunction with natural language has largely overlooked event descriptions and focused on tasks involving just numeric data like trend analysis or anomaly detection.
Approach: They propose a method for generating tasks that test a model’s ability to reason about events associated with time series data based on sports data and develop a benchmarking method.
Outcome: The proposed method can infer unobserved events from time series data, even when providing minimal context.
Large Language Models Can Learn Temporal Reasoning (2024.acl-long)

Copied to clipboard

Challenge: Temporal reasoning (TR) is a fundamental ability of large language models (LLMs) however, there is neo-standard methods to perform TR, which are not suitable for large language model applications.
Approach: They propose a framework to enhance temporal reasoning by using a latent representation, temporal graph (TG) instead of reasoning over the original context, they adopt a temporal representation that enhances TR learning.
Outcome: The proposed framework improves the learning of language-based TR by incorporating a latent representation, temporal graph (TG) a synthetic dataset is constructed for fine-tuning LLMs on text-to-TG translation tasks and benchmarks.
TRANSIENTTABLES: Evaluating LLMs’ Reasoning on Temporally Evolving Semi-structured Tables (2025.naacl-long)

Copied to clipboard

Challenge: a recent study shows that large language models are limited in their ability to reason over time due to static datasets.
Approach: They present a dataset that includes 3,971 questions derived from over 14,000 tables . they introduce a template-based question-generation pipeline that harnesses LLMs to refine questions .
Outcome: The proposed model improves on the TRANSIENTTABLES dataset . it demonstrates that the model can reason over time, even when it is not static .
TimeSAF: Towards LLM-Guided Semantic Asynchronous Fusion for Time Series Forecasting (2026.acl-long)

Copied to clipboard

Challenge: Existing time series forecasting methods use a deep synchronous fusion strategy . high-level abstract semantics are inappropriately entangled with low-level temporal dynamics .
Approach: They propose a framework based on hierarchical asynchronous fusion that decouples unimodal feature learning from cross-modal interaction.
Outcome: The proposed framework outperforms state-of-the-art approaches on long-term forecasting benchmarks.
NL2TL: Transforming Natural Languages to Temporal Logics using Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Temporal Logic (TL) can be used to specify complex high-level specifications for systems in many engineering domains.
Approach: They propose a framework for translation between NL and TL using Large Language Models . they use a dataset to create a model with 23K NL-TL pairs and human annotation .
Outcome: The proposed framework achieves higher accuracy (> 95%) using only 10% training data compared with baseline model.
CaTS-Bench: Can Language Models Describe Time Series? (2026.findings-acl)

Copied to clipboard

Challenge: Existing time series captioning benchmarks rely on fully synthetic or generic captions . authors propose a pipeline for generating high-fidelity synthetic captions, which is validated .
Approach: They propose a benchmark for Context-aware Time Series reasoning across 11 diverse domains . they evaluate leading Vision-Language Models on their benchmark .
Outcome: The proposed benchmark evaluates 1746 human-rewritten captions and shows they perform better than open-source models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations