DataTales: A Benchmark for Real-World Intelligent Data Narration (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks fail to capture the requisite analytical complexity for practical applications. |
| Approach: | They propose a benchmark to assess the proficiency of language models in data narration. |
| Outcome: | The proposed model combines financial reports with market data to demonstrate proficiency in data narration. |
Similar Papers
DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts (2024.emnlp-main)
Copied to clipboard
| Challenge: | Data-driven storytelling uses visual aids and visualizations to convey insights. |
| Approach: | They propose a task for data story generation using large language models and a benchmark containing 1,449 stories from diverse sources. |
| Outcome: | The proposed framework outperforms non-agentic counterparts in both model-based and human evaluations, but also reveals unique challenges in data story generation. |
NarraBench: A Comprehensive Framework for Narrative Benchmarking (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for narrative understanding are poorly aligned with existing metrics. |
| Approach: | They propose to use NarraBench to assess aspects of narrative understanding that are either overlooked in current work or are poorly aligned with existing metrics. |
| Outcome: | The proposed taxonomy and survey are useful to NLP researchers . they find that only 27% of tasks are well captured by existing benchmarks . |
Movie101v2: Improved Movie Narration Benchmark (2025.acl-long)
Copied to clipboard
| Challenge: | Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences. |
| Approach: | They propose to break down the ultimate goal of automatic movie narration into three stages . they propose a large-scale, bilingual dataset with enhanced data quality . |
| Outcome: | The proposed dataset breaks down the goal of automatic movie narration into three stages . achieving applicable movie narration is a fascinating goal that requires significant research . |
NarraSum: A Large-Scale Dataset for Abstractive Narrative Summarization (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies focus on summarizing news documents or structured documents. |
| Approach: | They propose to use a large-scale narrative summarization dataset to encourage research . they find there is a performance gap between humans and the models on NarraSum . |
| Outcome: | The proposed dataset shows that humans and state-of-the-art models perform poorly when summarizing a narrative . it contains 122K narratives collected from synopses of movies and TV episodes with diverse genres . |
Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora. |
| Approach: | They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch . |
| Outcome: | The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model. |
Chart Question Answering from Real-World Analytical Narratives (2025.acl-srw)
Copied to clipboard
| Challenge: | a dataset for chart question answering is constructed from visualization notebooks . data visualizations are an essential modality for communicating complex information about data. |
| Approach: | They propose a dataset for chart question answering constructed from visualization notebooks . they use real-world, multi-view charts paired with natural language questions . |
| Outcome: | The proposed dataset is constructed from student-authored visualization notebooks . it features real-world, multi-view charts paired with natural language questions . initial evaluations highlight significant performance gaps . |
ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding (2022.acl-long)
Copied to clipboard
| Challenge: | Large language models have shown exciting progress on several NLP benchmarks . however, evaluating their ability for complex analogical reasoning remains under-explored . |
| Approach: | They propose a dataset of narratives for employing proverbs in context as a benchmark for abstract language understanding. |
| Outcome: | The proposed dataset provides fine-grained annotation of aligned spans between proverbs and narratives and contains minimal overlaps between narratives with proverb . the results show that large language models struggle on these tasks compared to humans, and these tasks pose multiple learning challenges. |
Chart-to-Text: A Large-Scale Benchmark for Chart Summarization (2022.acl-long)
Copied to clipboard
Shankar Kantharaj, Rixie Tiffany Leong, Xiang Lin, Ahmed Masry, Megh Thakkar, Enamul Hoque, Shafiq Joty
| Challenge: | Inferring key insights from charts can be challenging and time-consuming. |
| Approach: | They propose a task where the goal is to explain a chart and summarize key takeaways from it in natural language. |
| Outcome: | The proposed model produces fluent summaries but suffers from hallucinations and factual errors . the proposed model is compared with other models and can be used to generate BLEU scores . |
Beyond Facts- Benchmarking Distributional Reading Comprehension in Large Language Models (2026.findings-acl)
Copied to clipboard
Pei-Fu Guo, Ya An Tsai, Chun-Chia Hsu, Kai-Xin Chen, Yun-Da Tsai, Kai-Wei Chang, Nanyun Peng, Mi-Yen Yeh, Shou-De Lin
| Challenge: | Existing reading comprehension benchmarks focus on factual information, but many real-world tasks require distributional knowledge expressed across text. |
| Approach: | They propose a reading comprehension benchmark for LLMs to evaluate their ability to infer distributional knowledge from natural language. |
| Outcome: | Experiments with multiple LLMs show that the model outperforms baselines, but performance varies widely across distribution types and characteristics. |
FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | FinChart-Bench is the first benchmark specifically focused on real-world financial charts. |
| Approach: | They propose a benchmark specifically focused on real-world financial charts. |
| Outcome: | The proposed benchmark evaluates 26 state-of-the-art LVLMs on FinChart-Bench. |