Automatic Slide Updating with User-Defined Dynamic Templates and Natural Language Instructions (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing automation methods follow fixed template filling and cannot support dynamic updates for diverse, user-authored decks. |
| Approach: | They propose a framework that combines multimodal slide parsing, natural language instruction grounding, and tool-augmented reasoning for tables, charts, and textual conclusions. |
| Outcome: | The proposed framework updates content while preserving layout and style while maintaining a strong reference baseline on DynaSlide. |
Similar Papers
Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation (2026.findings-acl)
Copied to clipboard
| Challenge: | Talk-to-Your-Slides is a high-efficiency slide editing agent that uses language-driven structured data manipulation instead of the image modality. |
| Approach: | They propose a language-driven slide editing agent that uses language-based structured data manipulation instead of image modality. |
| Outcome: | The proposed system achieves faster processing and better instruction fidelity than GUI-based agents. |
D2S: Document-to-Slide Generation Via Query-Based Text Summarization (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing research efforts to automate the document-to-slide generation process face a critical challenge: no publicly available dataset for training and benchmarking. |
| Approach: | They propose a dataset SciDuet that gathers papers and their corresponding slides from recent years’ NLP and ML conferences. |
| Outcome: | The proposed system outperforms state-of-the-art summarization baselines on both automated ROUGE metrics and qualitative human evaluation. |
SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding (2026.acl-long)
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) are a promising tool for document understanding, but they are not able to handle complex multi-page visual documents. |
| Approach: | They propose a flexible agentic framework for understanding multi-modal, multi-page, and multi-layout documents . SlideAgent employs specialized agents and decomposes reasoning into three specialized levels . |
| Outcome: | a new agentic framework improves accuracy over open-source and proprietary models . it decomposes reasoning into three levels to capture themes and visual cues . the framework is based on a multimodal large language model and a MLLM . |
Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to producing presentation slides rely on fixed templates or executable code . Existing methods rely only on predefined templates and emit executable codes . |
| Approach: | They propose a hierarchical slides generation workflow DeepSlides that organizes slide design tasks without any predefined template or style. |
| Outcome: | The proposed framework outperforms baseline methods on evaluated metrics and achieves superior performance in human preference evaluations. |
Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current automatic speech recognition systems rely on only audio information, ignoring multi-modal context. |
| Approach: | They propose to integrate visual context into existing automatic speech recognition systems to integrate presentation slides with multi-modal information. |
| Outcome: | The proposed model reduces word error rate by approximately 34% across all words and 35% for domain-specific terms compared to baseline model. |
Presentations by the Humans and For the Humans: Harnessing LLMs for Generating Persona-Aware Slides from Documents (2024.eacl-long)
Copied to clipboard
Ishani Mondal, Shwetha S, Anandhavelu Natarajan, Aparna Garimella, Sambaran Bandyopadhyay, Jordan Boyd-Graber
| Challenge: | Existing efforts to automate document-to-slide generation have failed to adapt to the persona of target audience or duration of presentation. |
| Approach: | They propose a concept of end-user specification-aware document to slides conversion that incorporates end- user specifications into the conversion process. |
| Outcome: | The proposed model can create persona-aware presentations tailored to the persona of target audience and cognitive abilities of target audiences. |
PresentAgent: Multimodal Agent for Presentation Video Generation (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Existing methods for generating static slides or text summaries are limited to producing narrated presentations. |
| Approach: | They propose a multimodal agent that transforms long-form documents into narrated presentations. |
| Outcome: | The present agent produces fully synchronized visual and spoken content that closely mimics human-style presentations. |
SlideGuard: AI-Driven Evaluation of Graduate Student Presentation Materials (2026.acl-demo)
Copied to clipboard
Nikolay Alekseevich Butakov, Maria Khodorchenko, Mazein Nikita, Daniil Gareev, Yuri Falevskiy, Georgii Konev, Denis Nasonov
| Challenge: | Effective communication is a core objective of graduate education in AI and machine learning (ML). |
| Approach: | They propose an evaluation agent that assesses slide decks against a framework of expert-defined criteria using a visual language model. |
| Outcome: | The evaluation agent detects the majority of expert-identified issues on a dataset of 150 annotated slide decks and shows strong results on structural and visual criteria and known limitations on subjective dimensions such as research quality. |
PreGenie: An Agentic Framework for High-quality Visual Presentation Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Visual presentations are vital for effective communication, but they are limited by their complexity and lack of visual understanding. |
| Approach: | a new framework is proposed to generate high-quality visual presentations using multimodal large language models. |
| Outcome: | The proposed framework outperforms existing models in multimodal understanding and content consistency. |
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides (2025.emnlp-main)
Copied to clipboard
Hao Zheng, Xinyan Guan, Hao Kong, Wenkai Zhang, Jia Zheng, Weixiang Zhou, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun
| Challenge: | Existing methods for generating presentations from documents focus on improving and evaluating content quality in isolation, overlooking visual appeal and structural coherence. |
| Approach: | They propose an edit-based presentation generation system that analyzes and iterates on slides to create new slides. |
| Outcome: | The proposed presentation generation tool outperforms existing methods in three dimensions . it analyzes slides, iterates and generates edit actions based on selected slides . |