Challenge: Existing benchmarks for scientific diagram generation rely on image-centric metrics or evaluation of intermediate symbolic representations rather than final rendered images.
Approach: They propose a structure-first benchmark for evaluating scientific diagram generation from pixel-level outputs.
Outcome: The proposed benchmark evaluates scientific diagram generation directly from pixel-level outputs.

Similar Papers

SciSketch: An Open-source Framework for Automated Schematic Diagram Generation in Scientific Papers (2025.emnlp-demos)

Copied to clipboard

Challenge: SCISKETCH is an open-source framework that supports two automated workflows for schematic diagram generation using foundation models.
Approach: They propose an open-source framework that supports two automated workflows for schematic diagram generation using foundation models.
Outcome: The open-source framework outperforms several state-of-the-art foundation models in generating schematic diagrams for scientific papers.
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement (2024.findings-emnlp)

Copied to clipboard

Challenge: Current text-to-image models struggle with generating accurate diagrams from long-context inputs.
Approach: They propose a task that extracts relevant information from scientific papers and generates diagrams based on user intentions using intermediate code generation.
Outcome: The proposed task outperforms existing models on factual correctness and visual appeal and outperfies existing ones on automatic and human judgement.
StarFlow: Generating Structured Workflow Outputs From Sketch Images (2026.eacl-long)

Copied to clipboard

Challenge: Despite being widely used, building workflows can be complex, often requiring manual configuration through low-code platforms or visual programming tools.
Approach: They propose a framework for generating structured workflow outputs from sketches using vision-language models to automate the process.
Outcome: The proposed framework outperforms large vision-language models in the task of generating structured workflow outputs from sketches and diagrams.
SciCompanion: Graph-Grounded Reasoning for Structured Evaluation of Scientific Arguments (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing retrieval-augmented generation (RAG) methods fail to provide deep, relational understanding of scientific literature.
Approach: They propose a graph-grounded reasoning framework for structured scientific evaluation that uses multi-hop reasoning to iteratively construct contextual graphs and generate structured critiques.
Outcome: The proposed framework reduces evaluation error by over 30% compared to baselines and allows smaller models to outperform larger models.
PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for scientific poster generation lack hierarchical document understanding and coherent content-layout planning.
Approach: They propose a training-free framework for scientific poster generation that captures document hierarchy and semantics across multiple levels.
Outcome: The proposed framework outperforms existing methods in both automatic and human evaluations without additional training or domain-specific supervision.
SciRepEval: A Multi-Format Benchmark for Scientific Document Representations (2023.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for evaluating scientific document representations fail to capture the diversity of relevant tasks.
Approach: They propose a benchmark for training and evaluating scientific document representations that includes 24 challenging and realistic tasks across four formats: classification, regression, ranking and search.
Outcome: The proposed model outperforms existing models by over 2 points absolute.
LEGOBench: Scientific Leaderboard Generation Benchmark (2024.findings-emnlp)

Copied to clipboard

Challenge: a growing number of papers make it difficult to stay informed about the latest state-of-the-art research.
Approach: They propose a benchmark to evaluate systems that generate scientific leaderboards . they use 22 years of submission data on arXiv and 11k machine learning leaderboard data on paperswithcode .
Outcome: The proposed model shows significant performance gaps in the LEGOBench model . the model is based on a language model and four graph-based leaderboard generation task configuration .
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature (2026.eacl-long)

Copied to clipboard

Challenge: Existing retrieval-augmented generation methods overlook citation graph structure, adapt poorly to complex queries, and yield fragmented, hard-to-verify syntheses.
Approach: They propose a retrieval-augmented generation framework that addresses these gaps by combining adaptive retrieval and symbolic reasoning.
Outcome: Extensive experiments show that SciRAG outperforms prior systems in factual accuracy and synthesis quality.
DIAGRAMS : A Review Framework for Reasoning-Level Attribution in Diagram QA (2026.acl-demo)

Copied to clipboard

Challenge: Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive the answer.
Approach: They propose a diagram question-answer review framework that decouples interface logic from dataset-specific JSON structures through an internal meta-schema and dataset adapters.
Outcome: The proposed framework achieves 85.39% precision and 75.30% recall across six diagram QA datasets.
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems (2026.acl-long)

Copied to clipboard

Challenge: Existing multimodal large language models (MLLMs) exhibit significant limitations when extracting essential information and reasoned properties from diagrams and performing complex reasoning based on these visual inputs.
Approach: They propose a benchmark that provides a fine-grained evaluation of MLLMs’ perception and reasoning capabilities.
Outcome: The proposed benchmark shows that existing MLLMs exhibit limitations when extracting essential information and reasoned properties from diagrams and performing complex reasoning based on these visual inputs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations