Papers by Chengye Wang

6 papers
FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents (2024.emnlp-main)

Copied to clipboard

Challenge: FinDVer is a benchmark to evaluate the explainable claim verification capabilities of LLMs . financial documents are typically long, intricate and dense, and they include both quantita and numerical reasoning.
Approach: They propose a benchmark to evaluate the explainable claim verification capabilities of LLMs . they assess 25 LLM systems under long-context and RAG settings .
Outcome: The proposed benchmark can be used to evaluate the explainable claim verification capabilities of LLMs in financial documents.
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction (2026.acl-long)

Copied to clipboard

Challenge: Existing document OCR largely targets plain text or Markdown, discarding structural and executable properties that make LaTeX essential for scientific publishing.
Approach: They propose a benchmark and a training corpus for document reconstruction . they train a 2B-parameter model using supervised fine-tuning and reinforcement learning .
Outcome: The proposed model improves on existing models using supervised fine-tuning and reinforcement learning with verifiable rewards.
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research (2025.acl-long)

Copied to clipboard

Challenge: a benchmark designed to evaluate the capabilities of LLMs in designing ablation studies for scientific research is available online.
Approach: They propose to use a benchmark to evaluate LLMs' ability to design ablation studies . they investigate whether current automated evaluation methods are not reliable .
Outcome: The benchmark compared leading LLMs with human experts on generating detailed ablation study designs . the results show that current evaluation methods are not reliable for the task .
SciMDR: Advancing Scientific Multimodal Document Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Current models struggle to provide reliable assistance in real-world scientific workflows because evidence is distributed across long, multimodal documents.
Approach: They propose a framework for QA Synthesis and document-scale regrounding that generates faithful, isolated QA pairs and reasoning on focused segments.
Outcome: The proposed framework achieves significant improvements across multiple QA benchmarks, particularly in tasks requiring complex document-level reasoning.
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers (2025.findings-acl)

Copied to clipboard

Challenge: MISS-QA is the first benchmark specifically designed to evaluate the ability of models to interpret schematic diagrams within scientific literature.
Approach: They propose an automated evaluation protocol powered by open-source LLMs trained on human-scored data to ensure reliable evaluation.
Outcome: The proposed protocol is powered by open-source LLMs trained on human-scored data.
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification (2025.acl-long)

Copied to clipboard

Challenge: Existing scientific claim verification benchmarks focus on textual content alone or on verifying claims based on a single table.
Approach: They propose to use SciVer to evaluate the ability of foundation models to verify claims within a multimodal scientific context.
Outcome: The proposed model outperforms 21 state-of-the-art models and human experts on SciVer.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations