Papers by Chengye Wang
FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents (2024.emnlp-main)
Copied to clipboard
Yilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu, Xiangru Tang, Yiming Zhang, Chen Zhao, Arman Cohan
| Challenge: | FinDVer is a benchmark to evaluate the explainable claim verification capabilities of LLMs . financial documents are typically long, intricate and dense, and they include both quantita and numerical reasoning. |
| Approach: | They propose a benchmark to evaluate the explainable claim verification capabilities of LLMs . they assess 25 LLM systems under long-context and RAG settings . |
| Outcome: | The proposed benchmark can be used to evaluate the explainable claim verification capabilities of LLMs in financial documents. |
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction (2026.acl-long)
Copied to clipboard
| Challenge: | Existing document OCR largely targets plain text or Markdown, discarding structural and executable properties that make LaTeX essential for scientific publishing. |
| Approach: | They propose a benchmark and a training corpus for document reconstruction . they train a 2B-parameter model using supervised fine-tuning and reinforcement learning . |
| Outcome: | The proposed model improves on existing models using supervised fine-tuning and reinforcement learning with verifiable rewards. |
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research (2025.acl-long)
Copied to clipboard
Yilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, Arman Cohan
| Challenge: | a benchmark designed to evaluate the capabilities of LLMs in designing ablation studies for scientific research is available online. |
| Approach: | They propose to use a benchmark to evaluate LLMs' ability to design ablation studies . they investigate whether current automated evaluation methods are not reliable . |
| Outcome: | The benchmark compared leading LLMs with human experts on generating detailed ablation study designs . the results show that current evaluation methods are not reliable for the task . |
SciMDR: Advancing Scientific Multimodal Document Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Current models struggle to provide reliable assistance in real-world scientific workflows because evidence is distributed across long, multimodal documents. |
| Approach: | They propose a framework for QA Synthesis and document-scale regrounding that generates faithful, isolated QA pairs and reasoning on focused segments. |
| Outcome: | The proposed framework achieves significant improvements across multiple QA benchmarks, particularly in tasks requiring complex document-level reasoning. |
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers (2025.findings-acl)
Copied to clipboard
| Challenge: | MISS-QA is the first benchmark specifically designed to evaluate the ability of models to interpret schematic diagrams within scientific literature. |
| Approach: | They propose an automated evaluation protocol powered by open-source LLMs trained on human-scored data to ensure reliable evaluation. |
| Outcome: | The proposed protocol is powered by open-source LLMs trained on human-scored data. |
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification (2025.acl-long)
Copied to clipboard
| Challenge: | Existing scientific claim verification benchmarks focus on textual content alone or on verifying claims based on a single table. |
| Approach: | They propose to use SciVer to evaluate the ability of foundation models to verify claims within a multimodal scientific context. |
| Outcome: | The proposed model outperforms 21 state-of-the-art models and human experts on SciVer. |