Papers with PDF
Autodive: An Integrated Onsite Scientific Literature Annotation Tool (2023.acl-demo)
Copied to clipboard
| Challenge: | Annotating scientific literature directly on PDF documents can greatly improve the labeling efficiency of scientists whose annotation costs are very high. |
| Approach: | They propose an integrated onsite scientific literature annotation tool for natural scientists and Natural Language Processing (NLP) researchers. |
| Outcome: | The proposed tool supports the whole lifecycle of corpus generation including i)project management, ii)resource management, and iv)ontology management, as well as manual annotation, onsite auto annotation, and vi)task statistic. |
ATLAS: A System for PDF-centric Human Interaction Data Collection (2024.naacl-demo)
Copied to clipboard
| Challenge: | Recent advances in AI only make the importance of high-quality data more pronounced. |
| Approach: | They propose to use the Portable Document Format (PDF) as a data format to better support researchers in collecting rich PDF-centric datasets from users. |
| Outcome: | The proposed toolkit and extensible schema allows researchers to customize the data collection tasks for a variety of purposes, including annotations, drawing, and reading behavior analytics. |
DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering (2024.naacl-demo)
Copied to clipboard
| Challenge: | DOCMASTER is a platform for annotating PDF documents, model training, and inference, tailored to document question-answering. |
| Approach: | They propose to integrate layout information into a unified platform for annotating PDF documents, model training, and inference tailored to document question-answering. |
| Outcome: | The proposed platform is designed for annotating PDF documents, model training, and inference, tailored to document question-answering. |
Arctic-TILT. Business Document Understanding at Sub-Billion Scale (2025.acl-industry)
Copied to clipboard
Łukasz Borchmann, Michał Pietruszka, Wojciech Jaśkowski, Dawid Jurkiewicz, Piotr Halama, Paweł Józiak, Łukasz Garncarek, Paweł Liskowski, Karolina Szyndler, Andrzej Gretkowski, Julita Ołtusek, Gabriela Nowakowska, Artur Zawłocki, Łukasz Duhr, Paweł Dyda, Michał Turski
| Challenge: | General-purpose LLMs and their multimodal counterparts provide a crucial advantage in process automation. |
| Approach: | They propose a model that can be finetuned and deployed on a single 24GB GPU . it provides reliable confidence scores and quick inferences for processing files in large-scale or time-sensitive environments. |
| Outcome: | The proposed model achieves state-of-the-art results on seven diverse benchmarks and provides reliable confidence scores and quick inferences. |
GovScape: A Public Multimodal Search System for 70 Million Pages of Government PDFs (2026.acl-demo)
Copied to clipboard
Ying-Hsiang Huang, Claire Gong, Shreya Shaji, Alison R Yan, Leslie Harka, Albert Du, Anjali Shubha Gopal, Samuel J Klein, Shannon Zejiang Shen, Mark E. Phillips, Trevor Owens, Kyle Deeds, Benjamin Charles Germain Lee
| Challenge: | Efforts over the past three decades have produced web archives containing billions of webpage snapshots and petabytes of data. |
| Approach: | They propose a public search system that supports multimodal searches across 10,015,993 federal government PDFs from the 2020 End of Term crawl. |
| Outcome: | The proposed system supports multimodal searches across 10,015,993 federal government PDFs from the 2020 End of Term crawl (70,958,487 total PDF pages) significant compute cost for GovScape’s pre-processing pipeline for 10 million PDFs was approximately 1,500, equivalent to 47,000 PDF pages per dollar spent on compute. |
PAWLS: PDF Annotation With Labels and Structure (2021.acl-demo)
Copied to clipboard
| Challenge: | Existing tools for annotation of PDFs are limited to a web browser, allowing users to extract semantically meaningful regions from PDFs. |
| Approach: | They propose an annotation tool specifically designed for Adobe’s Portable Document Format (PDF) PAWLS supports span-based textual annotation, N-ary relations and freeform, non-textual bounding boxes. |
| Outcome: | The proposed tool supports span-based textual annotation, N-ary relations and freeform, non-textual bounding boxes. |
XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser (2025.coling-main)
Copied to clipboard
Xianfu Cheng, Hang Zhang, Jian Yang, Xiang Li, Weixiao Zhou, Fei Liu, Kui Wu, Xiangyuan Guan, Tao Sun, Xianjie Wu, Tongliang Li, Zhoujun Li
| Challenge: | Document AI parsing semi-structured image form is a key information extraction task. |
| Approach: | They propose a multimodal and multilingual semi-structured FORM PARSER which integrates SER and relation extraction into a unified framework. |
| Outcome: | The proposed framework achieves up to 1.79% improvement on RE tasks in multilingual and zero-shot settings. |
FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosure (2026.acl-demo)
Copied to clipboard
Fan Zhang, Mingzi Song, Rania Elbadry, Yankai Chen, Shaobo Wang, Yixi Zhou, Xunwen Zheng, Yueru He, Yuyang Dai, Georgi Nenkov Georgiev, Ayesha Gull, Muhammad Usman Safder, Fan Wu, Liyuan Meng, Fengxian Ji, Junning Zhao, Xueqing Peng, Jimin Huang, YU Chen, Xue Liu, Preslav Nakov, Zhuohan Xie
| Challenge: | FinReporting is an agentic workflow for localized cross-jurisdiction financial reporting . existing approaches assume a single-market setting and overlook structural differences across jurisdictions . |
| Approach: | They propose a workflow that decomposes financial reporting into auditable stages . they use Large Language Models to extract and summarize corporate disclosures . |
| Outcome: | The proposed system decomposes reporting into auditable stages . it improves consistency and reliability under heterogeneous reporting regimes. |
Harnessing PDF Data for Improving Japanese Large Multimodal Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) have demonstrated strong performance in English, but their effectiveness in Japanese remains limited due to the lack of high-quality training data. |
| Approach: | They propose a pipeline that leverages pretrained models to extract image-text pairs from PDFs . they use layout analysis, OCR, and vision-language pairing to enrich the training data . |
| Outcome: | The proposed pipeline extracts image-text pairs from Japanese PDFs, eliminating manual annotations. |
CoFiF Plus: A French Financial Narrative Summarisation Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora for financial narrative summarisation exists in English but there is a significant lack of financial text resources in the French language. |
| Approach: | They propose to use natural language processing to analyse financial documents to find the best summarisation methods. |
| Outcome: | The proposed dataset is the first to provide a comprehensive set of financial text written in French. |
PDFAnno: a Web-based Linguistic Annotation Tool for PDF Documents (L18-1)
Copied to clipboard
| Challenge: | Currently, linguistic annotation tools for PDF documents focus on plain-text documents. |
| Approach: | They propose a web-based linguistic annotation tool for PDF documents . it offers functions for various types of linguistic annotations directly on PDF . |
| Outcome: | The proposed tool can annotate on PDF documents with named entity, dependency relation, and coreference chain. |
MathAlign: Linking Formula Identifiers to their Contextual Natural Language Descriptions (2020.lrec-1)
Copied to clipboard
Maria Alexeeva, Rebecca Sharp, Marco A. Valenzuela-Escárcega, Jennifer Kadowaki, Adarsh Pyarelal, Clayton Morrison
| Challenge: | Existing approaches to extract mathematical concepts and their descriptions are useful for a variety of tasks, including math information retrieval and accessibility efforts to make scientific documents available to the visually impaired. |
| Approach: | They propose a rule-based approach which extracts LaTeX representations of formula identifiers and links them to their in-text descriptions, given only the original PDF and the location of the formula of interest. |
| Outcome: | The proposed approach extracts LaTeX representations of formula identifiers and links them to their in-text descriptions, given only the original PDF and the location of the formula of interest. |
PDFdigest: an Adaptable Layout-Aware PDF-to-XML Textual Content Extractor for Scientific Articles (L18-1)
Copied to clipboard
| Challenge: | Existing tools to extract structured textual content from PDFs are essential to enable scientific text mining. |
| Approach: | They propose a PDF-to-XML textual content extraction tool that extracts structured textual contents from scientific articles in PDF format. |
| Outcome: | The proposed tool extracts structured textual content from scientific articles in PDF format while preserving both the textual contents and layout details of the input PDF document. |
TDDC: Timely Disclosure Documents Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | TDDC was prepared by manually aligning the sentences from past Japanese and English timely disclosure documents . tens of thousands of original Japanese documents are disclosed every year, but the availability of English disclosure documents is limited. |
| Approach: | They describe the details of the Timely Disclosure Documents Corpus (TDDC) TDDC was prepared by manually aligning the sentences from past Japanese and English timely disclosure documents . |
| Outcome: | The timely disclosure documents corpus (TDDC) was created by aligning sentences from past documents in Japanese and English. |
Musical Score Understanding Benchmark: Evaluating Large Language Models’ Comprehension of Complete Musical Scores (2026.acl-long)
Copied to clipboard
Congren Dai, Yue Yang, Krinos Li, Huichi Zhou, Shijie Liang, Zhang Bo, Enyang Liu, Ge Jin, Hongran An, Haosen Zhang, Peiyuan Jing, KinHei Lee, Zhenxuan Zhang, Xiaobing Li, Maosong Sun
| Challenge: | Existing benchmarks for musical score understanding are narrow in scope, focusing on isolated fragments, short excerpts, or multiple-choice formulations, rather than supporting holistic reasoning over entire scores. |
| Approach: | They propose a benchmark for score-level musical understanding across textual and visual modalities. |
| Outcome: | The musical score understanding benchmark contains 1,800 question-answer pairs from works by Bach, Beethoven, Chopin, Debussy, and others. |
Dataset Construction for Scientific-Document Writing Support by Extracting Related Work Section and Citations from PDF Papers (2022.lrec-1)
Copied to clipboard
| Challenge: | To augment datasets used for scientific-document writing support research, we extract texts from “Related Work” sections and citation information in PDF-formatted papers published in English. |
| Approach: | They propose to extract text from “Related Work” sections and citation information from PDF-formatted papers published in English. |
| Outcome: | The proposed dataset is based on a previously constructed dataset using only Tex papers and is compared with the existing one. |
PDF-to-Tree: Parsing PDF Text Blocks into a Tree (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies try to extract one universal reading order for PDF files, however, some applications, like Retrieval Augmented Generation, require breaking long articles into sections and subsections for better indexing. |
| Approach: | They propose a new task and dataset, PDF-to-Tree, which organizes the text blocks of a PDF into a tree structure. |
| Outcome: | The proposed parser achieves 93.93% accuracy, surpassing baseline methods by 6.72%. |
JDocQA: Japanese Document Question Answering Dataset for Generative Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites. |
| Approach: | They propose a large-scale document-based QA dataset that requires both visual and textual information to answer questions. |
| Outcome: | The proposed dataset incorporates multiple categories of questions and unanswerable questions from the document for realistic question-answering applications. |
MedQA-SWE - a Clinical Question & Answer Dataset for Swedish (2024.lrec-main)
Copied to clipboard
| Challenge: | MedQA-SWE is a clinical question & answering dataset in Swedish . it was created from exams aimed at evaluating doctors’ clinical understanding and decision making . |
| Approach: | They propose to create a multiple choice, clinical question & answering (Q&A) dataset in Swedish consisting of 3,180 questions. |
| Outcome: | The proposed dataset includes 3,180 questions and is the first open-source clinical Q&A dataset in Swedish. |