Challenge: a challenge in scientific literature mining is the difficulty of extracting high-quality text from formatted PDFs.
Approach: They propose a method to visually segment key regions of scientific articles using object detection augmented with contextual features.
Outcome: The proposed method improves the accuracy of the proposed method and the speed of the dataset.

Similar Papers

VILA: Improving Structured Content Extraction from Scientific PDFs Using Visual Layout Groups (2022.tacl-1)

Copied to clipboard

Challenge: Recent work has improved extraction accuracy by incorporating elementary layout information, for example, each token’s 2D position on the page, into language model pretraining.
Approach: They propose a method that explicitly models VIsual LAyout (VILA) groups, that is, text lines or text blocks, to further improve extraction accuracy.
Outcome: The proposed methods show that inserting special tokens denoting layout group boundaries can lead to a 1.9% Macro F1 improvement in token classification.
DocBank: A Benchmark Dataset for Document Layout Analysis (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches for document layout analysis are based on rule-based or machine learning methods that ignore textual information.
Approach: They present a benchmark document layout analysis dataset using a computer vision model . they build strong baselines and manually split train/dev/test sets for evaluation .
Outcome: The proposed model trains on DocBank accurately recognize layout information for a variety of documents.
Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis (2022.coling-1)

Copied to clipboard

Challenge: Document Layout Analysis tasks rely on visual cues to understand documents . traditional deep learning-based methods fail to recognize the layout and components of unstructured documents based on the document structure and the boundaries of each layout region.
Approach: They propose a way to harmonize and integrate heterogeneous aspects for Document Layout Analysis by using graph convolutional networks to enhance each aspect of features.
Outcome: The proposed task is based on three widely used datasets: PubLayNet, FUNSD, and DocBank.
Dataset Construction for Scientific-Document Writing Support by Extracting Related Work Section and Citations from PDF Papers (2022.lrec-1)

Copied to clipboard

Challenge: To augment datasets used for scientific-document writing support research, we extract texts from “Related Work” sections and citation information in PDF-formatted papers published in English.
Approach: They propose to extract text from “Related Work” sections and citation information from PDF-formatted papers published in English.
Outcome: The proposed dataset is based on a previously constructed dataset using only Tex papers and is compared with the existing one.
CiteWorth: Cite-Worthiness Detection for Improved Scientific Document Understanding (2021.findings-acl)

Copied to clipboard

Challenge: Scientific document understanding is challenging due to the highly domain specific nature of scientific language.
Approach: They propose a large, contextualized, rigorously cleaned labelled dataset for cite-worthiness detection built from extracted scientific documents.
Outcome: The proposed model improves on a paragraphlevel contextualized sentence labelling model based on Longformer . the model shows a 5 F1 point improvement over SciBERT which considers only individual sentences .
LayoutReader: Pre-training of Text and Layout for Reading Order Detection (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for reading order detection are too laborious to annotate large datasets.
Approach: They propose to use a large-scale dataset to annotate reading order information for document images . they use XML metadata to capture the reading order of WORD documents .
Outcome: The proposed model performs almost perfectly in reading order detection and improves both open-source and commercial OCR engines in ordering text lines in their results.
PLOD: An Abbreviation Detection Dataset for Scientific Documents (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for abbreviation detection and extraction are limited.
Approach: They propose to use a large-scale dataset for abbreviation detection and extraction that contains 160k+ segments automatically annotated with abbrevian and long forms.
Outcome: The proposed dataset has an F1 score of 0.92 for abbreviations and 0.89 for detecting their corresponding long forms.
DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for document hierarchy parsing are limited due to the small scale and inconsistency of datasets.
Approach: They propose a document hierarchy parsing dataset to compensate for the data scarcity problem and propose 'dHP' framework to grasp fine-grained text content and coarse-grounded pattern at layout element level.
Outcome: The proposed framework grasps both fine-grained text content and coarse-grounded pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling multi-page and multi-level challenges.
Dynamic Context Extraction for Citation Classification (2022.aacl-main)

Copied to clipboard

Challenge: Prior studies have focused on the application of fixed-size contiguous citation contexts or manually curated citation contextual contexts.
Approach: They propose an automated unsupervised approach for the selection of a dynamic-size and potentially non-contiguous citation context based on transformer-based document representations and embedding similarities.
Outcome: The proposed model improves on the domain-specific and multi-disciplinary datasets, irrespective of the dataset's domain.
LayoutPointer: A Spatial-Context Adaptive Pointer Network for Visual Information Extraction (2024.naacl-long)

Copied to clipboard

Challenge: Existing models inadequately utilize spatial information of entities, causing incorrectly linking spatially distant entities.
Approach: They propose a Spatial-Context Adaptive Pointer Network to restore semantic order among entities . they propose XFUND-based tail-to-head pointer to restore the semantic order .
Outcome: The proposed method outperforms existing state-of-the-art methods in F1 scores for RE tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations