Chargrid: Towards Understanding 2D Documents (D18-1)

Copied to clipboard

Challenge: a novel type of text representation preserves the 2D layout of a document . a computer vision algorithm that extracts text from image is not optimal for understanding semantics .
Approach: They propose a new type of text representation that preserves the 2D layout of a document . they demonstrate that it significantly outperforms approaches based on sequential text or document images .
Outcome: The proposed approach outperforms approaches based on sequential text or document images on an information extraction task from invoices.

Similar Papers

Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis (2022.coling-1)

Copied to clipboard

Challenge: Document Layout Analysis tasks rely on visual cues to understand documents . traditional deep learning-based methods fail to recognize the layout and components of unstructured documents based on the document structure and the boundaries of each layout region.
Approach: They propose a way to harmonize and integrate heterogeneous aspects for Document Layout Analysis by using graph convolutional networks to enhance each aspect of features.
Outcome: The proposed task is based on three widely used datasets: PubLayNet, FUNSD, and DocBank.
Representation Learning for Information Extraction from Form-like Documents (2020.acl-main)

Copied to clipboard

Challenge: Form-like documents like invoices, purchase orders, tax forms and insurance quotes are common in day-to-day business workflows, but current techniques for processing them largely still employ manual effort or brittle and error-prone heuristics for extraction.
Approach: They propose an extraction system that uses knowledge of the types of the target fields to generate extraction candidates and a neural network architecture that learns a dense representation of each candidate based on neighboring words in the document.
Outcome: The proposed system generates extraction candidates based on neighboring words in the document and is interpretable, as shown using loss cases.
Beyond Sequences: Two-dimensional Representation and Dependency Encoding for Code Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing code generation approaches represent code as a linear sequence of tokens, but positional encodings compromise generalization . explicit positional encoders sacrifice permutation invariance, imposes a strict order on the input sequence .
Approach: They propose to represent code snippets as two-dimensional entities with explicit encodings . they propose to use dictionary learning to perform semantic matching between code lines .
Outcome: The proposed model captures the hierarchical and spatial structure of code, especially the dependencies between code lines.
Intelligent Document Parsing: Towards End-to-end Document Parsing via Decoupled Content Parsing and Layout Grounding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods fragment document parsing into pipeline of separated subtasks, resulting in incomplete semantics and error propagation.
Approach: They propose an end-to-end document parsing framework that leverages vision-language priors of MLLMs.
Outcome: The proposed method surpasses existing methods significantly in document parsing . it leverages the vision-language priors of MLLMs to decouple parse and layout grounding based on visual information.
ConvTextTM: An Explainable Convolutional Tsetlin Machine Framework for Text Classification (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have reshaped the industry . complexity of such models makes them a “black box” and can cause ethical concerns .
Approach: They propose a convolutional TM architecture that breaks down text into a sequence of fragments . they propose to use a tokenization scheme to bind the tokens to the text fragments.
Outcome: The proposed architecture improves on a set of text fragments and eliminates the need for a corpus-specific vocabulary.
STable: Table Generation Framework for Encoder-Decoder Models (2024.eacl-long)

Copied to clipboard

Challenge: Existing approaches to infer text-to-table neural models are limited to raw text, but the proposed framework is capable of unifying a variety of problems involving natural language.
Approach: They propose a framework for text-to-table neural models that utilizes a generalized sequential method that comprehends information from all cells in the table.
Outcome: The proposed framework outperforms previous approaches on several challenging datasets and outperformed existing models by up to 15%.
DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for document hierarchy parsing are limited due to the small scale and inconsistency of datasets.
Approach: They propose a document hierarchy parsing dataset to compensate for the data scarcity problem and propose 'dHP' framework to grasp fine-grained text content and coarse-grounded pattern at layout element level.
Outcome: The proposed framework grasps both fine-grained text content and coarse-grounded pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling multi-page and multi-level challenges.
DocPolarBERT: A Pre-trained Model for Document Understanding with Relative Polar Coordinate Encoding of Layout Structures (2026.eacl-long)

Copied to clipboard

Challenge: Existing models that take text block positions into account are not efficient for document understanding.
Approach: They propose a layout-aware BERT model that takes into account text block positions in relative polar coordinate system rather than the Cartesian one.
Outcome: The proposed model eliminates the need for absolute positional embeddings on a dataset more than six times smaller than the widely used IIT-CDIP corpus.
A Spatial Model for Extracting and Visualizing Latent Discourse Structure in Text (P18-1)

Copied to clipboard

Challenge: Using sequences of sentences, we show that learning long-range latent discourse structure from large corpora can be useful for machine learning.
Approach: They propose a probabilistic model of documents as sequences of sentences with a 2- or 3-D spatial grid and embed sentences into a grid.
Outcome: The proposed model outperforms or is competitive with state-of-the-art generative approaches on tasks such as predicting the outcome of a story, and sentence ordering.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations