Challenge: a new open-source layout-aware IE test suite is available for download at https://github.com/gayecolakoglu/layIE-LLM.
Approach: They propose an open-source layout-aware IE test suite that provides a layout-based IE pipeline.
Outcome: The proposed method achieves 13.3–37.5 F1 points more than a baseline configuration using the same LLM.

Similar Papers

LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to enhance document comprehension require fine-tuning for each task and dataset, and are expensive to train and operate.
Approach: They propose a more flexible document analysis method that integrates visual-rich document understanding with large-scale language models (LLMs) by leveraging existing research in document image understanding and LLMs’ superior language understanding capabilities, the proposed model performs an understanding of document images in a single model.
Outcome: The proposed model improves on the baseline model in document image understanding tasks.
Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual Segments (2025.acl-long)

Copied to clipboard

Challenge: Specialized non-LLM NLP-based solutions lack reasoning and are not able to infer values not explicitly present in documents.
Approach: They propose a novel LLM-based approach that organizes VRDs into localized semantic textual segments called semantic blocks.
Outcome: The proposed approach outperforms the state-of-the-art on public VRD benchmarks by 1-3% in F1 scores and is resilient to document formats previously not encountered.
LMDX: Language Model-based Document Information Extraction and Localization (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models have revolutionized Natural Language Processing but their application in extracting information from visually rich documents has not been successful.
Approach: They propose a language model-based document information extraction and localization methodology to reframe the document information extract task for a LLM.
Outcome: The proposed method enables extraction of singular, repeated, and hierarchical entities with and without training data.
Massively Multilingual Instruction-Following Information Extraction (2025.findings-acl)

Copied to clipboard

Challenge: Past literature on information extraction (IE) has focused on a few high-resource languages, hindering their applications on multilingual corpora.
Approach: They propose a collection of data that unifies and standardizes instruction-following multilingual IE and introduce a structure-aware metric that captures partially matched spans.
Outcome: The proposed framework standardizes and unifies 215 manually annotated datasets, covering 96 typologically diverse languages from 18 language families.
A Bounding Box is Worth One Token - Interleaving Layout and Text in a Large Language Model for Document Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for integrating spatial layouts with text have limitations . existing methods produce overly long text sequences or lack autoregressive traits of LLMs .
Approach: They introduce Interleaving Layout and Text in a Large Language Model (LayTextLLM) they use OCR-derived text and spatial layouts to integrate with LLMs for document understanding .
Outcome: The proposed model shows an increase in performance in KIE and VQA tasks.
Large Language Model Is Not a Good Few-shot Information Extractor, but a Good Reranker for Hard Samples! (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have made remarkable strides in various tasks, but whether they are competitive few-shot solvers remains an open question.
Approach: They propose an adaptive filter-then-rerank paradigm to combine the strengths of LLMs and SLMs.
Outcome: The proposed system achieves promising improvements on various IE tasks with acceptable time and cost investment.
Intelligent Document Parsing: Towards End-to-end Document Parsing via Decoupled Content Parsing and Layout Grounding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods fragment document parsing into pipeline of separated subtasks, resulting in incomplete semantics and error propagation.
Approach: They propose an end-to-end document parsing framework that leverages vision-language priors of MLLMs.
Outcome: The proposed method surpasses existing methods significantly in document parsing . it leverages the vision-language priors of MLLMs to decouple parse and layout grounding based on visual information.
DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding (2024.acl-long)

Copied to clipboard

Challenge: Documents with rich layouts are a significant portion of enterprise corpora and document AI is still a challenge.
Approach: They propose a lightweight extension to traditional large language models for reasoning over visual documents that takes into account both textual semantics and spatial layout.
Outcome: The proposed model outperforms existing large language models on 14 out of 16 datasets and generalizes well to 4 out of 5 previously unseen datasets.
A Simple but Effective Approach to Improve Structured Language Model Output for Information Extraction (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models have impressive abilities in generating unstructured natural language . performance inconsistent when tasked with producing text that adheres to structured formats .
Approach: They propose a method to generate unstructured natural language using intermediate responses . they use the intermediate responses to organize the output into the desired structure .
Outcome: The proposed method improves performance on NER and RE tasks with minimal effort.
ADELIE: Aligning Large Language Models on Information Extraction (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) struggle to follow complex instructions of IE tasks due to not being aligned with humans.
Approach: They propose an aligned large language moDEL that effectively solves various IE tasks including closed IE, open IE and on-demand IE.
Outcome: The proposed model achieves state-of-the-art (SoTA) performance among open-source models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations