Papers by Feiyu Gao

6 papers
Graph Convolution for Multimodal Information Extraction from Visually Rich Documents (N19-2)

Copied to clipboard

Challenge: Visually rich documents (VRDs) present information in the form of both text and vision.
Approach: They propose a graph convolution based model to combine textual and visual information presented in VRDs.
Outcome: The proposed model outperforms existing models on two real-world datasets.
Intelligent Document Parsing: Towards End-to-end Document Parsing via Decoupled Content Parsing and Layout Grounding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods fragment document parsing into pipeline of separated subtasks, resulting in incomplete semantics and error propagation.
Approach: They propose an end-to-end document parsing framework that leverages vision-language priors of MLLMs.
Outcome: The proposed method surpasses existing methods significantly in document parsing . it leverages the vision-language priors of MLLMs to decouple parse and layout grounding based on visual information.
GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models that use plain HTMLs do not include crucial visual information in the rendered web.
Approach: They propose a Gestalt Enhanced Markup Language Model for hosting visual information without visual input.
Outcome: The proposed model can handle multiple downstream tasks without visual input.
Taming Text-to-Image Synthesis for Novices: User-centric Prompt Generation via Multi-turn Guidance (2025.emnlp-main)

Copied to clipboard

Challenge: Existing solutions for text-to-image synthesis are sensitive on textual prompts, posing a challenge for novice users.
Approach: They propose a dialogue-based TIS prompt generation model that emphasizes user experience for novice users.
Outcome: The proposed model emphasizes user experience for novice users . it improves user-centricity score while maintaining a competitive quality of synthesized images.
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding due to different types of annotation noise in training.
Approach: They propose a method to reduce C&P knowledge conflicts across all tested MLLMs . they propose to use annotation noise to train models to understand document content .
Outcome: The proposed method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks.
DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for document hierarchy parsing are limited due to the small scale and inconsistency of datasets.
Approach: They propose a document hierarchy parsing dataset to compensate for the data scarcity problem and propose 'dHP' framework to grasp fine-grained text content and coarse-grounded pattern at layout element level.
Outcome: The proposed framework grasps both fine-grained text content and coarse-grounded pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling multi-page and multi-level challenges.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations