Challenge: Documents may feature zero or more instances of a template of any given type, and the task of template extraction entails identifying the templates in a document and extracting each template’s slot values.
Approach: They propose to use iterative extraction to extract complex relations, i.e., N-tuples representing a mapping from named slots to spans of text within a document.
Outcome: The proposed model leads to state-of-the-art results on two established benchmarks and a strong baseline on the new BETTER Granular task.

Similar Papers

Representation Learning for Information Extraction from Form-like Documents (2020.acl-main)

Copied to clipboard

Challenge: Form-like documents like invoices, purchase orders, tax forms and insurance quotes are common in day-to-day business workflows, but current techniques for processing them largely still employ manual effort or brittle and error-prone heuristics for extraction.
Approach: They propose an extraction system that uses knowledge of the types of the target fields to generate extraction candidates and a neural network architecture that learns a dense representation of each candidate based on neighboring words in the document.
Outcome: The proposed system generates extraction candidates based on neighboring words in the document and is interpretable, as shown using loss cases.
SciREX: A Challenge Dataset for Document-Level Information Extraction (2020.acl-main)

Copied to clipboard

Challenge: Conventional datasets and methods for information extraction focus on within-sentence relations from general Newswire text.
Approach: They propose a document-level IE dataset that integrates automatic and human annotations to annotate entities and document- level N-ary relation identification from scientific articles.
Outcome: The proposed dataset extends state-of-the-art IE models to document-level IE.
Document-level Entity-based Extraction as Template Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Document-level entity-based extraction (EE) tasks extract entity-centric information from unstructured text across multiple sentences.
Approach: They propose a generative framework for two document-level EE tasks: role-filler entity extraction (RE) and relation extraction ( RE).
Outcome: The proposed framework captures cross-entity dependencies and avoids exponential computation complexity of identifying N-ary relations.
Automatic Error Analysis for Document-level Information Extraction (2022.acl-long)

Copied to clipboard

Challenge: Document-level information extraction (IE) tasks have been revisited in earnest . evaluation of the approaches has been limited in a number of dimensions .
Approach: They propose a transformation-based framework for automating error analysis in document-level event and (N-ary) relation extraction.
Outcome: The proposed framework compares two state-of-the-art document-level template-filling approaches on datasets from three domains and four systems from the MUC-4 evaluation.
Towards Better Document-level Relation Extraction via Iterative Inference (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods only consider feature information of entity pairs, but our model exploits both feature information and previous predictions of entity pair.
Approach: They propose a document-level relation extraction model with iterative inference to extract relations between entities from raw texts.
Outcome: The proposed model outperforms existing methods on three commonly-used datasets.
ITER: Iterative Transformer-based Entity Recognition and Relation Extraction (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in NLP generate structured information in an autoregressive manner, causing low throughput . authors propose an efficient encoder-based relation extraction model that performs the task in three parallelizable steps.
Approach: They propose an efficient encoder-based relation extraction model that performs the task in three parallelizable steps.
Outcome: The proposed model achieves state-of-the-art on two datasets and is faster than existing models.
Set Learning for Generative Information Extraction (2023.emnlp-main)

Copied to clipboard

Challenge: Recent efforts to employ sequence-to-sequence models to solve IE tasks have been focused on a single problem: structured objects are an unordered set, resulting in a potential order bias.
Approach: They propose a sequence-to-sequence (Seq2Sequen) model that considers multiple permutations of structured objects to optimize set probability approximately.
Outcome: The proposed model improves existing frameworks on vast tasks and datasets.
Easy-to-Hard Learning for Information Extraction (2023.findings-acl)

Copied to clipboard

Challenge: Existing models for information extraction (IE) use a one-stage learning strategy to extract the target structure from unstructured text data.
Approach: They propose a unified easy-to-hard learning framework that mimics the human learning process by breaking down the learning process into multiple stages.
Outcome: The proposed framework enables the model to acquire general IE task knowledge and improve its generalization ability on 13 out of 17 datasets.
Enhanced Language Representation with Label Knowledge for Span Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to extract text spans from plain text do not fully exploit label knowledge.
Approach: They propose a model to integrate label knowledge into text representations by encoding texts and annotations independently and then integrating label knowledge with an elaborate-designed semantics fusion module.
Outcome: The proposed model achieves state-of-the-art performance on four benchmarks and reduces training time and inference time by 76% and 77% on average compared with the existing paradigm.
Iterative Document Representation Learning Towards Summarization with Polishing (D18-1)

Copied to clipboard

Challenge: Existing summarization methods read through document only once to generate a document representation, resulting in a sub-optimal representation.
Approach: They propose an iterative model for supervised extractive text summarization which polishes the document representation on many passes through the document.
Outcome: The proposed model outperforms state-of-the-art extractive systems on CNN/DailyMail and DUC2002 datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations