Challenge: Information extraction (IE) is the process of automatically extracting structural information from unstructured or semi-structured data.
Approach: This tutorial will provide an introduction to recent advances in IE by answering several important research questions.
Outcome: The tutorial will address several important research questions and outline directions for further investigation.

Similar Papers

Easy-to-Hard Learning for Information Extraction (2023.findings-acl)

Copied to clipboard

Challenge: Existing models for information extraction (IE) use a one-stage learning strategy to extract the target structure from unstructured text data.
Approach: They propose a unified easy-to-hard learning framework that mimics the human learning process by breaking down the learning process into multiple stages.
Outcome: The proposed framework enables the model to acquire general IE task knowledge and improve its generalization ability on 13 out of 17 datasets.
Scalable Construction and Reasoning of Massive Knowledge Bases (N18-6)

Copied to clipboard

Challenge: Existing knowledge mining systems assume abundant human annotations for training high quality machine learning models, which is impractical when trying to deploy IE systems to a broad range of domains, settings and languages.
Approach: They introduce how to extract structured facts from text corpora to construct knowledge bases.
Outcome: The proposed methods are weakly-supervised and domain-independent for knowledge base construction across various domains.
A Survey of Generative Information Extraction (2025.coling-main)

Copied to clipboard

Challenge: Information Extraction (IE) is a popular and fundamental task in natural language processing.
Approach: They first review generative information extraction methods based on pre-trained language models and large language models focusing on their adaptation and generalization capabilities.
Outcome: The proposed methods are based on pre-trained language models and large language models, and emphasize the importance of model collaboration.
A Survey on Open Information Extraction (C18-1)

Copied to clipboard

Challenge: Existing approaches to open information extraction (Open IE) focus on narrow, well-defined requests over a predefined set of target relations on small, homogeneous corpora.
Approach: They propose to use unsupervised methods to extract all types of relations found in text . they propose to implement a system that can be automated to detect possible relations .
Outcome: The proposed approaches have been compared with existing methods and are based on the results of a literature review.
Multi-modal Information Extraction from Text, Semi-structured, and Tabular Data on the Web (2020.acl-tutorials)

Copied to clipboard

Challenge: a tutorial explores the commonalities in the challenges and solutions developed to address information extraction from the World Wide Web.
Approach: This tutorial examines methods for extracting information from the World Wide Web . it explores the commonalities in the challenges and solutions developed to address these different forms of text .
Outcome: This paper examines the commonalities in the challenges and solutions developed to address the World Wide Web.
ADAPTIVE IE: Investigating the Complementarity of Human-AI Collaboration to Adaptively Extract Information on-the-fly (2025.coling-main)

Copied to clipboard

Challenge: Existing IE systems are either fully supervised, requiring expensive human annotations, or fully unsupervised, extracting information that often do not cater to user’s needs.
Approach: They propose a framework that uses human-in-the-loop refinement to adapt to changing user questions.
Outcome: The proposed framework is domain-agnostic, responsive, efficient for helping users access useful information while quickly reorganizing information in response to evolving information needs.
Latent Structure Models for Natural Language Processing (P19-4)

Copied to clipboard

Challenge: Latent structure models are a powerful tool for compositional data modeling and pipelines.
Approach: This tutorial will cover recent advances in discrete latent structure models . it will discuss their motivation, potential, and limitations .
Outcome: This tutorial will cover recent advances in discrete latent structure models . it will discuss their motivation, potential, and limitations .
Unsupervised Natural Language Parsing (Introductory Tutorial) (2021.eacl-tutorials)

Copied to clipboard

Challenge: Unsupervised parsing learns a syntactic parser from training sentences without parse tree annotations.
Approach: This tutorial will introduce what unsupervised parsing does and how it can be useful for and beyond syntactic parse.
Outcome: This paper will provide an overview of major approaches to unsupervised parsing and analyze their strengths and weaknesses.
IEPile: Unearthing Large Scale Schema-Conditioned Information Extraction Corpus (2024.acl-short)

Copied to clipboard

Challenge: Large Language Models exhibit a significant performance gap in Information Extraction (IE) high-quality instruction data is the vital key for enhancing LLMs' specific capabilities .
Approach: They propose a bilingual (English and Chinese) IE instruction corpus that contains 0.32B tokens.
Outcome: The proposed model improves the performance of LLMs for IE with zero-shot generalization.
Cost-effective End-to-end Information Extraction for Semi-structured Document Images (2021.emnlp-main)

Copied to clipboard

Challenge: a real-world information extraction system for semi-structured document images often involves a long pipeline of multiple modules, which can lead to unstable performance if not designed carefully.
Approach: They propose to use a sequence generation task to build an end-to-end IE system . they propose to combine three manually engineered modules with one data-driven module .
Outcome: The proposed system can be easily replaced and deployed in large-scale production.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations