Papers by Colin Lockard

8 papers
OpenKI: Integrating Open Information Extraction and Knowledge Bases with Relation Inference (N19-1)

Copied to clipboard

Challenge: Existing methods for knowledge extraction and alignment are limited in quality and performance.
Approach: They propose to integrate OpenIE extractions in the form of (subject, predicate, object) triples with Knowledge Bases (KB)
Outcome: The proposed method improves state-of-the-art for OpenIE extractions and boosts performance on OpenIE from semi-structured data.
Train a Unified Multimodal Data Quality Classifier with Synthetic Data (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal Large Language Models are pre-trained on image-text caption data and interleaved document data.
Approach: They propose to train an efficient MLLM as a Unified Mulitmodal Data Quality Classifier to filter image-text caption and interleaved data.
Outcome: The proposed method enables efficient creation of sample-score pairs for caption and interleaved data to train UniFilter.
OpenCeres: When Open Information Extraction Meets the Semi-Structured Web (N19-1)

Copied to clipboard

Challenge: Open Information Extraction (OpenIE) is a problem of extracting triples from natural language text whose predicate relations are not aligned to any pre-defined ontology.
Approach: They propose an open-source method to extract triples from semi-structured websites . they use a semi-supervised label propagation technique to create training data for relations .
Outcome: The proposed method extracts over 2 million triples from 31 websites in the movie vertical.
ZeroShotCeres: Zero-Shot Relation Extraction from Semi-Structured Webpages (2020.acl-main)

Copied to clipboard

Challenge: Existing work on information extraction from semi-structured websites has relied on manual data annotation and learning a model specific to a given template.
Approach: They propose a solution for “zero-shot” open-domain relation extraction from webpages with previously unseen templates using a graph neural network-based approach.
Outcome: The proposed model provides a 31% gain over baseline for zero-shot extraction in a new subject vertical.
Extracting Shopping Interest-Related Product Types from the Web (2023.findings-acl)

Copied to clipboard

Challenge: Existing e-commerce products are limited in their ability to assist customers in interest-oriented shopping.
Approach: They propose to extract PTs from Web pages containing hand-crafted PT recommendations for SIs . they propose to use tree-transformer encoders for node classification to improve inter-node dependency modeling .
Outcome: The proposed model outperforms the best baseline model by 2.37 F1 points on a WebPT dataset.
PLAtE: A Large-scale Dataset for List Page Web Extraction (2023.acl-industry)

Copied to clipboard

Challenge: Existing methods for web extraction are limited by the limited number of available large-scale datasets.
Approach: They introduce a dataset that focuses on shopping data and a list page web extraction task.
Outcome: The proposed dataset is the first large-scale list page web extraction dataset . it contains 52,898 items and 156,014 attributes, making it the first dataset based on this task .
Multi-modal Information Extraction from Text, Semi-structured, and Tabular Data on the Web (2020.acl-tutorials)

Copied to clipboard

Challenge: a tutorial explores the commonalities in the challenges and solutions developed to address information extraction from the World Wide Web.
Approach: This tutorial examines methods for extracting information from the World Wide Web . it explores the commonalities in the challenges and solutions developed to address these different forms of text .
Outcome: This paper examines the commonalities in the challenges and solutions developed to address the World Wide Web.
Semi-Supervised Event Extraction with Paraphrase Clusters (N18-2)

Copied to clipboard

Challenge: Existing event extraction systems are limited in their accuracy due to the lack of available training data.
Approach: They propose a method for self-training event extraction systems by bootstrapping additional training data.
Outcome: The proposed method improves on ACE 2005 and TAC-KBP 2015 datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations