Papers by Colin Lockard
OpenKI: Integrating Open Information Extraction and Knowledge Bases with Relation Inference (N19-1)
Copied to clipboard
| Challenge: | Existing methods for knowledge extraction and alignment are limited in quality and performance. |
| Approach: | They propose to integrate OpenIE extractions in the form of (subject, predicate, object) triples with Knowledge Bases (KB) |
| Outcome: | The proposed method improves state-of-the-art for OpenIE extractions and boosts performance on OpenIE from semi-structured data. |
Train a Unified Multimodal Data Quality Classifier with Synthetic Data (2025.findings-emnlp)
Copied to clipboard
Weizhi Wang, Rongmei Lin, Shiyang Li, Colin Lockard, Ritesh Sarkhel, Sanket Lokegaonkar, Jingbo Shang, Xifeng Yan, Nasser Zalmout, Xian Li
| Challenge: | Multimodal Large Language Models are pre-trained on image-text caption data and interleaved document data. |
| Approach: | They propose to train an efficient MLLM as a Unified Mulitmodal Data Quality Classifier to filter image-text caption and interleaved data. |
| Outcome: | The proposed method enables efficient creation of sample-score pairs for caption and interleaved data to train UniFilter. |
OpenCeres: When Open Information Extraction Meets the Semi-Structured Web (N19-1)
Copied to clipboard
| Challenge: | Open Information Extraction (OpenIE) is a problem of extracting triples from natural language text whose predicate relations are not aligned to any pre-defined ontology. |
| Approach: | They propose an open-source method to extract triples from semi-structured websites . they use a semi-supervised label propagation technique to create training data for relations . |
| Outcome: | The proposed method extracts over 2 million triples from 31 websites in the movie vertical. |
ZeroShotCeres: Zero-Shot Relation Extraction from Semi-Structured Webpages (2020.acl-main)
Copied to clipboard
| Challenge: | Existing work on information extraction from semi-structured websites has relied on manual data annotation and learning a model specific to a given template. |
| Approach: | They propose a solution for “zero-shot” open-domain relation extraction from webpages with previously unseen templates using a graph neural network-based approach. |
| Outcome: | The proposed model provides a 31% gain over baseline for zero-shot extraction in a new subject vertical. |
Extracting Shopping Interest-Related Product Types from the Web (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing e-commerce products are limited in their ability to assist customers in interest-oriented shopping. |
| Approach: | They propose to extract PTs from Web pages containing hand-crafted PT recommendations for SIs . they propose to use tree-transformer encoders for node classification to improve inter-node dependency modeling . |
| Outcome: | The proposed model outperforms the best baseline model by 2.37 F1 points on a WebPT dataset. |
PLAtE: A Large-scale Dataset for List Page Web Extraction (2023.acl-industry)
Copied to clipboard
Aidan San, Yuan Zhuang, Jan Bakus, Colin Lockard, David Ciemiewicz, Sandeep Atluri, Kevin Small, Yangfeng Ji, Heba Elfardy
| Challenge: | Existing methods for web extraction are limited by the limited number of available large-scale datasets. |
| Approach: | They introduce a dataset that focuses on shopping data and a list page web extraction task. |
| Outcome: | The proposed dataset is the first large-scale list page web extraction dataset . it contains 52,898 items and 156,014 attributes, making it the first dataset based on this task . |
Multi-modal Information Extraction from Text, Semi-structured, and Tabular Data on the Web (2020.acl-tutorials)
Copied to clipboard
| Challenge: | a tutorial explores the commonalities in the challenges and solutions developed to address information extraction from the World Wide Web. |
| Approach: | This tutorial examines methods for extracting information from the World Wide Web . it explores the commonalities in the challenges and solutions developed to address these different forms of text . |
| Outcome: | This paper examines the commonalities in the challenges and solutions developed to address the World Wide Web. |
Semi-Supervised Event Extraction with Paraphrase Clusters (N18-2)
Copied to clipboard
| Challenge: | Existing event extraction systems are limited in their accuracy due to the lack of available training data. |
| Approach: | They propose a method for self-training event extraction systems by bootstrapping additional training data. |
| Outcome: | The proposed method improves on ACE 2005 and TAC-KBP 2015 datasets. |