Scalable Knowledge Graph Construction from Text Collections (D19-66)

Copied to clipboard

Challenge: Existing open-source solutions for analyzing unstructured text are lacking in the field of knowledge graph construction.
Approach: They propose a scalable open-source platform that "distills" a text collection into a knowledge graph . they scale out the Stanford CoreNLP toolkit via Apache Spark integration .
Outcome: The proposed platform scales out the Stanford CoreNLP toolkit via Apache Spark integration . it extracts mentions and relations from documents and then ingests them into a knowledge graph .

Similar Papers

CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era (2025.acl-long)

Copied to clipboard

Challenge: Graphs are used for storing open-domain knowledge and domain-specific enterprise data.
Approach: They propose to use property graph views on top of the underlying RDF graph to efficiently query LLMs.
Outcome: The proposed graph views can be efficiently queried by LLMs using Cypher . the proposed graphs have a large schema, overlapping and ambiguous relation types and lack of normalization.
End-to-End Construction of NLP Knowledge Graph (2021.findings-acl)

Copied to clipboard

Challenge: a new schema for NLP knowledge about tasks, datasets and metrics is proposed.
Approach: They propose a new schema that represents knowledge about tasks, datasets and metrics in the NLP domain.
Outcome: The proposed framework can be automatically built into scientific leaderboards . the proposed system achieves reasonable results for all relation types on this small-scale graph .
Knowledge Graph Based Synthetic Corpus Generation for Knowledge-Enhanced Language Model Pre-training (2021.naacl-main)

Copied to clipboard

Challenge: Existing work on data-to-text generation focused on domain-specific benchmark datasets.
Approach: They use a KG-Wikipedia text aligned corpus to verbalize the entire English Wikidata KG . they show that this approach can be used to integrate structured KGs and natural language corpora .
Outcome: The proposed method improves on open domain QA and the LAMA knowledge probe.
A High Precision Pipeline for Financial Knowledge Graph Construction (2020.coling-main)

Copied to clipboard

Challenge: Knowledge graphs are a standard for structured knowledge representation in the Semantic Web.
Approach: They propose to extract financial news articles into a knowledge graph by using a financial dictionary.
Outcome: The proposed pipeline extracts 342,000 financial news articles with a precision of 78% at the top-100 extractions.
Tree-KG: An Expandable Knowledge Graph Construction Framework for Knowledge-intensive Domains (2025.acl-long)

Copied to clipboard

Challenge: Knowledge graphs are a useful tool for organizing complex data in knowledge-intensive domains.
Approach: They propose an expandable framework that combines structured domain texts with advanced semantic techniques to create a tree-like graph from textbooks.
Outcome: The proposed framework surpasses competing methods in the text-Annotated dataset with high scores on the Text-Annalytated data.
Text Generation from Knowledge Graphs with Graph Transformers (N19-1)

Copied to clipboard

Challenge: Existing methods for generating text with structured inputs are expensive and require manual annotation.
Approach: They propose a graph transforming encoder which leverages relational structure of knowledge graphs without imposing linearization or hierarchical constraints.
Outcome: The proposed system produces more informative texts than competing methods.
Extracting Fine-Grained Knowledge Graphs of Scientific Claims: Dataset and Transformer-Based Results (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches focus on high-level description of how research is carried out . instead, we focus on the subtleties of how experimental associations are presented .
Approach: They propose a transformer-based approach to relational scientific information extraction that captures associations over experimental variables and their qualifications, subtypes, and evidence.
Outcome: The proposed schema captures causal, comparative, predictive, statistical, and proportional associations over experimental variables along with qualifications, subtypes, and evidence.
Multiple Knowledge GraphDB (MKGDB) (2020.lrec-1)

Copied to clipboard

Challenge: ConceptNet, DBpedia, WebIsAGraph, WordNet and Wikipedia category hierarchy are used to create a large-scale graph database.
Approach: They propose to use multiple taxonomy backbones extracted from 5 existing knowledge graphs to create a large-scale graph database.
Outcome: The proposed database is intended to favour and support the development of open-domain natural language processing applications relying on knowledge bases.
Mind the Query: A Benchmark Dataset towards Text2Cypher Task (2025.emnlp-industry)

Copied to clipboard

Challenge: Graph databases store data in nodes and relationships, enabling more natural modeling of complex, interconnected data.
Approach: They present a high-quality dataset for the Text2Cypher task . it is enabling the translation of natural language (NL) questions into executable Cypher queries over graph databases.
Outcome: The proposed dataset includes 27,529 NL queries and corresponding Cyphers spanning across 11 real-world graph datasets.
CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data (2026.acl-long)

Copied to clipboard

Challenge: Existing datasets are small, domain-limited, and lack diversity, constraining LLM progress.
Approach: They propose a knowledge Graph retrieval tool that can translate natural language questions into structured queries.
Outcome: Extensive experiments show that CypherSmith achieves state-of-the-art performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations