Papers with DBpedia

23 papers
The DBpedia Databus Tutorial: Increase the Visibility and Usability of Your Data (2024.lrec-tutorials)

Copied to clipboard

Challenge: Linked Open Data tutorial introduces DBpedia Databus, a FAIR data publishing platform . aimed at addressing data production and consumption challenges faced by knowledge graph stakeholders .
Approach: This tutorial introduces DBpedia Databus, a FAIR data publishing platform . it addresses challenges faced by data producers and consumers .
Outcome: This tutorial addresses challenges faced by data producers and consumers using DBpedia Databus.
Improving Knowledge Graph Embedding Using Simple Constraints (P18-1)

Copied to clipboard

Challenge: Recent efforts focused on designing more complicated models or incorporating extra information beyond triples.
Approach: They propose to use non-negativity constraints on entity representations and approximate entailment constraints on relation representations to improve KG embedding.
Outcome: The proposed model outperforms baseline models on WordNet, Freebase, and DBpedia.
An Industry Evaluation of Embedding-based Entity Alignment (2020.coling-industry)

Copied to clipboard

Challenge: Knowledge graphs (KGs) are increasingly important in various applications such as question answering and search engines.
Approach: They propose to use a supervised learning environment with unbiased seed mappings for training and validation to evaluate alignment methods in an industrial context.
Outcome: The proposed methods are evaluated in an industrial context and are compared with DBpedia and Wikidata benchmarks.
Improving Knowledge Base Construction from Robust Infobox Extraction (N19-2)

Copied to clipboard

Challenge: Existing knowledge bases are incomplete, resulting in poor answers and incompleteness.
Approach: They propose a method to extract Wikipedia infobox tables to populate an existing KB.
Outcome: The proposed method improves accuracy and completeness of the final KB significantly compared to DBpedia's baseline method.
Cross-lingual Knowledge Graph Alignment via Graph Convolutional Networks (D18-1)

Copied to clipboard

Challenge: Existing approaches to align multilingual knowledge graphs with counterparts in different languages are not effective.
Approach: They propose a novel approach for cross-lingual KG alignment via graph convolutional networks . they train GCNs to embed entities of each language into a unified vector space .
Outcome: The proposed approach gets the best performance on real multilingual KGs compared with other embedding-based approaches.
A Semantics-aware Transformer Model of Relation Linking for Knowledge Base Question Answering (2021.acl-short)

Copied to clipboard

Challenge: Existing knowledge base question answering systems do not leverage the explicit semantic parse of the question text.
Approach: They propose a transformer-based neural model that leverages the AMR semantic parse of a sentence.
Outcome: The proposed model outperforms the state-of-the-art on 4 popular benchmark datasets.
WeDH - a Friendly Tool for Building Literary Corpora Enriched with Encyclopedic Metadata (2020.lrec-1)

Copied to clipboard

Challenge: Linked Open Data repositories are difficult to use for text corpora enriched with metadata . a collaborative project aims to fill the access to textual resources available on the web and the possibility of combining these resources with sources of metadata extending the life and maintenance of the data itself.
Approach: They propose a web interface that allows users to leverage encyclopedic knowledge from DBpedia, wikidata and VIAF to enrich texts with bibliographical and exegetical knowledge.
Outcome: WeDH aims to fill the access to textual resources available on the web and the possibility of combining these resources with sources of metadata that can enrich the texts with useful information.
Location Name Extraction from Targeted Text Streams using Gazetteer-based Statistical Language Models (C18-1)

Copied to clipboard

Challenge: Location name extraction tool (LNEx) is a statistical language for extracting location names from informal and unstructured social media data.
Approach: They propose a location name extraction tool that extracts location names from social media data . they use n-gram statistics and location-related dictionaries to evaluate an observed n in targeted text .
Outcome: The proposed tool outperforms state-of-the-art taggers on 4,500 event-specific tweets . it improves the average F-Score by 33-179%, outperforming all tagger .
When Shallow is Good Enough: Automatic Assessment of Conceptual Text Complexity using Shallow Semantic Features (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches to automatic assessment of text complexity focus on syntactic and lexical complexity.
Approach: They propose to use graph-based deep semantic features to automatically assess conceptual text complexity by using DBpedia as a proxy to human knowledge.
Outcome: The proposed features outperform the state-of-the-art features on pairwise comparison of two versions of the same text and five-level classification task.
Farewell Freebase: Migrating the SimpleQuestions Dataset to DBpedia (C18-1)

Copied to clipboard

Challenge: Existing datasets for question answering over knowledge graphs lack answer triples from Freebase . a defunct knowledge graph makes it difficult to build "real-world" question answering systems .
Approach: They propose a benchmark dataset for simple question answering over knowledge graphs that maps SimpleQuestions entities and predicates from Freebase to DBpedia.
Outcome: The proposed dataset provides simple yet strong baselines with and without neural networks.
A State-transition Framework to Answer Complex Questions over Knowledge Base (D18-1)

Copied to clipboard

Challenge: Existing methods for complex question answering have some limitations . existing methods employ predefined patterns or templates to understand complex questions.
Approach: They propose a state transition-based approach to translate a natural language question to a semantic query graph.
Outcome: The proposed approach outperforms state-of-the-art methods on several benchmarks with two knowledge bases.
Multiple Knowledge GraphDB (MKGDB) (2020.lrec-1)

Copied to clipboard

Challenge: ConceptNet, DBpedia, WebIsAGraph, WordNet and Wikipedia category hierarchy are used to create a large-scale graph database.
Approach: They propose to use multiple taxonomy backbones extracted from 5 existing knowledge graphs to create a large-scale graph database.
Outcome: The proposed database is intended to favour and support the development of open-domain natural language processing applications relying on knowledge bases.
SYGMA: A System for Generalizable and Modular Question Answering Over Knowledge Bases (2022.findings-emnlp)

Copied to clipboard

Challenge: Knowledge Base Question Answering (KBQA) systems have limited generalizability across knowledge bases and multiple reasoning types.
Approach: They propose a modular approach for KBQA that is built on a framework adaptable to multiple knowledge bases and reasoning types.
Outcome: The proposed approach is generalized across multiple knowledge bases and reasoning types.
KORE 50ˆDYWC: An Evaluation Data Set for Entity Linking Based on DBpedia, YAGO, Wikidata, and Crunchbase (2020.lrec-1)

Copied to clipboard

Challenge: A major domain of research in natural language processing is named entity recognition and disambiguation (NERD).
Approach: They extend a widely-used data set to include NERD tasks for DBpedia and YAGO, Wikidata and Crunchbase.
Outcome: The extended data set allows for a broader spectrum of evaluation.
Leveraging Abstract Meaning Representation for Knowledge Base Question Answering (2021.findings-acl)

Copied to clipboard

Challenge: Existing approaches face challenges including complex question understanding and lack of large end-to-end training datasets.
Approach: They propose a modular knowledge base question answering system that leverages AMR parses for task-independent question understanding.
Outcome: The proposed system achieves state-of-the-art performance on two prominent KBQA datasets based on DBpedia.
A Two-Stage Approach towards Generalization in Knowledge Base Question Answering (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for Knowledge Base Question Answering focus on a specific knowledge base or evaluating it on underlying knowledge base requires non-trivial changes.
Approach: They propose a framework that separates semantic parsing from knowledge base interaction . they propose KBQA framework that allows generalization across knowledge bases .
Outcome: The proposed framework achieves comparable or state-of-the-art performance on datasets with a different knowledge base.
A Large Multilingual and Multi-domain Dataset for Recommender Systems (L18-1)

Copied to clipboard

Challenge: Existing algorithms for recommending items are limited and focused on specific domains.
Approach: They propose a multi-domain interests dataset to train and test Recommender Systems . the english dataset includes an average of 90 preferences per user on music, books, movies, celebrities, sport, politics .
Outcome: The proposed method exploits popular services such as Spotify, Goodreads and others to extract preferences from Twitter messages in Italian and English.
Aligning Cross-Lingual Entities with Multi-Aspect Information (D19-1)

Copied to clipboard

Challenge: Existing knowledge graphs that represent entities in different languages are not covered by existing systems.
Approach: They propose two ways to embed entities from multilingual knowledge graphs into the same vector space, where equivalent entities are close to each other.
Outcome: The proposed method significantly outperforms existing systems on two benchmark datasets.
Unsupervised Fine-tuning for Text Clustering (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to text clustering fine-tune pre-trained models have been limited.
Approach: They propose a method to fine-tune pre-trained models unsupervisedly for text clustering by learning text representations and cluster assignments using a clustering oriented loss.
Outcome: The proposed model outperforms baseline methods and achieves state-of-the-art results on three text clustering datasets.
Enhancing Task-Specific Distillation in Small Data Regimes through Language Generation (2022.coling-1)

Copied to clipboard

Challenge: Large-scale pretrained language models have led to significant improvements in Natural Language Processing, but they come at the cost of high computational and storage requirements.
Approach: They propose to distill knowledge from larger models to smaller ones through pseudo-labels on task-specific datasets.
Outcome: The proposed approach improves on the SST-2, MRPC, YELP-2, and TREC-6 datasets.
Cross-Lingual Knowledge Projection and Knowledge Enhancement for Zero-Shot Question Answering in Low-Resource Languages (2025.coling-main)

Copied to clipboard

Challenge: Knowledge bases (KBs) in low-resource languages are often incomplete, restricting the ability to do zero-shot question answering using multilingual language models.
Approach: They propose a novel cross-lingual mapping technique which improves word alignments extracted from parallel English-LRL text by combining lexical alignment, named entity recognition, and semantic alignment.
Outcome: The proposed approach improves zero-shot question answering accuracy by up to 17% compared to baselines without KB access.
A Large Interlinked Knowledge Graph of the Italian Cultural Heritage (2022.lrec-1)

Copied to clipboard

Challenge: Existing efforts to create knowledge bases are limited to relatively small resources, such as entities from libraries, archeological sites and museums.
Approach: They propose to create a large knowledge graph linking Italian cultural heritage entities with concepts defined on well-known knowledge bases.
Outcome: The proposed graph shows that the Italian cultural heritage entities are interlinked with concepts defined on well-known knowledge bases.
PESCO: Prompt-enhanced Self Contrastive Learning for Zero-shot Text Classification (2023.acl-long)

Copied to clipboard

Challenge: Existing text classification frameworks require large amounts of human-labeled documents to train .
Approach: They propose a contrastive learning framework that improves zero-shot text classification . they add prompts to enhance label retrieval and use retrieved labels to enrich training .
Outcome: The proposed framework achieves state-of-the-art on four benchmark text classification datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations