Papers by Johannes Hoffart

5 papers
A Study of the Importance of External Knowledge in the Named Entity Recognition Task (P18-2)

Copied to clipboard

Challenge: Existing studies have shown that external knowledge is important for Named Entity Recognition .
Approach: They propose a modular framework that divides knowledge into four categories according to depth . they show the effects when incrementally adding deeper knowledge .
Outcome: The proposed framework outperforms agnostic frameworks with more external knowledge . the proposed frameworks outperformed agrarian frameworks on two standard datasets .
diaNED: Time-Aware Named Entity Disambiguation for Diachronic Corpora (P18-2)

Copied to clipboard

Challenge: Named Entity Disambiguation (NED) systems perform well on news articles but quality drops when inputs span long time periods.
Approach: They propose a time-aware method that resolves ambiguities even when mention contexts give only few cues.
Outcome: The proposed method improves on a newly created diachronic corpus.
KGPool: Dynamic Knowledge Graph Context Selection for Relation Extraction (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for relation extraction (RE) use only expanded facts from the knowledge graph .
Approach: They propose a method for relation extraction from a single sentence . they use a neural network to expand the context with additional facts from the KG .
Outcome: The proposed method is more accurate than state-of-the-art methods on standard datasets.
CHOLAN: A Modular Approach for Neural Entity Linking on Wikipedia and Wikidata (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to target end-to-end entity linking over knowledge bases are not efficient.
Approach: They propose a modular approach to target end-to-end entity linking over knowledge bases.
Outcome: The proposed approach outperforms state-of-the-art approaches on two well-known knowledge bases.
Unsupervised Multi-View Post-OCR Error Correction With Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Prior work used text generation techniques or redundancy in similar passages for OCR error correction, which is not appropriate in cases of low corpus redundancies or weak document contextual information.
Approach: They propose to use a pretrained language model to reconcile different OCR views in unsupervised way so that their combination contains fewer errors than each individual view.
Outcome: The proposed model can reconcile multiple OCR views so that their combined version contains fewer errors than the best OCR view.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations