Challenge: Existing datasets that provide alignments between natural language and knowledge bases (KB) triples are limited in size, lack coverage and are of unreported quality.
Approach: They propose to build a large scale dataset of alignments between Wikipedia abstracts and Wikidata triples that is two orders of magnitude larger than the largest available alignments dataset.
Outcome: The proposed dataset is two orders of magnitude larger than the largest available dataset and covers 2.5 times more predicates.

Similar Papers

A Survey on Training-free Alignment of Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a survey of large language models (LLMs) aims to ensure outputs adhere to human values, ethical standards, and legal norms.
Approach: They present the first systematic review of TF alignment methods . they categorize them by stages of pre-decoding, in-decoder and post-decoration .
Outcome: The proposed methods are based on training-free (TF) alignment techniques . they are able to be used in open-source and closed-source environments without retraining .
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment (2022.emnlp-main)

Copied to clipboard

Challenge: Existing word alignment models capture few interactions between input sentence pairs, which severely degrades the word alignment quality.
Approach: They propose to model deep interactions between input and target sentences using a two-stage training framework to train the model.
Outcome: The proposed model achieves the state-of-the-art (SOTA) performance on four out of five language pairs.
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Knowledge graphs provide structured, verifiable grounding for large language models . current LLMs use KGs as auxiliary structures for text retrieval .
Approach: They propose a pipeline that constructs KGs from open-domain texts using triplets and qualifiers.
Outcome: The proposed pipeline outperforms existing methods in retrieval-augmented generation.
Alignment Data base for a Sign Language Concordancer (2020.lrec-1)

Copied to clipboard

Challenge: a new study examines the need for sign language translators to have tools similar to text-to-text translation.
Approach: They propose to use a concordancer to search for parallel Franch-LSF segments . they use dozens of short news clips and 120 SL videos to align them manually .
Outcome: The proposed data base will be searched using a concordancer and expand in the future.
Learning to Map Natural Language Statements into Knowledge Base Representations for Knowledge Base Construction (L18-1)

Copied to clipboard

Challenge: Currently, the construction and updating of knowledge bases rely on human labor.
Approach: They propose to map relational phrases in triples from natural language to knowledge base predicate format.
Outcome: The proposed mapping results show high quality and promising coverage on relational phrases compared to previous research.
Better Together: Modern Methods Plus Traditional Thinking in NP Alignment (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that end-to-end systems are not structurally free.
Approach: They propose to use dictionary- and word vector-based baselines to align NPs in the bitext . they argue that alignment of NP's in MT can be improved by using old-fashioned methods .
Outcome: a new study shows that alignment of NPs in the bitext is relevant even in an end-to-end paradigm . the proposed system can be improved by bringing in old-fashioned methods, the authors argue .
Hierarchical Relation-Guided Type-Sentence Alignment for Long-Tail Relation Extraction with Distant Supervision (2022.findings-naacl)

Copied to clipboard

Challenge: Distant supervision uses triple facts to label corpus for relation extraction, leading to wrong labeling and long-tail problems.
Approach: They propose a model to enrich distantly-supervised sentences with entity types by injecting context-free and -related backgrounds into sentences to alleviate sentence-level wrong labeling.
Outcome: The proposed model achieves state-of-the-art on benchmarks and in overall and long-tail performance.
BinaryAlign: Word Alignment as Binary Sequence Labeling (2024.acl-long)

Copied to clipboard

Challenge: State-of-the-art word alignment training methods require a different class depending on the availability of gold data for a particular language pair.
Approach: They propose a novel word alignment technique based on binary sequence labeling that outperforms existing approaches in both scenarios.
Outcome: The proposed method outperforms existing models on non-English language pairs and performs stratified error analysis over alignment error type.
TaKG: A New Dataset for Paragraph-level Table-to-Text Generation Enhanced with Knowledge Graphs (2022.findings-aacl)

Copied to clipboard

Challenge: Existing table-to-text generation benchmarks have some limitations, such as E2E and ToTTo focusing on singlesentence generation tasks.
Approach: They propose a new table-to-text generation dataset called TaKG that uses a set of knowledge graphs to enhance table input.
Outcome: The proposed model outperforms existing models for short-text generation tasks and shows reliable performance on long-text generated across a variety of metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations