Challenge: Existing approaches to measure document similarity are inadequate for document pairs with non-comparable lengths, such as a long document and its summary.
Approach: They propose a document matching approach to bridge the gap between long documents and their abstract information in a common space of hidden topics.
Outcome: The proposed approach outperforms strong baselines on two matching tasks and incorporates domain knowledge to gain further performance improvement.

Similar Papers

Matching Varying-Length Texts via Topic-Informed and Decoupled Sentence Embeddings (2024.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to matching text with non-comparable lengths are limited due to truncation issues.
Approach: They propose a model that decouples sentences and embeds them into natural sentences for matching texts of significantly different lengths.
Outcome: The proposed model matches texts of significantly different lengths across three well-studied datasets.
Self-Supervised Document Similarity Ranking via Contextualized Language Models and Hierarchical Inference (2021.findings-acl)

Copied to clipboard

Challenge: Existing approaches to document-to-document similarity ranking are limited to relatively short documents or lack similarity labels.
Approach: They propose a self-supervised method for document similarity ranking that can be applied to documents of arbitrary length.
Outcome: The proposed model outperforms existing methods on large documents datasets.
Bridging the Gap between Relevance Matching and Semantic Matching for Short Text Similarity Modeling (D19-1)

Copied to clipboard

Challenge: Existing techniques for relevance and semantic matching cannot be easily adapted to the other.
Approach: They propose a model that incorporates a hybrid encoder module, a relevance matching module and co-attention mechanisms that capture context-aware semantic relatedness.
Outcome: The proposed model incorporates a hybrid encoder module, a relevance matching module and co-attention mechanisms that capture context-aware semantic relatedness.
Deep Relevance Ranking Using Enhanced Document-Query Interactions (D18-1)

Copied to clipboard

Challenge: Document relevance ranking is the task of ranking documents from a large collection using the query and the text of each document only.
Approach: They propose to use convolutional n-gram matching to inject rich context-sensitive encodings into their models, inspired by PACRR's convolution-based ngram matching features.
Outcome: The proposed models outperform baselines, DRMM, and PACRR on the BIOASQ and TREC ROBUST questions and document inputs.
Simple and Effective Text Matching with Richer Alignment Features (P19-1)

Copied to clipboard

Challenge: Existing models only use a single inter-sequence alignment layer to make full use of this process.
Approach: They propose to keep three key features available for inter-sequence alignment . they conduct experiments on four well-studied benchmark datasets .
Outcome: The proposed model is able to perform on four well-studied datasets with fewer parameters and the inference speed is at least 6 times faster than similar models.
Unsupervised Concept Representation Learning for Length-Varying Text Similarity (2021.naacl-main)

Copied to clipboard

Challenge: Existing document similarity approaches suffer from the information gap caused by context and vocabulary mismatches when comparing varying-length texts.
Approach: They propose an unsupervised concept representation learning approach to address this issue . they propose a concept-based document matching method to leverage recognition of local phrase features .
Outcome: The proposed method achieves a better F1 score than baseline models on real-world data sets.
Matching Article Pairs with Graphical Decomposition and Convolutions (P19-1)

Copied to clipboard

Challenge: Existing methods for matching sentence pairs do not perform well in longer documents . Existing approaches for matching sentences do not work in longer document understanding tasks .
Approach: They propose to model article pairs by comparing sentences that enclose same concept vertex . they propose to use a concept interaction graph to match articles by encoding sentences .
Outcome: The proposed methods show significant improvements over existing methods . the proposed datasets consist of 30K pairs of breaking news articles .
Interpretable Text Embeddings and Text Similarity Explanation: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging.
Approach: They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned .
Outcome: The proposed methods are compared with existing models and compare them with existing ones.
A Document Descriptor using Covariance of Word Vectors (P18-2)

Copied to clipboard

Challenge: Existing methods for retrieving documents using vectors have been used to model documents and queries using bag-of-words (BOW) representations.
Approach: They propose to use the word embeddings of a document to define a novel document descriptor.
Outcome: The proposed descriptor performs well against state-of-the-art methods in supervised and unsupervised environments.
Aspect-based Document Similarity for Research Papers (2020.coling-main)

Copied to clipboard

Challenge: Traditional document similarity measures do not consider in what aspects two documents are similar.
Approach: They extend document similarity with aspect information by performing a pairwise document classification task.
Outcome: The proposed approach is best performing on 172,073 research paper pairs from the ACL Anthology and CORD-19 corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations