| Challenge: | Existing approaches to measure document similarity are inadequate for document pairs with non-comparable lengths, such as a long document and its summary. |
| Approach: | They propose a document matching approach to bridge the gap between long documents and their abstract information in a common space of hidden topics. |
| Outcome: | The proposed approach outperforms strong baselines on two matching tasks and incorporates domain knowledge to gain further performance improvement. |
Similar Papers
Matching Varying-Length Texts via Topic-Informed and Decoupled Sentence Embeddings (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing approaches to matching text with non-comparable lengths are limited due to truncation issues. |
| Approach: | They propose a model that decouples sentences and embeds them into natural sentences for matching texts of significantly different lengths. |
| Outcome: | The proposed model matches texts of significantly different lengths across three well-studied datasets. |
Self-Supervised Document Similarity Ranking via Contextualized Language Models and Hierarchical Inference (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to document-to-document similarity ranking are limited to relatively short documents or lack similarity labels. |
| Approach: | They propose a self-supervised method for document similarity ranking that can be applied to documents of arbitrary length. |
| Outcome: | The proposed model outperforms existing methods on large documents datasets. |
Bridging the Gap between Relevance Matching and Semantic Matching for Short Text Similarity Modeling (D19-1)
Copied to clipboard
| Challenge: | Existing techniques for relevance and semantic matching cannot be easily adapted to the other. |
| Approach: | They propose a model that incorporates a hybrid encoder module, a relevance matching module and co-attention mechanisms that capture context-aware semantic relatedness. |
| Outcome: | The proposed model incorporates a hybrid encoder module, a relevance matching module and co-attention mechanisms that capture context-aware semantic relatedness. |
Deep Relevance Ranking Using Enhanced Document-Query Interactions (D18-1)
Copied to clipboard
| Challenge: | Document relevance ranking is the task of ranking documents from a large collection using the query and the text of each document only. |
| Approach: | They propose to use convolutional n-gram matching to inject rich context-sensitive encodings into their models, inspired by PACRR's convolution-based ngram matching features. |
| Outcome: | The proposed models outperform baselines, DRMM, and PACRR on the BIOASQ and TREC ROBUST questions and document inputs. |
Simple and Effective Text Matching with Richer Alignment Features (P19-1)
Copied to clipboard
| Challenge: | Existing models only use a single inter-sequence alignment layer to make full use of this process. |
| Approach: | They propose to keep three key features available for inter-sequence alignment . they conduct experiments on four well-studied benchmark datasets . |
| Outcome: | The proposed model is able to perform on four well-studied datasets with fewer parameters and the inference speed is at least 6 times faster than similar models. |
Unsupervised Concept Representation Learning for Length-Varying Text Similarity (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing document similarity approaches suffer from the information gap caused by context and vocabulary mismatches when comparing varying-length texts. |
| Approach: | They propose an unsupervised concept representation learning approach to address this issue . they propose a concept-based document matching method to leverage recognition of local phrase features . |
| Outcome: | The proposed method achieves a better F1 score than baseline models on real-world data sets. |
Matching Article Pairs with Graphical Decomposition and Convolutions (P19-1)
Copied to clipboard
| Challenge: | Existing methods for matching sentence pairs do not perform well in longer documents . Existing approaches for matching sentences do not work in longer document understanding tasks . |
| Approach: | They propose to model article pairs by comparing sentences that enclose same concept vertex . they propose to use a concept interaction graph to match articles by encoding sentences . |
| Outcome: | The proposed methods show significant improvements over existing methods . the proposed datasets consist of 30K pairs of breaking news articles . |
Interpretable Text Embeddings and Text Similarity Explanation: A Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Text embeddings are a fundamental component in many NLP tasks, but their interpretation and explanation remain challenging. |
| Approach: | They propose a framework for interpretable text embeddings and text similarity explanation . they characterize the main ideas, approaches, and trade-offs and discuss lessons learned . |
| Outcome: | The proposed methods are compared with existing models and compare them with existing ones. |
A Document Descriptor using Covariance of Word Vectors (P18-2)
Copied to clipboard
| Challenge: | Existing methods for retrieving documents using vectors have been used to model documents and queries using bag-of-words (BOW) representations. |
| Approach: | They propose to use the word embeddings of a document to define a novel document descriptor. |
| Outcome: | The proposed descriptor performs well against state-of-the-art methods in supervised and unsupervised environments. |
Aspect-based Document Similarity for Research Papers (2020.coling-main)
Copied to clipboard
| Challenge: | Traditional document similarity measures do not consider in what aspects two documents are similar. |
| Approach: | They extend document similarity with aspect information by performing a pairwise document classification task. |
| Outcome: | The proposed approach is best performing on 172,073 research paper pairs from the ACL Anthology and CORD-19 corpus. |