Challenge: Sentence pair modeling is critical for many NLP tasks, such as paraphrase identification and semantic textual similarity.
Approach: They propose to use subwords to represent sentences without pretrained word embeddings . they find that subword models can achieve new state-of-the-art results without pretraining .
Outcome: The proposed models can achieve state-of-the-art results on two social media datasets and competitive results on news data for paraphrase identification.

Similar Papers

Neural Network Models for Paraphrase Identification, Semantic Textual Similarity, Natural Language Inference, and Question Answering (C18-1)

Copied to clipboard

Challenge: Sentence pair modeling is a fundamental technique underlying many NLP tasks.
Approach: They analyze several neural network designs for sentence pair modeling and compare their performance extensively across eight datasets.
Outcome: The proposed models perform well across eight datasets including paraphrase identification, semantic textual similarity, natural language inference, and question answering tasks.
What is the best recipe for character-level encoder-only modelling? (2023.acl-long)

Copied to clipboard

Challenge: aims to benchmark recent progress in language understanding models that output contextualised representations at the character level.
Approach: They aim to find the best way to build and train character-level BERT-like models by comparing architectural innovations with pretraining objectives.
Outcome: The proposed model outperforms a token-based model on a set of evaluation tasks with a fixed training procedure.
CharBERT: Character-aware Pre-trained Language Model (2020.coling-main)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) construct word representations at subword level with Byte-Pair Encoding (BPE) or its variations . but these methods split a word into subword units and make it incomplete and fragile .
Approach: They propose a character-aware pre-trained language model to tackle OOV problems . they construct contextual word embedding for each token from sequential character representations .
Outcome: The proposed model improves on the existing models on multiple NLP benchmarks.
Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models (D18-1)

Copied to clipboard

Challenge: Recent advances in text normalization have limited applications in other languages . a novel approach to text normalizing uses character embeddings and word embedds .
Approach: They propose a sequence-to-sequence model with character-based attention that uses pre-trained word embeddings to model subword information.
Outcome: The proposed model achieves state-of-the-art F1 score on Arabic spelling correction task despite being small and unsuited for the task.
Are the Best Multilingual Document Embeddings simply Based on Sentence Embeddings? (2023.findings-eacl)

Copied to clipboard

Challenge: obtaining document embeddings at document level is challenging due to computational requirements and lack of appropriate data.
Approach: They compare methods to produce document-level representations from sentences based on LASER, LaBSE, and Sentence BERT pre-trained multilingual models.
Outcome: The proposed methods produce document-level representations from sentences in 8 languages . the results show that a clever combination of sentence embeddings is usually better than encoding the full document as a single unit.
Neural Machine Translation without Embeddings (2021.naacl-main)

Copied to clipboard

Challenge: Existing models operate over subword tokens, but byte-based models employ a different approach . a one-hot representation of each byte does not hurt performance, but it improves BLEU scores .
Approach: They propose to represent every computerized text as a sequence of bytes via UTF-8 . this eliminates the need for an embedding layer and improves performance .
Outcome: The proposed model improves BLEU scores on byte-to-byte translation models compared to character-level models . the proposed model does not require an embedding layer and does not drop out of the decoder .
Pretrained Language Models for Sequential Sentence Classification (D19-1)

Copied to clipboard

Challenge: Recent successful models for document-level understanding have used hierarchical encoding and CRFs to capture dependencies between subsequent labels.
Approach: They propose a pretrained language model that captures contextual dependencies without hierarchical encoding nor a CRF.
Outcome: The proposed model captures contextual dependencies without hierarchical encoding nor a CRF on four datasets, including a new dataset of structured scientific abstracts.
Paraphrase-based Contrastive Learning for Sentence Pair Modeling (2025.naacl-srw)

Copied to clipboard

Challenge: Existing methods to improve performance of sentence pair modeling are not available on a large-scale for non-English languages.
Approach: They propose a method to apply contrastive learning to pre-trained masked language models . they use sentence embeddings of paraphrase pairs to make similar sentences .
Outcome: The proposed method can be used on four sentence pair modeling tasks in English and Japanese.
Generalizing Word Embeddings using Bag of Subwords (D18-1)

Copied to clipboard

Challenge: Existing word embeddings techniques have a fixed vocabulary, i.e., they can only provide vectors over a finite set of common words that appear frequently in a given corpus.
Approach: They propose a subword-level word vector generation model that views words as bags of character n-grams and provides good vectors for rare or unseen words.
Outcome: The proposed model performs state-of-the-art in English word similarity task and in joint prediction of part-of speech tag and morphosyntactic attributes in 23 languages.
Gating Mechanisms for Combining Character and Word-level Word Representations: an Empirical Study (N19-3)

Copied to clipboard

Challenge: Existing studies show that combining character and word-level representations improves word and sentence representations . however, word-based embeddings do not account for derivational processes resulting in syntactically-similar words with different meanings.
Approach: They propose to combine character and word-level representations to improve word and sentence representations.
Outcome: The proposed method performed well in several word similarity datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations