Challenge: Existing word representation models for morphologically rich languages use subword-level information, but their systematic comparative analysis across typologically diverse languages and tasks is still missing.
Approach: They propose a framework for learning subword-informed word representations that allows for easy experimentation with different segmentation and composition components.
Outcome: The proposed framework allows for easy experimentation with different segmentation and composition components, as well as advanced techniques based on position embeddings and self-attention.

Similar Papers

Understanding Subword Compositionality of Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) take sequences of subwords as input, requiring them to compose subword representations into meaningful word-level representations.
Approach: They propose to probe how large language models compose subword information . they find structural similarity, semantic decomposability, and form retention are key aspects .
Outcome: The proposed models can be classified into three distinct groups, the authors show . they show that they can achieve great performance when probing layer by layer their sensitivity to semantic decompositionality .
Generalizing Word Embeddings using Bag of Subwords (D18-1)

Copied to clipboard

Challenge: Existing word embeddings techniques have a fixed vocabulary, i.e., they can only provide vectors over a finite set of common words that appear frequently in a given corpus.
Approach: They propose a subword-level word vector generation model that views words as bags of character n-grams and provides good vectors for rare or unseen words.
Outcome: The proposed model performs state-of-the-art in English word similarity task and in joint prediction of part-of speech tag and morphosyntactic attributes in 23 languages.
How Suitable Are Subword Segmentation Strategies for Translating Non-Concatenative Morphology? (2021.findings-emnlp)

Copied to clipboard

Challenge: Data-driven subword segmentation is the default strategy for open-vocabulary machine translation but may not be sufficiently generic for learning non-concatenative morphology.
Approach: They propose to test data-driven subword segmentation on non-concatenative morphological phenomena in a controlled, semi-synthetic setting.
Outcome: The proposed model can translate non-concatenative morphological phenomena in a controlled, semi-synthetic setting.
Segmentation-free compositional n-gram embedding (N19-1)

Copied to clipboard

Challenge: Existing word embedding models depend on word segmentation, but this method is difficult when corpora written in noisy or unsegmented languages.
Approach: They propose a new method that models words, phrases and sentences seamlessly without word segmentation.
Outcome: The proposed method is very effective for noisy corpora written in unsegmented languages such as Chinese and Japanese.
Subword models struggle with word learning, but surprisal hides it (2025.acl-short)

Copied to clipboard

Challenge: Subword LMs struggle to discern words and non-words with high accuracy, character LM models do this easily and consistently.
Approach: They propose to model word learning in subword and character language models with the psycholinguistic lexical decision task.
Outcome: The results suggest that word learning and syntactic learning are separable in character LMs.
Reusing Weights in Subword-Aware Neural Language Models (N18-1)

Copied to clipboard

Challenge: a statistical language model assigns a probability to a sequence of words . data sparsity is a major problem in building traditional n-gram language models .
Approach: They propose several ways to reuse subword embeddings and other weights in subword-aware neural language models.
Outcome: The proposed techniques do not benefit a competitive character-aware model . but they show significant reductions in model sizes and performance.
Unsupervised Cross-Lingual Representation Learning (P19-4)

Copied to clipboard

Challenge: a comprehensive survey of cutting-edge weakly-supervised and unsupervised cross-lingual word representations is presented .
Approach: This tutorial provides a comprehensive survey of recent work on weakly-supervised and unsupervised cross-lingual word representations.
Outcome: This tutorial provides a comprehensive survey of cutting-edge weakly-supervised and unsupervised word representations.
Gating Mechanisms for Combining Character and Word-level Word Representations: an Empirical Study (N19-3)

Copied to clipboard

Challenge: Existing studies show that combining character and word-level representations improves word and sentence representations . however, word-based embeddings do not account for derivational processes resulting in syntactically-similar words with different meanings.
Approach: They propose to combine character and word-level representations to improve word and sentence representations.
Outcome: The proposed method performed well in several word similarity datasets.
PBoS: Probabilistic Bag-of-Subwords for Generalizing Word Embedding (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing word embeddings assume fixed finite-size vocabularies, hindering their ability to provide useful word representations for out-of-vocaulary words.
Approach: They propose a model that generalizes word embeddings without extra contextual information . they use the spellings of words to model subword segmentation and compute subword-based compositional word embeds.
Outcome: The proposed model can generate meaningful subword segmentations without any source of explicit morphological knowledge.
Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations (D18-1)

Copied to clipboard

Challenge: Existing approaches to generalization to resource-rich languages are difficult . a recent study shows that word representations can be useful in low resource languages .
Approach: They propose two approaches for improving generalization to low-resource languages by adapting continuous word representations using linguistically motivated subword units.
Outcome: The proposed method improves generalization to low resource languages . it requires neither parallel corpora nor bilingual dictionaries and requires no parallel training .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations