Challenge: Multilingual language models (LMs) have become a powerful tool in NLP, especially for non-English languages.
Approach: They propose a method to reduce a multilingual LM vocabulary to a target language by deleting potentially irrelevant tokens from its vocabulary.
Outcome: The proposed method can retain the original performance of the multilingual LM while being considerably smaller in size than the original model.

Similar Papers

XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Large multilingual models rely on a single vocabulary shared across 100+ languages . this vocabulary bottleneck limits the representational capabilities of multilingual model XLM-R .
Approach: They propose a new approach for scaling to large multilingual vocabularies by de-emphasizing token sharing between languages with little lexical overlap and assigning vocabulary capacity to achieve sufficient coverage for each individual language.
Outcome: The proposed model outperforms XLM-R on all language tasks and is particularly effective on low-resource tasks.
Efficient Vocabulary Reduction for Small Language Models (2025.coling-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have high computational costs and energy consumption, making their deployment in industrial settings difficult.
Approach: They propose a small language model that compresses the embedding layer and reduces model size without significant loss of performance.
Outcome: The proposed model reduces the embedding layer while maintaining performance while improving accuracy and performance.
Accelerating Multilingual Language Model for Excessively Tokenized Languages (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown a significant degree of multilingual proficiency on a variety of tasks in multiple languages.
Approach: They propose a framework to fine-tune a language model head and fine-track it while preserving its performance.
Outcome: The proposed framework increases the generation speed by 1.7 while maintaining the performance of pre-trained multilingual models on target monolingual tasks.
Improving Multilingual Models with Language-Clustered Vocabularies (2020.emnlp-main)

Copied to clipboard

Challenge: State-of-the-art multilingual models depend on vocabularies that cover all languages . but the methods for generating those vocalaries are not ideal for massively multilingual applications.
Approach: They propose a procedure for multilingual vocabulary generation that combines separately trained vocabularies of several automatically derived language clusters.
Outcome: The proposed procedure shows improvements across languages on multilingual benchmark tasks . the proposed procedure reduces out-of-vocabulary rate by a factor of 8 .
VEEF-Multi-LLM: Effective Vocabulary Expansion and Parameter Efficient Finetuning Towards Multilingual Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have a significant disadvantage for low-resource languages . VEEF-Multi-LLM-8B excels in multilingual instruction-following tasks .
Approach: They propose a low-resource multilingual large language model that expands the vocabulary for multilingual support.
Outcome: The proposed model outperforms existing models in multilingual instruction-following tasks, but lags behind English-centric models in some tasks.
Fast Vocabulary Transfer for Language Model Compression (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing methods to reduce model size and size are expensive and inefficient for some applications.
Approach: They propose a method that relies on vocabulary transfer to reduce model size and inference time while compromising on performance.
Outcome: The proposed method reduces model size and inference time while compromising on performance.
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior studies show that large language models map multilingual content into English-aligned representations at intermediate layers before projecting them back into target-language token spaces in the later layers.
Approach: They propose a method to identify and manipulate dimensions that are sparse and sparsity-based . they propose to use as few as 50 sentences of either parallel or monolingual data to manipulate these dimensions .
Outcome: Experiments on a multilingual generation control task show the interpretability of these dimensions.
Vocabulary-informed Language Encoding (2022.coling-1)

Copied to clipboard

Challenge: A Multilingual model relies on language encodings to identify input languages . a method to compute a vocabulary-informed language coding can improve multilingual models .
Approach: They propose a method to compute a vocabulary-informed language encoding as the language representation for a required language.
Outcome: The proposed method improves performance on unsupervised translation and cross-lingual embedding.
An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference (2024.findings-emnlp)

Copied to clipboard

Challenge: Cross-lingual vocabulary adaptation (CVA) methods have been proposed for adapting models to a target language . but effectiveness of these methods on increasing inference efficiency of generative large language models has not been explored.
Approach: They propose to use cross-lingual vocabulary adaptation methods to adapt models to a target language to improve downstream performance.
Outcome: The proposed methods significantly speed up models in four languages and four natural language understanding tasks.
What Do Compressed Multilingual Machine Translation Models Forget? (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that pre-trained models achieve state-of-the-art results in NLP tasks but their size makes it more challenging to apply them in resource-constrained environments.
Approach: They assess the impact of compression methods on multilingual Neural Machine Translation models for various language groups, gender, and semantic biases.
Outcome: The proposed compression methods improve models on different benchmarks for language groups, gender, and semantic biases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations