Challenge: Neural machine translation models learn to map from source language sentences to target language sentences via continuous-space intermediate representations.
Approach: They propose an encoder with character attention which augments the (sub)word-level representation with character-level information and a decoder with multiple attentions that enable the representations from different levels of granularity to control the translation cooperatively.
Outcome: The proposed model outperforms the standard word-based model, subword-based models, and strong character-based ones on translation tasks.

Similar Papers

Multi-Granularity Self-Attention for Neural Machine Translation (D19-1)

Copied to clipboard

Challenge: Existing neural machine translation models use a deep multi-head self-attention network with no explicit phrase information.
Approach: They propose a neural network that combines multi-head self-attention and phrase modeling to train attention heads to attend to phrases in either n-gram or syntactic formalisms.
Outcome: The proposed approach improves on English-to-German and NIST Chinese-to English translation tasks.
Character-Level Translation with Self-attention (2020.acl-main)

Copied to clipboard

Challenge: Existing models for character-level neural machine translation operate on word-level, which makes them memory inefficient because of large vocabulary sizes.
Approach: They propose a transformer-based model and a novel variant that uses convolutions to combine information from nearby characters to facilitate character interactions.
Outcome: The proposed model outperforms the standard transformer model and learns more robust character alignments on bilingual and multilingual translation datasets.
Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English (2020.coling-main)

Copied to clipboard

Challenge: Recent work shows that deeper character-based neural machine translation models outperform subword-based models.
Approach: They propose to investigate the ability of character-based models to learn word senses and morphological inflections and the attention mechanism in Finnish into English translation.
Outcome: The character-based models outperform subword-based model in Finnish to English translation.
Learning Mutually Informed Representations for Characters and Subwords (2024.findings-naacl)

Copied to clipboard

Challenge: Pretrained language models rely on subword tokenization to process text as a sequence of subwords.
Approach: They propose a character-subword language model that integrates character and subword modalities into one model.
Outcome: The proposed model outperforms its backbone language models on English sequence labeling and classification tasks.
On the Importance of Word Boundaries in Character-level Neural Machine Translation (D19-56)

Copied to clipboard

Challenge: Neural Machine Translation models typically use a fixed-size lexical vocabulary . subword segmentation methods rely on statistical heuristics that lack any linguistic notion .
Approach: They propose a hierarchical decoding architecture for character-level NMT using subwords . they propose fewer parameters and a more efficient approach to perform translation at the level of words .
Outcome: The proposed model can reach higher translation accuracy than the subword-level model with fewer parameters while maintaining longer-distance contextual and grammatical dependencies.
Seeing Both the Forest and the Trees: Multi-head Attention for Joint Classification on Different Compositional Levels (2020.coling-main)

Copied to clipboard

Challenge: Neural networks can capture expressive language features, but insights into the link between words and sentences are difficult to acquire automatically.
Approach: They propose a deep neural network architecture that explicitly wires lower and higher linguistic components and evaluate its ability to perform the same task at different hierarchical levels.
Outcome: The proposed model outperforms equivalent models that are not incentivized towards compositional representations.
Combining Subword Representations into Word-level Representations in the Transformer Architecture (2020.acl-srw)

Copied to clipboard

Challenge: Currently dominant approaches use word-level tokens, but this increases the length of the sequences and makes it difficult to profit from word-based information.
Approach: They propose to combine subword-level representations into word-level ones in the first layers of the encoder, reducing the effective length of the sequences in the following layers.
Outcome: The proposed model maintains translation quality with no extra word-level information . it is superior to the current dominant method for incorporating word- level source language information a priori .
Improving Language and Modality Transfer in Translation by Character-level Modeling (2025.acl-long)

Copied to clipboard

Challenge: Current translation systems cover only 5% of the world's languages . expanding to the long-tail of low-resource languages requires data-efficient methods that rely on cross-lingual and cross-modal knowledge transfer.
Approach: They propose a character-based approach to improve adaptability to new languages and modalities by using a teacher-student approach and parallel translation data to obtain a SONAR character-level encoder.
Outcome: The proposed model outperforms subword-based models in speech-to-text translation on the FLEURS benchmark on 33 languages and achieves state-of-the-art generalizability to unseen languages.
Document-Level Neural Machine Translation with Hierarchical Attention Networks (D18-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) can be improved by including document-level contextual information.
Approach: They propose a hierarchical attention model that captures document-level contextual information and conditioning on the NMT model’s own hidden states.
Outcome: The proposed model improves the BLEU score over a strong NMT baseline with the state-of-the-art in context-aware methods and that both the encoder and decoder benefit from context in complementary ways.
Controlling Text Complexity in Neural Machine Translation (D19-1)

Copied to clipboard

Challenge: Prior work on text complexity has focused on simplifying input text in one language, primarily English.
Approach: They propose a method to align news articles written for different levels of target language proficiency.
Outcome: The proposed model outperforms pipeline approaches that translate and simplify text independently.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations