Challenge: morphological inflection generation and historical text normalization tasks are character-level tasks that outperform recurrent models.
Approach: They propose a technique to handle feature-guided character-level transduction that further improves performance.
Outcome: The transformer outperforms recurrent models on morphological inflection and historical text normalization tasks.

Similar Papers

Rethinking Document-level Neural Machine Translation (2022.findings-acl)

Copied to clipboard

Challenge: Neural machine translation models are weak enough for document-level translation . current models only translate sentences individually, resulting in poor document coherence .
Approach: They propose to use the original Transformer model to test document-level neural machine translation . they find that the original transformer models can achieve strong results for document translation if trained properly .
Outcome: The proposed model outperforms sentence-level models on nine datasets and two sentence- level datasets across six languages.
Improving the Transformer Translation Model with Document-Level Context (D18-1)

Copied to clipboard

Challenge: Existing models for document-level context translation ignore documentlevel context.
Approach: They propose a document-level context encoder to represent document- level context and integrate it into the Transformer model.
Outcome: Experiments on NIST Chinese-English and IWSLT French-English datasets show that the proposed translation model outperforms the Transformer model significantly.
Towards Reasonably-Sized Character-Level Transformer NMT by Finetuning Subword Systems (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to train character-level models require very deep architectures that are difficult and slow to train.
Approach: They propose to fine tune a Transformer token-based model to get a model without token segmentation.
Outcome: The proposed model improves translation quality and robustness to noise while requiring less token segmentation.
The NLP Task Effectiveness of Long-Range Transformers (2023.eacl-main)

Copied to clipboard

Challenge: Existing benchmarks on long-range attention models have not been sufficient to develop efficient Transformers and their practical application on complex NLP tasks.
Approach: They propose to benchmark 7 Transformer variants on 5 difficult NLP tasks and 7 datasets to examine their capacity for long-range attention.
Outcome: The proposed models have advantages on content selection and query-guided decoding, but they come with previously unrecognized drawbacks such as insufficient attention to distant tokens and accumulated approximation error.
RingFormer: Rethinking Recurrent Transformer with Adaptive Level Signals (2025.findings-emnlp)

Copied to clipboard

Challenge: Transformers have shown strong performance in processing sequential data, but their parameters are larger . a novel approach to reduce the model parameters while maintaining high performance is proposed .
Approach: They propose a transformer-based model that processes input repeatedly in a circular, ring-like manner.
Outcome: The proposed approach reduces model parameters while maintaining high performance . the proposed approach is validated in the experiments.
Getting The Most Out of Your Training Data: Exploring Unsupervised Tasks for Morphological Inflection (2024.emnlp-main)

Copied to clipboard

Challenge: Pre-trained transformers have been shown to be effective in many natural language tasks, but are under-explored for character-level sequence to sequence tasks.
Approach: They propose to use pre-trained transformers for character-level morphological inflection in several languages to train models for unsupervised tasks.
Outcome: The proposed model outperforms the best two shared tasks on morphological inflection and graphemeto-phoneme conversion benchmarks.
Hierarchical Transformers Are More Efficient Language Models (2022.findings-naacl)

Copied to clipboard

Challenge: Transformers are impressive but inefficient and costly, which limits their applications and accessibility.
Approach: They first use different ways to downsample and upsamplify activations in Transformers to make them hierarchical.
Outcome: The proposed model outperforms Transformers on the ImageNet32 and enwik8 benchmarks.
Character-Level Translation with Self-attention (2020.acl-main)

Copied to clipboard

Challenge: Existing models for character-level neural machine translation operate on word-level, which makes them memory inefficient because of large vocabulary sizes.
Approach: They propose a transformer-based model and a novel variant that uses convolutions to combine information from nearby characters to facilitate character interactions.
Outcome: The proposed model outperforms the standard transformer model and learns more robust character alignments on bilingual and multilingual translation datasets.
TranSFormer: Slow-Fast Transformer for Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Prior work has focused on treating subwords as basic units in developing such systems.
Approach: They propose a slow-fast two-stream learning model that uses a “slow” branch to deal with subword sequences and a "fast" branch to cope with longer character sequences.
Outcome: The proposed model shows consistent BLEU improvements (larger than 1 BLUE point) on several machine translation benchmarks.
Learning Deep Transformer Models for Machine Translation (P19-1)

Copied to clipboard

Challenge: Neural machine translation models have advanced the previous state-of-the-art by learning mappings between sequences via neural networks and attention mechanisms.
Approach: They propose to use layer normalization to pass the combination of previous layers to the next layer to improve the model.
Outcome: The proposed model outperforms the shallow Transformer-Big/Base baseline model on English-German and Chinese-English tasks by 0.4-2.4 BLEU points.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations