| Challenge: | Using a simplified version of GRU, we replace the GRUs at the middle layers of hierarchical recurrent models with Fixed-size Ordinally-Forgetting Encoding (FOFE). |
| Approach: | They propose to make the lower layers simpler than the upper ones to simplify two typical hierarchical recurrent models, namely Hierarchical Recurrent Encoder-Decoder (HRED) and R-NET, whose basic building block is GRU. |
| Outcome: | The proposed models contain less trainable parameters, consume less training time, and achieve slightly better performance than baseline models. |
Similar Papers
Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent Networks (D18-1)
Copied to clipboard
| Challenge: | Existing gated recurrent networks have a vanishing gradient, allowing for more matrix transformations and less transparent functions. |
| Approach: | They propose an additionsubtraction twin-gated recurrent network (ATR) to simplify neural machine translation. |
| Outcome: | The proposed system is more transparent than LSTM/GRU due to the simplification. |
Reproducing and Regularizing the SCRN Model (C18-1)
Copied to clipboard
| Challenge: | Recurrent neural networks (RNNs) have demonstrated tremendous success in sequence modeling . naive dropout, variational dropout and weight tying are common techniques used to regularize the SCRN model . |
| Approach: | They propose a Structurally Constrained Recurrent Network (SCRN) model and regularize it using existing techniques. |
| Outcome: | The proposed model outperforms the LSTM model on non-English data while being much simpler. |
Hybrid-Regressive Paradigm for Accurate and Speed-Robust Neural Machine Translation (2023.findings-acl)
Copied to clipboard
| Challenge: | Autoregressive translation (NAT) is less robust in decoding batch size and hardware settings than NAT. |
| Approach: | They propose a two-stage translation prototype that prompts a small number of AT predictions and fills in previously skipped tokens at once. |
| Outcome: | The proposed translation prototype achieves comparable translation quality with AT while having 1.5x faster inference speed regardless of batch size and device. |
From Bottom to Top: Extending the Potential of Parameter Efficient Fine-Tuning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to fine-tune large language models primarily focus on the interaction between different layers, ignoring the fact that different layers store different information. |
| Approach: | They propose a Parameter Efficient Fine-Tuning method which freeze pre-trained parameters and fine-tunes only a few task-specific parameters. |
| Outcome: | The proposed methods reduce parameter count to nearly half by omitting fine-tuning in the middle layers. |
SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs (2026.findings-acl)
Copied to clipboard
Jie Sun, Yu Liu, Lu Han, Qiwen Deng, Xiang Shu, Yang Xiao, Lintao Ma, Xingyu Lu, Jun Zhou, Pengfei Liu, Jiancan Wu, Xiang Wang
| Challenge: | Existing large-scale large-context models suffer from performance degradation when processing long numerical sequences. |
| Approach: | They propose a framework to mitigate attention dispersion by strategically inserting separator tokens into the model to recalibrat attention to local segments while preserving global context. |
| Outcome: | The proposed framework improves accuracy and reduces inference token consumption by 16.4% on 9 widely-adopted LLMs. |
Decoupling Generalization and Adaptation in Meta-Learning for Large Language Models (2026.acl-short)
Copied to clipboard
| Challenge: | Adapting large language models to specific downstream tasks requires multi-step fine-tuning with substantial training data, incurring significant computational overhead. |
| Approach: | They propose a framework that separates learning generalizable initializations and adaptation through dedicated parameter spaces. |
| Outcome: | The proposed framework outperforms existing meta-learning and standard multi-task baselines on common-sense reasoning, mathematics, logic, medical and coding benchmarks. |
MoDification: Mixture of Depths Made Easy (2025.naacl-long)
Copied to clipboard
Chen Zhang, Meizhi Zhong, Qimeng Wang, Xuantao Lu, Zheyu Ye, Chengqiang Lu, Yan Gao, Yao Hu, Kehai Chen, Min Zhang, Dawei Song
| Challenge: | Long-context efficiency is a trending topic in large language model (LLM) serving. |
| Approach: | They propose a method to combine long-context efficiency and mixture of depths to bring down both latency and memory. |
| Outcome: | The proposed method achieves 1.2 speedup in latency and 1.8 reduction in memory compared to original LLMs especially in long-context applications. |
Sparse Attention with Linear Units (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have suggested that sparse attention mechanisms can be made more interpretable by replacing the softmax activation with its sparser variants. |
| Approach: | They propose a method to replace softmax activation with a ReLU to achieve sparsity in attention by layer normalization with either a specialized initialization or an additional gating function. |
| Outcome: | The proposed model is easy to implement and more efficient than previously proposed sparse attention mechanisms. |
DenseLoRA: Dense Low-Rank Adaptation of Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Low-rank adaptation (LoRA) is an efficient approach for adapting large language models (LLMs) but many of the weights in these matrices are redundant, leading to inefficiencies in parameter utilization. |
| Approach: | They propose a low-rank adaptation approach that fine-tunes two low-ranked matrices and adapts them through a dense low-Rank matrix, improving parameter utilization and adaptation efficiency. |
| Outcome: | The proposed approach achieves 83.8% accuracy with only 0.01% of trainable parameters compared to LoRA's 80.8% with 0.70% of trainability parameters on LLaMA3-8B. |
LoRMA: Low-Rank Multiplicative Adaptation for LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models have shown impressive generalization capabilities, but can be expensive to fine-tune due to high computational costs. |
| Approach: | They propose a low-rank multiplicative Adaptation technique that shifts the paradigm of additive updates to a richer space of matrix multiplicative transformations. |
| Outcome: | The proposed approach overcomes computational complexity and rank bottlenecks in terms of matrix multiplication metrics. |