| Challenge: | Language models are a major class of natural language processing (NLP) models whose development has led to major progress in many areas like translation, speech recognition or summarization. |
| Approach: | They propose a hierarchical multi-scale language model where short time-scale dependencies are encoded in the hidden state of a lower-level recurrent neural network while longer time- scale dependencies can be encoded into the dynamic of the lower- level network. |
| Outcome: | The proposed model uses a meta-learner to update the weights of the lower-level neural network in an online meta-learning fashion to prevent catastrophic forgetting in the continuous learning framework. |
Similar Papers
Revisiting the Hierarchical Multiscale LSTM (C18-1)
Copied to clipboard
| Challenge: | Hierarchical Multiscale LSTM model learns structure from character input . high complexity of architecture, training and implementations might hinder its applicability . |
| Approach: | They propose to reproduce and ablate hierarchical multiscale LSTM language model and show that simplifying certain aspects of the architecture can improve its performance. |
| Outcome: | The proposed model performs better when simplified and linguistic units are learned by different levels of the model. |
RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging (2025.emnlp-main)
Copied to clipboard
Bowen Wang, Haiyuan Wan, Liwen Shi, Chen Yang, Peng He, Yue Ma, Haochen Han, Wenhao Li, Tiao Tan, Yongjian Li, Fangming Liu, Gong Yifan, Sheng Zhang
| Challenge: | Existing models that require task labels or performance trade-offs are susceptible to catastrophic forgetting. |
| Approach: | They propose a representation-aware model merging framework for continual learning without access to historical data. |
| Outcome: | The proposed framework outperforms baselines in knowledge retention and generalization across five NLP tasks and multiple continual learning scenarios. |
The Importance of Being Recurrent for Modeling Hierarchical Structure (D18-1)
Copied to clipboard
| Challenge: | Recent work shows that recurrent neural networks can implicitly capture hierarchical information when trained to solve common natural language processing tasks. |
| Approach: | They propose a convolutional sequence-to-sequence model that exploits hierarchical information implicitly. |
| Outcome: | The proposed model is recurrent and non-recurrent, and it can model hierarchical structure implicitly. |
SCALE: Upscaled Continual Learning of Large Language Models (2026.findings-acl)
Copied to clipboard
Jin-woo Lee, Junhwa Choi, Bongkyu Hwang, Jinho Choo, Bogun Kim, Jeongseon Yi, Joonseok Lee, DongYoung Jung, Jaeseon Park, Kyoungwon Park, Suk-hoon Jung
| Challenge: | Recent discussions suggest that further progress will come from scaling the right structure, not merely parameters or data, while preserving acquired knowledge. |
| Approach: | They propose a width upscaling architecture that inserts lightweight expansions into linear modules while freezing all pre-trained parameters. |
| Outcome: | The proposed architecture reduces severe forgetting while learning new knowledge on a controlled synthetic biography benchmark. |
Recurrent Neural Networks with Mixed Hierarchical Structures and EM Algorithm for Natural Language Processing (2022.lrec-1)
Copied to clipboard
| Challenge: | A variety of hierarchical RNN models have been proposed to incorporate hierarchically-based hierarchic information in modeling languages in the literature. |
| Approach: | They propose a latent indicator layer approach to identify and learn hierarchical information and develop an EM algorithm to handle the latent indicators layer in training. |
| Outcome: | The proposed approach outperforms other RNN-based models in document classification tasks. |
On Efficiently Representing Regular Languages as RNNs (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent work by Hewitt et al. (2020) provides an interpretation of the empirical success of recurrent neural networks (RNNs) as language models (LMs). |
| Approach: | They generalize their construction and show that RNNs can efficiently represent a larger class of LMs than previously claimed. |
| Outcome: | The results suggest that RNNs can represent a larger class of LMs than previously claimed . |
Hierarchical Phrase-Based Sequence-to-Sequence Learning (2022.emnlp-main)
Copied to clipboard
| Challenge: | a neural transducer that incorporates hierarchical phrases as a source of inductive bias during training and as explicit constraints during inference is described. |
| Approach: | They propose a neural transducer that incorporates hierarchical phrases as a source of inductive bias during training and as explicit constraints during inference. |
| Outcome: | The proposed model performs well on small scale machine translation benchmarks. |
Hierarchical Bracketing Encodings Work for Dependency Graphs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Sequence labeling (SL) is a simple yet effective paradigm for a wide range of natural language problems. |
| Approach: | They propose a new bracketing approach for dependency graph parsing that encodes graphs as sequences and n tagging actions. |
| Outcome: | The proposed approach significantly reduces label space while preserving structural information. |
Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods encode text and label hierarchy separately and mix their representations for classification, where the hierarchy remains unchanged for all input text. |
| Approach: | They propose to embed hierarchy into a text encoder by combining input and output data to generate a hierarchy-aware representation. |
| Outcome: | Extensive experiments on three benchmark datasets verify the effectiveness of the proposed model. |
Deep RNNs Encode Soft Hierarchical Syntax (P18-2)
Copied to clipboard
| Challenge: | Existing studies show that syntactic information is useful for a wide variety of NLP tasks. |
| Approach: | They propose to use word-level representations to learn internal representations that capture soft hierarchical notions of syntax from highly varied supervision. |
| Outcome: | The proposed model encodes significant amounts of syntax even without explicit supervision. |