Reconsidering the Past: Optimizing Hidden States in Language Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | In this paper, we present Hidden-State Optimization (HSO) for language models at inference time. |
| Approach: | They propose a method that uses the log-probability gradient to update hidden states rather than the model parameters to improve the performance of transformer language models. |
| Outcome: | The proposed method improves performance of transformer language models at inference time. |
Similar Papers
Hidden State Variability of Pretrained Language Models Can Guide Computation Reduction for Transfer Learning (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to transfer a pretrained language model include fine-tuning all the parameters in the language model and adapting all its subsets. |
| Approach: | They propose to select layers based on the variability of their hidden states given a task-specific corpus. |
| Outcome: | The proposed model reduces the computational cost of transfer learning methods without sacrificing performance. |
HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation Perturbation (2023.acl-long)
Copied to clipboard
| Challenge: | Existing techniques to fine-tune pre-trained language models on downstream tasks are inadequate. |
| Approach: | They propose a technique to perturb hidden Transformers representations by enhancing generalization of hidden representations from different layers. |
| Outcome: | The proposed technique outperforms vanilla fine-tuning and enhances generalization of hidden representations from different layers. |
Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model Predictions (2023.acl-long)
Copied to clipboard
| Challenge: | Recent work on why Transformer-based large language models make predictions has made their behavior opaque due to the complexity of the computations performed within each layer. |
| Approach: | They propose a linear decomposition of final hidden states from autoregressive language models based on each initial input token, which is exact for virtually all contemporary Transformer architectures. |
| Outcome: | The proposed method analyzes the influence of input tokens on model probabilities over a sequence of upcoming words with only one forward pass from the model. |
Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Instruction-fine-tuned large language models (LLMs) under 14B parameters underperform on NLU tasks . we explore a framework to improve the NLU capabilities of LLMs . |
| Approach: | They propose to use Proximal Policy Optimization to improve NLU capabilities . they frame NLU as a reinforcement learning environment and optimize for reward signals . |
| Outcome: | The proposed framework outperforms supervised fine-tuning on GLUE and superGLUE tasks. |
Scaling Hidden Markov Language Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Hidden Markov models are a fundamental tool for sequence modeling that separates the hidden state from the emission structure. |
| Approach: | They propose methods for scaling hidden Markov models to massive state spaces while maintaining efficient exact inference and effective regularization. |
| Outcome: | The proposed methods are much more accurate than previous HMMs and n-gram-based methods, making progress towards the performance of state-of-the-art NN models. |
Hidden States as Early Signals: Step-level Trace Evaluation and Pruning for Efficient Test-Time Scaling (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to speed up parallel scaling have relied on similarity-based or confidence-based pruning, but these signals do not reliably indicate trace quality. |
| Approach: | They propose a pruning framework that evaluates reasoning steps using hidden states and dynamically prunes unpromising traces during generation. |
| Outcome: | The proposed framework reduces end-to-end inference latency by 45%–70% on average compared to self-consistency while improving reasoning accuracy. |
An Empirical Study on Hyperparameter Optimization for Fine-Tuning Pre-trained Language Models (2021.acl-long)
Copied to clipboard
| Challenge: | In the recent years, pre-trained language models have achieved great success in the NLP community. |
| Approach: | They propose two general strategies and an experimental procedure to troubleshoot HPO’s failure cases. |
| Outcome: | The proposed methods outperform grid search on two state-of-the-art language models using the same time budget and overfitting. |
Better Exploiting Latent Variables in Text Modeling (P19-1)
Copied to clipboard
| Challenge: | Consistent gains in performance on two datasets, Penn Treebank and Yahoo, indicate the generalizability of our method. |
| Approach: | They propose a method to exploit latent variables through hidden state averaging by sampling latent variable multiple times at a gradient step. |
| Outcome: | The proposed method shows consistent gains on two datasets showing that it is generalizable. |
Multi-armed bandits for resource efficient, online optimization of language model pre-training: the use case of dynamic masking (2023.findings-acl)
Copied to clipboard
| Challenge: | Using a Bayesian optimization framework, we pre-train Transformer-based language models (TLMs) using a multi-armed bandit framework requires high computational resources and introduces many unresolved design choices. |
| Approach: | They propose a Bayesian optimization framework for resource efficient pre-training of Transformer-based language models. |
| Outcome: | The proposed framework achieves lower MLM loss in fewer epochs, across settings, while avoiding expensive hyperparameter grid search. |
DenseSSM: State Space Models with Dense Hidden Connection for Efficient Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) face excessive computational and memory requirements due to the commonly used Transformer architecture. |
| Approach: | They propose a method to enhance the flow of hidden information between layers in large language models by selectively integrating shallow-layer hidden states into deeper layers. |
| Outcome: | The proposed method maintains parallelizability and inference efficiency of SSMs while significantly boosting performance on public benchmarks. |