What Part of the Neural Network Does This? Understanding LSTMs by Measuring and Dissecting Neurons (D19-1)
Copied to clipboard
| Challenge: | Biological neural systems consist of a huge number of neurons, and can react to the environment in complicated ways. |
| Approach: | They propose a metric to quantify the sensitivity of neurons to each label and conduct experiments to prove it. |
| Outcome: | The proposed metric is based on a set of experiments that show that dropping an arbitrary neuron significantly degrades the accuracy of the model. |
Similar Papers
From Language to Language-ish: How Brain-Like is an LSTM’s Representation of Nonsensical Language Stimuli? (2020.findings-emnlp)
Copied to clipboard
| Challenge: | LSTMs are often used to measure event related potentials, but are they able to generalize to new data in a human-like way? |
| Approach: | They asked whether an LSTM model represents a language sample with degraded semantic or syntactic information and whether it resembles the brain's reaction to the stimuli. |
| Outcome: | The results suggest that LSTMs and human brain handle nonsensical data similarly. |
LSTMs Compose—and Learn—Bottom-Up (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work in NLP shows that LSTMs capture compositional structure in language data. |
| Approach: | They propose to measure the decompositional interdependence between word meanings in an LSTM based on their gate interactions. |
| Outcome: | The proposed model can model syntactic relationships rather than learning the longer-range relations independently. |
The emergence of number and syntax units in LSTM language models (N19-1)
Copied to clipboard
| Challenge: | a recent study shows that LSTMs can capture syntax-sensitive generalizations such as long-distance number agreement. |
| Approach: | They investigate the inner mechanics of number tracking in LSTMs at the single neuron level . they find that long-distance number information is largely managed by two "number units" importantly, the behaviour of these units is partially controlled by other units to track syntactic structure . |
| Outcome: | The proposed model is based on a language model with a long-distance number agreement task. |
How much complexity does an RNN architecture need to learn syntax-sensitive dependencies? (2020.acl-srw)
Copied to clipboard
| Challenge: | Long-term memory (LSTM) networks are capable of encapsulating long-range dependencies . but simple recurrent networks (SRNs) have been less successful at capturing long-term dependencies and loci of grammatical errors in an unsupervised setting. |
| Approach: | They propose a new architecture that incorporates the decaying nature of neuronal activations and models the excitatory and inhibitory connections in a population of neurons. |
| Outcome: | The proposed architecture shows competitive performance relative to LSTMs on subject-verb agreement, sentence grammaticality, and language modeling tasks. |
How LSTM Encodes Syntax: Exploring Context Vectors and Semi-Quantization on Natural Text (2020.coling-main)
Copied to clipboard
| Challenge: | LSTMs are widely used to capture informative long-term syntactic dependencies, but how they are reflected in their internal vectors for natural text has not been adequately investigated. |
| Approach: | They analyze how syntactic dependencies are reflected in LSTM's internal gates by learning a language model where syntaktic structures are implicitly given. |
| Outcome: | The proposed model can predict whether a word is inside a phrase structure or not from a small number of components of the context-update vector. |
Exploiting Document Knowledge for Aspect-level Sentiment Classification (P18-2)
Copied to clipboard
| Challenge: | Existing public aspect-level datasets for aspect-based sentiment classification are small . existing methods for aspect level sentiment classification require annotation of all opinion targets . |
| Approach: | They propose two approaches that transfer knowledge from document-level data to improve aspect-level sentiment classification. |
| Outcome: | The proposed methods improve aspect-level sentiment classification on 4 public datasets. |
Simpler neural networks prefer subregular languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Inductive biases of neural networks are still poorly understood, says dr. johansen . subregular languages are thought to form a bound on human phonological patterns . |
| Approach: | They apply a relaxation of L0 regularization which induces sparsity to study inductive biases of LSTMs. |
| Outcome: | The proposed method is based on a relaxation of L0 regularization, which induces sparsity, and a subregular language bias in LSTMs is related to the cognitive bias observed in human phonology. |
Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models (2024.lrec-main)
Copied to clipboard
Boxi Cao, Qiaoyu Tang, Hongyu Lin, Shanshan Jiang, Bin Dong, Xianpei Han, Jiawei Chen, Tianshu Wang, Le Sun
| Challenge: | Pre-trained language models have shown remarkable memory formation, but vanilla networks without pre-training suffer catastrophic forgetting problem. |
| Approach: | They conduct experiments to investigate the retentive-forgetful contradiction between vanilla and pre-trained language models by controlling the target knowledge types, learning strategies and learning schedules. |
| Outcome: | The results show that pre-trained language models are forgetful and pre-training leads to retentive models . |
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions (2025.emnlp-main)
Copied to clipboard
Seyedali Mohammadi, Bhaskara Hanuma Vedula, Hemank Lamba, Edward Raff, Ponnurangam Kumaraguru, Francis Ferraro, Manas Gaur
| Challenge: | Exact label definitions are considered as clues to disambiguate unclear labels, helping models perform their tasks more effectively. |
| Approach: | They conducted controlled experiments on multiple explanation benchmark datasets and label definition conditions using expert-curated, LLM-generated, perturbed, and swapped definitions. |
| Outcome: | The results suggest that models often default to internal representations, particularly in general tasks, while domain-specific tasks benefit more from explicit definitions. |
Neuron-Level Knowledge Attribution in Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for attribution of knowledge in large language models struggle to operate at neuron level due to computational constraints. |
| Approach: | They propose a static method for pinpointing significant neurons using three metrics . they also propose identifying "query neurons" which activate these "value neurons" |
| Outcome: | The proposed method shows superior performance across three metrics compared to seven other methods . it analyzes six types of knowledge across attention and feed-forward network layers . |