Challenge: Large Language Models (LLMs) encode and store contextual information, but internal mechanisms are opaque.
Approach: They propose a toolkit that assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions and measures intrinsic dimensionality of representations.
Outcome: The proposed framework assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions, and measures the intrinsic dimensionality of representations.

Similar Papers

Punctuations and Predicates in Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Recent work has shown that LLMs perform tasks in ways that diverge significantly from human reasoning.
Approach: They examine the computational importance of punctuation tokens in large language models . they use zeroing and layer-swapping techniques to examine their necessity and sufficiency .
Outcome: The proposed model differs in GPT-2, DeepSeek, and Gemma in that punctuation is necessary and sufficient in multiple layers . the findings offer new insight into the internal mechanisms of punctuations in LLMs and have implications for interpretability and model analysis.
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Current language models process text as sequences of tokens that roughly correspond to words . individual tokens are often semantically unrelated to the meanings of the words/concepts they comprise .
Approach: They propose a method to "read out" the implicit vocabulary of an autoregressive LLM by examining differences in token representations across layers.
Outcome: The proposed method "reads out" the implicit vocabulary of an autoregressive LLM by examining differences in token representations across layers.
A Multi-Perspective Analysis of Memorization in Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) can generate the same sequences contained in the pre-train corpus, known as memorization.
Approach: They analyze the relationship between memorization and outputs from Large Language Models (LLMs) they show a sudden drop and increase in the frequency of input tokens when generating memorized/unmemorized sequences .
Outcome: The proposed model can generate the same sequences contained in the pre-train corpus, and it can predict unmemorized tokens.
Spelling-out is not Straightforward: LLMs’ Capability of Tokenization from Token to Characters (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) can spell out tokens character by character with high accuracy, yet struggle with more complex character-level tasks.
Approach: They examine how large language models internally represent character-level information during the spelling-out process.
Outcome: The embedding layer does not fully encode character-level information, especially beyond the first character.
How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have focused on enhancing the factualness of large language models using context knowledge.
Approach: They propose to use ChatGPT to construct probing datasets that provide diverse and coherent evidence corresponding to various facts.
Outcome: The proposed model can encode knowledge across different layers, and it is compared with existing models.
Frozen LLMs are Native Decoders for High-Norm Semantic Vectors (2026.acl-long)

Copied to clipboard

Challenge: Existing compression methods selectively prune tokens based on information-theoretic metrics, resulting in interpretability but risking the loss of fine-grained information.
Approach: They propose a landmark-based compression framework for long contexts that captures global dependencies over landmark tokens.
Outcome: The proposed framework outperforms soft compression baselines on four QA benchmarks.
Insights into LLM Long-Context Failures: When Transformers Know but Don’t Tell (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit positional bias, struggling to utilize information from the middle or end of long contexts.
Approach: They propose to examine LLMs' long-context generalizations by probing their hidden representations.
Outcome: The proposed models excel at processing extended contexts while preserving their positional bias.
Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model Predictions (2023.acl-long)

Copied to clipboard

Challenge: Recent work on why Transformer-based large language models make predictions has made their behavior opaque due to the complexity of the computations performed within each layer.
Approach: They propose a linear decomposition of final hidden states from autoregressive language models based on each initial input token, which is exact for virtually all contemporary Transformer architectures.
Outcome: The proposed method analyzes the influence of input tokens on model probabilities over a sequence of upcoming words with only one forward pass from the model.
Exploring Context Strategies in LLMs for Discourse-Aware Machine Translation (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models excel at machine translation, but the impact of how LLMs utilize different forms of contextual information on discourse-level phenomena remains underexplored.
Approach: They examine how different forms of context influence standard MT metrics and specific discourse phenomena such as formality, pronoun selection, and lexical cohesion.
Outcome: Evaluating multiple LLMs across multiple domains and language pairs, the findings consistently show that context boosts translation and discourse-specific performance.
Can Large Language Models Understand Context? (2024.findings-eacl)

Copied to clipboard

Challenge: Existing evaluation methodologies for Large Language Models (LLMs) have been inadequate to evaluate their ability to understand contextual features.
Approach: They propose a benchmark to assess large language models' ability to understand context by adapting existing datasets to suit their evaluation.
Outcome: The proposed model performs better under the in-context learning pretraining scenario than state-of-the-art models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations