Challenge: Existing methods for extracting explanations from complex models are based on discovering a large number of features, and this affects interpretability.
Approach: They propose a model that leverages Large Language Models and clustering algorithms to discover a compact set of interpretable features.
Outcome: The proposed model reduces the number of features on 3 Style Classification tasks by 85–99% while reducing the number by 85.

Similar Papers

Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on explainable AI focus on post-hoc explanation methods that interpret trained models through external approximations.
Approach: They propose to categorize existing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction.
Outcome: The proposed approaches are categorized into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction.
LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown unprecedented performance across various tasks.
Approach: They propose an easy-to-use framework that integrates adapters into LLMs . they evaluate adapters on 14 datasets from two different reasoning tasks .
Outcome: The proposed framework can be used to fine-tune open-access language models with task-specific data and instruction data.
RANCC: Rationalizing Neural Networks via Concept Clustering (2020.coling-main)

Copied to clipboard

Challenge: Existing models that construct explanations concurrently with classification predictions are opaque.
Approach: They propose a self-explainable model for Natural Language Processing (NLP) text classification tasks . they extract a rationale from the text and use it to predict a concept of interest .
Outcome: The proposed model can be compressed without complicated compression techniques.
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders (2026.findings-eacl)

Copied to clipboard

Challenge: Concept-based explanations for large language models are not well understood in text classification.
Approach: They propose a model with a specialized classifier head and activation rate sparsity loss for sentence classification . they compare it to existing models with HI-Concept and ConceptShap .
Outcome: The proposed model improves both the causality and interpretability of the extracted features.
Explainable Text Classification with LLMs: Enhancing Performance through Dialectical Prompting and Explanation-Guided Training (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing explanation methods that generate keywords may be less effective due to missing critical contextual information.
Approach: They propose a new method to generate explanations for possible labels using LLMs and a dialectical prompt.
Outcome: The proposed method significantly improves accuracy and explanation quality over state-of-the-art methods on multiple datasets from diverse domains.
Sparse Autoencoder Features for Classifications and Transferability (2025.emnlp-main)

Copied to clipboard

Challenge: Sparse Autoencoders (SAEs) provide potential for uncovering structured, human-interpretable representations in Large Language Models (LLMs).
Approach: They analyze SAEs for interpretable feature extraction from Large Language Models in safety-critical classification tasks.
Outcome: The proposed framework outperforms hidden-state and BoW models while demonstrating cross-lingual toxicity detection and visual classification tasks.
Enabling LLM Knowledge Analysis via Extensive Materialization (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have majorly advanced NLP and AI, and a major success factor is their internalized factual knowledge.
Approach: They propose a method to comprehensively materialize an LLM’s factual knowledge through recursive querying and result consolidation.
Outcome: The proposed method provides constructive insights into the scope and structure of LLM knowledge (or beliefs) it provides scale, accuracy, bias, cutoff and consistency at the same time.
Can LLMs Augment Low-Resource Reading Comprehension Datasets? Opportunities and Challenges (2024.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive zero-shot performance on a wide range of NLP tasks.
Approach: They propose to use large language models to augment extractive reading comprehension datasets by fine-tuning their annotations and comparing their performance to human annotators.
Outcome: The proposed model can be used to augment extractive reading comprehension datasets.
Characterizing Large Language Models as Rationalizers of Knowledge-intensive Tasks (2024.findings-acl)

Copied to clipboard

Challenge: Large language models generate fluent text with minimal task-specific supervision, but their ability to generate rationales for knowledge-intensive tasks (KITs) remains under-explored.
Approach: They propose to generate retrieval-augmented rationalization of KIT model predictions via external knowledge guidance within a few-shot setting.
Outcome: The proposed rationales were compared with crowd-sourced rationale models on factuality, sufficiency, and convincingness.
Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: specialized LLMs are often limited in domain-specific applications that require specialized knowledge.
Approach: They provide a comprehensive overview of four key methods to enhance large language models by integrating domain-specific knowledge.
Outcome: The proposed methods are categorized into four key approaches: dynamic knowledge injection, static knowledge embedding, modular adapters, and prompt optimization.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations