Papers by Britney Whyte
Locating and Extracting Relational Concepts in Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing knowledge recall models lack interpretability for relational concepts . a hidden state expresses causal effects of relational concept in input prompts . |
| Approach: | They propose to use causal mediation analysis to find hidden states that express relational concepts in LLMs. |
| Outcome: | The proposed representations exhibit high credibility and can be flexibly transplanted into other recall processes. |