Papers by Jacob Matthews
Semantics or spelling? Probing contextual word embeddings with orthographic noise (2024.findings-acl)
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) are used to generate contextual word embeddings . linguistics research has focused on semantic information in hidden states . |
| Approach: | They investigate whether a single character swap in the input word will not affect the resulting representation . they find that PLM-derived contextual word embeddings are highly sensitive to noise . |
| Outcome: | The results show that the PLM-derived representations are highly sensitive to noise . the fewer tokens used to represent a word at input, the more sensitive their CWE is . |