Papers by Marisa Hudspeth
Contextual morphologically-guided tokenization for Latin encoder models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing tokenization methods focus on information-theoretical goals like high compression and low fertility rather than linguistic goals like morphological alignment. |
| Approach: | They propose to incorporate morphological knowledge into tokenization to improve both morphology and downstream performance. |
| Outcome: | The proposed tokenization improves overall performance on four downstream tasks. |
Automated main concept generation for narrative discourse assessment in aphasia (2025.findings-acl)
Copied to clipboard
| Challenge: | Several advances have been made towards developing theoretical and computational methods for understanding narratives. |
| Approach: | They propose a method that generates MCs from novel stories that experts can edit manually. |
| Outcome: | The proposed method can generate most of the gold standard MCs for stories from an existing narrative summarization dataset. |