Papers by Armin Hoenen
Knowing the Author by the Company His Words Keep (L18-1)
Copied to clipboard
| Challenge: | In traditional linguistics, there exists a famous saying that one should know a word by the company it keeps. |
| Approach: | They propose a method which uses word embeddings to identify pairwise relational features in the context of authorship attribution. |
| Outcome: | The proposed method is based on three literary corpora and shows that word similarity is a key feature in the authorship attribution task. |
From Manuscripts to Archetypes through Iterative Clustering (L18-1)
Copied to clipboard
| Challenge: | philologists have since the beginnings of the age of print attempted to provide one textual representation of this variety using stemmata . stemmats are trees depicting the copy history (manuscripts = nodes, Copy processes = edges) |
| Approach: | They propose to use stemmata to extract the most likely common ancestor of all observed variants and then iteratively cluster them to create a single textual representation. |
| Outcome: | The proposed method uses trees depicting the copy history to find the base text which is most likely the latest common ancestor of all observed variants. |
Multi Modal Distance - An Approach to Stemma Generation With Weighting (L18-1)
Copied to clipboard
| Challenge: | Stemma generation is a task where manuscripts are copied and copied from each other and from M. Existing methods to generate stemma using unweighted token similarity weighting have been used. |
| Approach: | They propose to use a distance model to weight the texts of M1 and M2 to estimate the most likely tree from a series of mapping processes. |
| Outcome: | The proposed method is small in the experimental scenario(s) it is based on psycholinguistically gained distance matrices of letters in three modalities: vision, audition and motorics. |