Papers by Stephen Bothwell
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin (2024.lrec-main)
Copied to clipboard
| Challenge: | Historical linguists hypothesize systems of sound change to explain the evolution of language over time, but the evidence is limited. |
| Approach: | They propose a dataset that consists of roughly 3,000 pairs of forms from Proto-Italic and Latin. |
| Outcome: | The proposed dataset enables historical linguists to enhance other datasets by enhancing them with the existing datasets. |
Introducing Rhetorical Parallelism Detection: A New Task with Datasets, Metrics, and Baselines (2023.emnlp-main)
Copied to clipboard
| Challenge: | Parallelism is a common stylistic tool in rhetorical structures, but it is rarely investigated in the field of natural language processing. |
| Approach: | They propose a task of rhetorical parallelism detection to investigate its structure and meaning . they use a Latin and adapted Chinese dataset to define parallelise and define it using a family of metrics . |
| Outcome: | The proposed method achieves F1 scores on Latin and Chinese datasets. |