Papers by Andrin Büchler
The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks (2026.findings-eacl)
Copied to clipboard
| Challenge: | a small-scale human evaluation confirms that the segments are highly parallel, making the dataset suitable for NLP applications. |
| Approach: | They present a first parallel corpus of Romansh idioms from 291 schoolbooks . they use automatic alignment methods to extract 207k multi-parallel segments from the books . |
| Outcome: | The proposed corpus is based on 291 schoolbook volumes, which are comparable in content for the five idioms. |