Papers by Kabilan Prasanna
IruMozhi: Automatically classifying diglossia in Tamil (2024.findings-naacl)
Copied to clipboard
| Challenge: | Literary Tamil is highly diglossic, with two very different registers in everyday use . Spoken Tamil is under-studied in modern NLP systems compared to Literary Tamil written in the Tamil script . |
| Approach: | They present a human-translated dataset of parallel text in Literary and Spoken Tamil. |
| Outcome: | The proposed model trains classifiers on the task of identifying which Tamil variety a text belongs to. |