Papers by Anika Harju
data2lang2vec: Data Driven Typological Features Completion (2025.coling-main)
Copied to clipboard
| Challenge: | Language typology databases are useful for multilingual Natural Language Processing (NLP) but their coverage is limited, with only 28.9% of all possible combinations specified in the database. |
| Approach: | They propose to use textual data to improve feature prediction by using a multi-lingual Part-of-Speech tagger and a more realistic evaluation setup to focus on likely to be missing typology features. |
| Outcome: | The proposed model outperforms previous studies on missing features in 1,749 languages and with external statistical features and machine learning algorithms. |