Papers by Sjur Moshagen

3 papers
The Ethical Question – Use of Indigenous Corpora for Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Creating language technology based on language data is becoming more popular . indigenous language resources are not comparable in that they would encode the most recent normativised language .
Approach: They describe an ethical way to work with indigenous languages based on language data . they say data driven methods make assumptions based upon majority languages they work with . authors say data-driven methods are not ethical or beneficial .
Outcome: The proposed method is ethical and sustainable, and can be applied to indigenous languages in an ethical way.
Modeling Northern Haida Verb Morphology (L18-1)

Copied to clipboard

Challenge: a computational model of the verbal morphology of Northern Haida is being developed . the model is capable of handling complex affixation patterns and morphophonological alternations .
Approach: They propose a computational model of the verbal morphology of Northern Haida based on finite state machines with a focus on verbs.
Outcome: The proposed model can handle complex affixation patterns and morphophonological alternations in the native language.
Unmasking the Myth of Effortless Big Data - Making an Open Source Multi-lingual Infrastructure and Building Language Resources from Scratch (2022.lrec-1)

Copied to clipboard

Challenge: During the last two decades, machine learning approaches have dominated the field of natural language processing (NLP) weak literary traditions give rise to corpora too unreliable to function as a model for NLP tools.
Approach: They propose an alternative to corpus-based language technology that can provide language technology solutions for minority languages.
Outcome: The proposed approach can provide language technology solutions for minority languages outside the reach of corpus-based language technology.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations