Papers by Kunzang Namgyal
Building a Part-of-Speech Tagged Corpus for Drenjongke (Bhutia) (2020.aacl-srw)
Copied to clipboard
| Challenge: | a corpus of sentences and 1379 tokens were generated for the first Drenjongke corpus . the language is considered "vulnerable," "definitely endangered" and "severely endangered." |
| Approach: | They propose to generate the first corpus of the Tibetan language using a phrase book . they propose to use 34 Part-of-Speech (PoS) tags to define the first Drenjongke corpus . |
| Outcome: | The first corpus of the Drenjongke language comprises 275 sentences and 1379 tokens . the paper plans to expand with other materials to promote further studies of the language . |