Papers by Sin-En Lu
Exploring Methods for Building Dialects-Mandarin Code-Mixing Corpora: A Case Study in Taiwanese Hokkien (2022.findings-emnlp)
Copied to clipboard
| Challenge: | CM is a challenging task when mixed languages include dialects. |
| Approach: | They propose to construct a Hokkien-Mandarin CM dataset to overcome the limitation . they propose to use a linguistics-based toolkit to train the model for translation tasks . |
| Outcome: | The proposed model achieves good results on CM data translation while maintaining monolingual translation quality. |
BRCC and SentiBahasaRojak: The First Bahasa Rojak Corpus for Pretraining and Sentiment Analysis Dataset (2022.coling-1)
Copied to clipboard
| Challenge: | Code-mixing is prevalent in multilingual societies and is challenging to train . we use data augmentation to build a model to deal with code-mixed inputs . |
| Approach: | They propose to train a model to deal with code-mixing phenomena of Bahasa Rojak using data augmentation to construct a Bahasan Rojakin corpus and a pre-trained model to process input tokens. |
| Outcome: | The proposed model can tag the language of the input token automatically to process code-mixing input. |