Papers by Sin-En Lu

2 papers
Exploring Methods for Building Dialects-Mandarin Code-Mixing Corpora: A Case Study in Taiwanese Hokkien (2022.findings-emnlp)

Copied to clipboard

Challenge: CM is a challenging task when mixed languages include dialects.
Approach: They propose to construct a Hokkien-Mandarin CM dataset to overcome the limitation . they propose to use a linguistics-based toolkit to train the model for translation tasks .
Outcome: The proposed model achieves good results on CM data translation while maintaining monolingual translation quality.
BRCC and SentiBahasaRojak: The First Bahasa Rojak Corpus for Pretraining and Sentiment Analysis Dataset (2022.coling-1)

Copied to clipboard

Challenge: Code-mixing is prevalent in multilingual societies and is challenging to train . we use data augmentation to build a model to deal with code-mixed inputs .
Approach: They propose to train a model to deal with code-mixing phenomena of Bahasa Rojak using data augmentation to construct a Bahasan Rojakin corpus and a pre-trained model to process input tokens.
Outcome: The proposed model can tag the language of the input token automatically to process code-mixing input.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations