Papers by Yi-Hsuan Lin
Enhancing Taiwanese Hokkien Dual Translation by Exploring and Standardizing of Four Writing Systems (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, machine translation systems cater to high-resource languages (HRLs), while low-resourced languages (LRLs) like Taiwanese Hokkien are relatively under-explored. |
| Approach: | They propose to use a pre-trained LLaMA 2-7B model specialized in Traditional Mandarin Chinese to leverage orthographic similarities between Taiwanese Hokkien Han and Traditional Mandarin China. |
| Outcome: | The proposed model bridges the gap between Taiwanese Hokkien and other low-resource languages by using a pre-trained LLaMA 2-7B model and a monolingual corpus. |
MingOfficial: A Ming Official Career Dataset and a Historical Context-Aware Representation Learning Framework (2023.emnlp-main)
Copied to clipboard
You-Jun Chen, Hsin-Yi Hsieh, Yu Lin, Yingtao Tian, Bert Chan, Yu-Sin Liu, Yi-Hsuan Lin, Richard Tsai
| Challenge: | In Chinese studies, understanding the nuanced traits of historical figures can be challenging due to the need for domain expertise, specialist knowledge, and context-specific insights. |
| Approach: | They propose a large-scale multi-modal dataset for Chinese officials from the Ming Dynasty that integrates structured and text data to enable investigation of social structures. |
| Outcome: | The proposed dataset could enable exploratory analysis of official identities and significantly boost performance in tasks such as identifying nuance identities from 24.6% to 98.2% F1 score in hold-out test set. |