Papers by Shu-Kai Hsieh
CxLM: A Construction and Context-aware Language Model (2022.lrec-1)
Copied to clipboard
| Challenge: | Constructions are direct form-meaning pairs with possible schematic slots . however, these slots are constrained by the embedded construction and the context . we propose that a conditional probability distribution could be described but language models cannot capture this distribution. |
| Approach: | They propose that a conditional probability distribution could describe constructions’ schematic slots. |
| Outcome: | The proposed model predicts masked slots more accurately than baselines and generates structurally and semantically plausible word samples. |
Computational Modeling of Affixoid Behavior in Chinese Morphology (2020.coling-main)
Copied to clipboard
| Challenge: | affixoid behavior in Mandarin Chinese is unclear due to polysemy and diachronic dynamics. |
| Approach: | They propose to use three quantitative features to model affixoid behavior in Mandarin Chinese to determine its status. |
| Outcome: | The proposed model shows that there are no clear criteria that can be used to identify an affix’s status in an isolating language like Mandarin Chinese. |
Character Jacobian: Modeling Chinese Character Meanings with Deep Learning Model (2022.coling-1)
Copied to clipboard
| Challenge: | Compounding is a prevalent word-formation process in Chinese morphology, where each character is bound and free when treated as a morpheme. |
| Approach: | They propose a model that learns non-linear relations between constituents and words and a character Jacobians model that describes character’s role in each word. |
| Outcome: | The proposed model predicts embeddings of real words from constituents but helps account for behavioral data of pseudowords. |
Fluid Annotation: A Granularity-aware Annotation Tool for Chinese Word Fluidity (L18-1)
Copied to clipboard
| Challenge: | Using word segmentation, we propose a wordhood annotation framework for Chinese language . word segmentations have been used for years in preprocessing NLP tasks for languages without explicit word delimiter. |
| Approach: | They propose a word-granularity-aware annotation framework for Chinese language . they argue that word segmentation is fluid in nature and that it rearranges the boundary of word segmentations and linguistic annotation. |
| Outcome: | The proposed framework rearranges the boundary between word segmentation and linguistic annotation and supports flexible annotation tasks for various linguistic and affective phenomena. |
Eigencharacter: An Embedding of Chinese Character Orthography (D19-64)
Copied to clipboard
| Challenge: | Chinese characters encode world knowledge through thousands of years evolution . |
| Approach: | They propose an embedding approach to encode Chinese orthography knowledge using eigencharacter space. |
| Outcome: | The proposed representations encode lexical knowledge embedded in Chinese characters and integrate with other computational models. |
Do You Believe It Happened? Assessing Chinese Readers’ Veridicality Judgments (2020.lrec-1)
Copied to clipboard
| Challenge: | Using data from news datasets, we examine readers' veridicality judgments to news events at sentence level. |
| Approach: | They collect and study Chinese readers’ veridicality judgments to news events . goal is to observe pragmatic behaviors of linguistic features under context . |
| Outcome: | The aim is to observe the pragmatic behaviors of linguistic features under context which affects readers in making veridicality judgments. |