Papers by Chu-Ren Huang
Annotating Chinese Light Verb Constructions according to PARSEME guidelines (L18-1)
Copied to clipboard
| Challenge: | Using existing resources, we can annotate Chinese multiword expressions using PARSEME guidelines. |
| Approach: | They propose to use an existing resource containing Chinese light verbs to make an annotation of a Chinese UD treebank in two steps. |
| Outcome: | The proposed annotations are based on an existing treebank containing Chinese light verbs and are consistent with the proposed guidelines. |
EmbodiedBERT: Cognitively Informed Metaphor Detection Incorporating Sensorimotor Information (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for metaphor detection rely on heuristics such as Metaphor Identification Procedure (MIP) and Selection Preference Violation (SPV). |
| Approach: | They propose a cognitively motivated module that leverages the cognitive information of embodiment that can be derived from word embeddings and explicitly models the process of sensorimotor change that has been demonstrated as essential for metaphor processing. |
| Outcome: | The proposed module can improve metaphor detection compared with the heuristic MIP that has been applied previously. |
Be Helpful but Don’t Talk too Much - Enhancing Helpfulness in Conversations through Relevance in Multi-Turn Emotional Support (2024.emnlp-main)
Copied to clipboard
| Challenge: | a helpful speaker should maintain an "effect-effort" tradeoff for a conversation to help and support . a study aimed to cultivate the awareness of "optimal relevance" into the cognitive process of conversation agents . |
| Approach: | They integrate the "Cognitive Relevance Principle" into emotional support agents . they found that the "relevance principle" is effective in generating human-like, helpful, harmless conversations . |
| Outcome: | The proposed method improves human-likedness and support in multi-turn conversations . the source code will be available at https://github.com/CN-Eyetk/VLESA-ORL.git . |
Ciron: a New Benchmark Dataset for Chinese Irony Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | Automatic Chinese irony detection often lacks labeled benchmark datasets . despite its pervasive nature, irony is a trope whose actual meaning differs from what is literally enunciated. |
| Approach: | They propose to use a Chinese benchmark dataset for automatic Chinese irony detection to provide a benchmark for machine learning models. |
| Outcome: | The proposed dataset includes more than 8.7K posts, collected from Weibo, a micro blogging platform. |
Automatic Learning of Modality Exclusivity Norms with Crosslingual Word Embeddings (2020.starsem-1)
Copied to clipboard
| Challenge: | Normative studies on modality for English words are relatively common . however, they are limited to a relatively small number of languages and require costly ratings. |
| Approach: | They aim to learn a mapping between word embeddings and modality norms by training on a high-resource language and testing on . monolingual and crosslingual word embeds are used to predict modality association scores . |
| Outcome: | The proposed model predicts modality associations even when trained on an English resource and tested on a completely unseen language. |
Exploring Hybrid Sampling Inference for Aspect-based Sentiment Analysis (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for inference require multiple sampling with preset size . however, it is a high-cost method that requires multiple sampling . |
| Approach: | They propose a method that combines multiple and single sampling to greatly reduce the cost of multiple sampling without sacrificing performance. |
| Outcome: | The proposed method greatly reduces the cost of multiple sampling without sacrificing performance. |
Sentimental Image Generation for Aspect-based Sentiment Analysis (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent work on textual Aspect-Based Sentiment Analysis (ABSA) has demonstrated promising performance, but limited semantics derived from raw data. |
| Approach: | They propose a method that provides visual semantics to reinforce textual ABSA by adding additional augmentations to the input data. |
| Outcome: | The proposed method can provide visual semantics to reinforce the textual extraction. |
Comparing Probabilistic, Distributional and Transformer-Based Models on Logical Metonymy Interpretation (2020.aacl-main)
Copied to clipboard
| Challenge: | Logical metonymies are type clashes between an event-selecting verb and an entity-denoting noun . they are typically interpreted by inferring a hidden event on the basis of contextual cues . |
| Approach: | They propose to use probabilistic and distributional models to model logical metonymy interpretation . they compare models with the best Transformer-based models and some traditional distributional ones . |
| Outcome: | The proposed models perform well on a complex scenario, but low performance on some datasets suggests that logical metonymy is still a challenging phenomenon for computational modeling. |
Revisiting Classical Chinese Event Extraction with Ancient Literature Information (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies on classical Chinese event extraction focus on grafting the complex modeling from English or modern Chinese works, neglecting the unique characteristic of this language. |
| Approach: | They propose a Literary Vision-Language Model (VLM) for classical Chinese event extraction . they integrate annotations, historical background and character glyphs to capture the inner- and outer-context information from the sequence. |
| Outcome: | The proposed model can capture the inner- and outer-context information at nearly zero cost. |
Employing Glyphic Information for Chinese Event Extraction with Vision-Language Model (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies on event extraction have incorporated a variety of features, including textual elements and annotations. |
| Approach: | They propose a glyphic multi-modal Chinese event extraction model with hieroglyphic images to capture morphological structure from the sequence. |
| Outcome: | The proposed model can extract events from a Chinese and KBP Eval datasets at low cost. |
Modeling the Influence of Verb Aspect on the Activation of Typical Event Locations with BERT (2021.findings-acl)
Copied to clipboard
| Challenge: | Prior studies have shown that aspect of the main verb plays an important role in non-core semantic roles such as locations. |
| Approach: | They tested the popular language model BERT to determine whether its predictions of prototypical locations were influenced by aspect. |
| Outcome: | The language model BERT modelled the typicality of locations independently of the verb aspect. |
Affection Driven Neural Networks for Sentiment Analysis (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing deep neural network models lack mechanisms to highlight important sentiment terms. |
| Approach: | They propose a method to incorporate affective knowledge into deep neural network models by mapping affective influence vectors to an affective impact value and integrating them into long-term memory models to highlight affective terms. |
| Outcome: | The proposed approach improves on three large datasets by 1.0% to 1.5% on the benchmark datasets. |
Sparse Brains are Also Adaptive Brains: Cognitive-Load-Aware Dynamic Activation for LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing sparsity methods lack adaptivity to contextual or model structural demands or incur prohibitive computational overhead. |
| Approach: | They propose a Cognitive-Load-Aware Dynamic Activation framework that synergizes statistical sparsity with semantic adaptability. |
| Outcome: | The proposed framework achieves 20% average speedup with less than 2% accuracy degradation outperforming Griffin and TT. |
Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives (2026.acl-long)
Copied to clipboard
| Challenge: | a new study examines whether large language models acquire embodied cognition and cultural conventions from training data . demonstratives are a natural lens for evaluating linguistic phenomena that reflect cultural variation . aaron e. duan and j. nà: "the complexity of the language model is a major challenge for LLMs" |
| Approach: | They introduce demonstratives as a probe for grounded knowledge by analyzing 6,400 responses from 320 native speakers. |
| Outcome: | The proposed model fails to understand proximal–distal contrast and shows no cultural differences . the proposed model is a new probe for evaluating embodied cognition and cultural conventions . |
From Text to Historical Ecological Knowledge: The Construction and Application of the Shan Jing Knowledge Base (2024.lrec-main)
Copied to clipboard
| Challenge: | Traditional Ecological Knowledge (TEK) is a shared cultural heritage and crucial instrument to tackle environmental challenges. |
| Approach: | They propose to build a language resource based on Shanhai Jing (the classic of mountains and seas) written 2000 years ago and uses a stylized narrative and juxtaposition of knowledge from multiple domains to build the knowledge base. |
| Outcome: | The proposed knowledge base contains 1432 systematically classified entities and 3294 relationships. |
Are Word Embeddings Really a Bad Fit for the Estimation of Thematic Fit? (2020.lrec-1)
Copied to clipboard
| Challenge: | In recent years, vectors derived from neural network training have replaced count-based distributional semantic models as a de facto standard for word representation in NLP. |
| Approach: | They propose to evaluate count models and word embeddings on thematic fit estimation by taking into account a larger number of parameters and verb roles and introducing dependency-based embedders in the comparison. |
| Outcome: | The proposed model outperforms count models and word embeddings in thematic fit estimation tasks while introducing dependency-based embedders. |
New Compendium of a Myriad of Plants: A New Dataset Describing Ancient Chinese Plants (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to digitize ancient Chinese texts and extract information from them are shallow and inconsistent with modern realities. |
| Approach: | They propose to expand ancient Chinese datasets using large language model . they focus on Great Compendium of Myriad Flowers, an ancient plants dataset . |
| Outcome: | The proposed model can extract plant-related information from classical Chinese poetry and prose. |
CalligraphicOCR for Chinese Calligraphy Recognition (2025.emnlp-main)
Copied to clipboard
| Challenge: | Increasing efforts to digitize calligraphy have rely on isolated character recognition, requiring expensive manual splitting into single characters. |
| Approach: | They propose a calligraphicOCR model with calligraphy image augmentation and action-based corrector targeting the root of the problem. |
| Outcome: | The proposed model outperforms baseline models due to visual variations and domain shifts in semantics and is more accurate than previous models. |
Sina Mandarin Alphabetical Words:A Web-driven Code-mixing Lexical Resource (2020.aacl-main)
Copied to clipboard
| Challenge: | Mandarin Alphabetical Words (MAWs) are a key component of Modern Chinese . they are characterized by unique code-mixing idiosyncrasies influenced by language exchanges . |
| Approach: | They propose to construct a large collection of Mandarin Alphabetic Words from Sina Weibo . they propose to use a web-based technique to identify and validate MAWs . |
| Outcome: | The proposed method identifies 16,207 Mandarin Alphabetic Words (MAWs) using a web-based technique . the results show that the proposed method is useful for linguistic research and inquiries . |
An Effective Incorporating Heterogeneous Knowledge Curriculum Learning for Sequence Labeling (2025.acl-short)
Copied to clipboard
| Challenge: | Existing approaches to enhance sequence labeling models require data heterogeneity and additional modules. |
| Approach: | They propose a dual-stage curriculum learning framework specifically designed for sequence labeling tasks. |
| Outcome: | The proposed model improves training and accelerates training, mitigating the slow training issue of complex models. |