Papers by Yuan Zhuang

14 papers
Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for text regression lack local grounding and rely on shared representations.
Approach: They propose a distributional regression model with quantile tokens that insert dedicated quantiles into the input sequence.
Outcome: The proposed method outperforms baseline models on the inside Airbnb and StackSample datasets.
NAP2: A Benchmark for Naturalness and Privacy-Preserving Text Rewriting by Learning from Human (2025.findings-emnlp)

Copied to clipboard

Challenge: a large number of large language models are being used to protect user privacy . sanitizing sensitive text using two common strategies is the answer .
Approach: They propose sanitizing sensitive text using deleting expressions and abstracting them . they propose a tool for text rewriting that uses crowdsourcing and large language models .
Outcome: The proposed approach protects privacy before sending sensitive data to large language models . it combines crowdsourcing and large language modeling to create a text rewrite tool .
CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages (2025.coling-main)

Copied to clipboard

Challenge: Existing multilingual models such as XLM-R support only approximately 100-200 languages, leaving nearly 7,000 low-resource languages untapped.
Approach: They construct and open-source a dataset of four-language corpora obtained through machine translation into Chinese, Uyghur and Tibetan.
Outcome: The proposed dataset includes two resource-rich languages and two low-resource languages.
Rhetorical Questions in LLM Representations: A Linear Probing Study (2026.acl-long)

Copied to clipboard

Challenge: Rhetorical questions are asked not to seek information, but to persuade or signal stance . how large language models internally represent rhetorical questions remains unclear .
Approach: They analyze rhetorical questions in LLM representations using linear probes on two social-media datasets with different discourse contexts.
Outcome: The results show that rhetorical signals emerge early and are most stably captured by last-token representations.
Eliciting Affective Events from Language Models by Multiple View Co-prompting (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to generate training data using weakly labeled data are costly and limited .
Approach: They propose a method for acquiring and labeling affective events with multiple view co-prompting using pre-trained language models.
Outcome: The proposed approach improves state-of-the-art affective event classifier on two datasets.
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration (2025.acl-long)

Copied to clipboard

Challenge: Efficient data selection is crucial to accelerate the pretraining of language models . limited research has addressed the inherent conflicts between data selection methods .
Approach: They propose a multi-actor collaborative data selection mechanism that prioritizes data based on its specific criterion and updates prioritization rules using the current state of the model.
Outcome: The proposed model accelerates convergence in LM pretraining and achieves an average relative performance gain of 10.5% across multiple language model benchmarks.
SOLAR: Serendipity Optimized Language Model Aligned for Recommendation (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have shown strong potential in recommendation tasks . however, their application to serendipity-oriented recommendations remains challenging .
Approach: They propose a domain-adaptive instruction tuning method that aligns Large Language Models with recommendation tasks.
Outcome: The proposed framework bridges the domain gap between LLMs and recommendation tasks.
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for parameter-efficient fine-tuning (PEFT) are limited by computational costs and performance degradation.
Approach: They propose a method that integrates Low-Rank Adaptation and Mixture-of-Experts (MoE) they propose combining expert load imbalance and representation collapse to improve LLM performance .
Outcome: The proposed method outperforms homogeneous MoE-LoRA architectures in performance and parameter efficiency.
Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages (2025.acl-long)

Copied to clipboard

Challenge: Large language models demonstrate cross-lingual transfer capabilities, but these capabilities often fail to extend to low-resource languages, especially those utilizing non-Latin scripts.
Approach: They propose to combine character transliteration with Huffman coding to create a complete transliterations framework that can be extended to other low-resource languages.
Outcome: The proposed framework reduces storage requirements and improves accuracy and accuracy across multiple downstream tasks while maintaining performance on high-resource languages.
Affective Event Classification with Discourse-enhanced Self-training (2020.emnlp-main)

Copied to clipboard

Challenge: Prior work on recognizing affective events focused on producing lexical resources of verbs or event phrases with corresponding affective polarity values.
Approach: They propose a BERT-based model for affective event classification and a discourse-enhanced self-training method that iteratively improves the classifier with unlabeled data.
Outcome: The proposed model outperforms existing models with unlabeled data and improves recall and precision.
Recognizing Social Cues in Crisis Situations (2024.lrec-main)

Copied to clipboard

Challenge: During natural disasters, observations of other people's behavior can play an essential role in a person's decision-making.
Approach: They propose a task to categorize social cues in tweets during crisis situations using an annotated dataset of 6,000 tweets.
Outcome: The proposed task is challenging for existing systems and a manual task is based on a dataset of 6,000 tweets labeled with eight social cue categories.
My Heart Skipped a Beat! Recognizing Expressions of Embodied Emotion in Natural Language (2024.naacl-long)

Copied to clipboard

Challenge: a new task is needed to recognize physical manifestations of emotions in natural language . physical manifestation of emotions affects not only our mental state but also our physical state .
Approach: They propose a task to recognize expressions of embodied emotion in natural language . they use body part mentions with human annotations to extract emotional manner expressions .
Outcome: The proposed model can train without gold data and improve performance with gold data.
PLAtE: A Large-scale Dataset for List Page Web Extraction (2023.acl-industry)

Copied to clipboard

Challenge: Existing methods for web extraction are limited by the limited number of available large-scale datasets.
Approach: They introduce a dataset that focuses on shopping data and a list page web extraction task.
Outcome: The proposed dataset is the first large-scale list page web extraction dataset . it contains 52,898 items and 156,014 attributes, making it the first dataset based on this task .
Exploring the Role of Context to Distinguish Rhetorical and Information-Seeking Questions (2020.acl-srw)

Copied to clipboard

Challenge: Social media posts often contain questions, but many of them are rhetorical and do not seek information.
Approach: They propose a dataset containing questions in tweets paired with their prior tweets to provide context . they find that prior tweet and topic features can improve performance on this task .
Outcome: The proposed dataset compares questions in tweets with their prior tweets to provide context . it shows that prior tweet and topic features can improve performance on this task .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations