Papers by Zilu Tang

5 papers
A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information (2025.findings-acl)

Copied to clipboard

Challenge: Prior research has focused on toxicity and polarization as separate problems . extreme polarizing deepens divisions, often leading to hostility and fragmentation .
Approach: They propose to use a multi-label Indonesian dataset annotated for toxicity, polarization, and annotator demographic information to study polarizing language and toxicity.
Outcome: The proposed dataset shows that polarization cues improve toxicity classification and vice versa.
Mitigating Hallucinated Translations in Large Language Models with Hallucination-focused Preference Optimization (2025.naacl-long)

Copied to clipboard

Challenge: Machine Translation (MT) systems based on fine-tuned large language models (LLMs) are at a higher risk of generating hallucinations, which can severely undermine user’s trust and safety.
Approach: They propose a method that intrinsically learns to mitigate hallucinations during the model training phase.
Outcome: The proposed method reduces hallucinations by 89% on an average across three unseen target languages while preserving translation quality.
Disentangling Text and Math in Word Problems: Evidence for the Bidimensional Structure of Large Language Models’ Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that LLMs struggle with text interpretation and equation solving, despite distinct proficiencies in textual and mathematical components.
Approach: They disentangle textual interpretation and mathematical solving steps in word problems drawn from Brazil's largest college entrance exam and popular grade school-level benchmark GSM8K.
Outcome: The proposed model outperforms LLMs in Brazil's largest college entrance exam and popular grade school-level benchmark.
AugCSE: Contrastive Sentence Embedding with Diverse Augmentations (2022.aacl-main)

Copied to clipboard

Challenge: Similar work has shown that a single augmentation can be used to learn a robust generalpurpose representation with contrastive learning.
Approach: They propose a unified framework to utilize diverse sets of data augmentations to achieve a better, general-purpose sentence embedding model.
Outcome: The proposed framework achieves state-of-the-art results on downstream transfer tasks and performs competitively on semantic textual similarity tasks, using only unsupervised data.
Explain-then-translate: an analysis on improving program translation with self-generated explanations (2023.findings-emnlp)

Copied to clipboard

Challenge: Using self-generated natural language explanations improves zero-shot performance by 12% on average.
Approach: They propose to use self-generated natural language explanations as an intermediate step for code-to-code translation with language models.
Outcome: The proposed approach improves zero-shot performance by 12% on average . the proposed approach is not evaluated on a broader set of languages including low-resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations