Papers by Tianhui Zhang
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages (2025.acl-long)
Copied to clipboard
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad
| Challenge: | Emotion recognition is an umbrella term for several NLP tasks, but most work on high-resource languages has focused on low-resourced languages. |
| Approach: | They propose to use emotion recognition to describe perceived emotions in 28 different languages and across several domains to identify and annotate the datasets. |
| Outcome: | The proposed datasets cover low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers. |
C-World: A Computer Use Agent Environment Creator (2026.acl-long)
Copied to clipboard
Ziqiao Xi, Shuang Liang, Qi Liu, Jiaqing Zhang, Letian Peng, Fang Nan, Meshal Nayim, Tianhui Zhang, Rishika Mundada, Lianhui Qin, Biwei Huang, Kun Zhou
| Challenge: | C-World enables users to build agent environments on demand. |
| Approach: | They propose a system that enables users to build agent environments on demand. |
| Outcome: | The proposed system outperforms baselines on 119k samples and achieves Spearman = 0.883 ranking correlation with real execution. |
Evaluating the Evaluation of Diversity in Commonsense Generation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation metrics for commonsense generation are unclear on which metrics are best suited for evaluating the diversity of outputs. |
| Approach: | They propose to use a large language model to analyze commonsense generation data to determine which diversity metrics are best suited for commonsensing. |
| Outcome: | The proposed metrics outperform form-based metrics and show high correlations with the LLM-based ratings. |
Improving Diversity of Commonsense Generation by Large Language Models via In-Context Learning (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown proficiency in enhancing the generation quality across various tasks without the need for any fine-tuning. |
| Approach: | They propose a method that diversifies the LLM generations while preserving their quality. |
| Outcome: | The proposed method can be used as training data to improve diversity in existing commonsense generators. |
Synthetic Data Generation for Training Diversified Commonsense Reasoning Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing Generative Commonsense Reasoning datasets are created using a small number of human annotators, covering only a narrow set of commonsense scenarios. |
| Approach: | They propose to use a synthetic dataset to train diverse commonsense generators. |
| Outcome: | The proposed model improves both generation diversity and quality compared with vanilla models and human-crafted datasets across different size Large Language Models (LLMs). |
Evaluating the Effect of Retrieval Augmentation on Social Biases (2026.eacl-long)
Copied to clipboard
| Challenge: | RAG is a popular method for injecting up-to-date knowledge into LLMs. |
| Approach: | They examine how RAG modulates social biases across three languages and four categories . they find that biased documents are amplified even when base LLM has low-level of intrinsic bias . |
| Outcome: | The proposed method can enhance factual accuracy but its effect on social biases is not well understood. |