Papers by Ming-Bin Chen
Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection (2026.findings-acl)
Copied to clipboard
Zhiwei Liu, Yupeng Cao, Yuechen Jiang, Mohsinul Kabir, Polydoros Giannouris, Chen Xu, Ziyang Xu, Tianlei Zhu, Md. Tariquzzaman, Triantafillos Papadopoulos, Yan Wang, Lingfei Qian, Xueqing Peng, Zhuohan Xie, Ye Yuan, Saeed Almheiri, Abdulrazzaq Alnajjar, Ming-Bin Chen, Harry Stuart, Paul Thompson, Prayag Tiwari, Alejandro Lopez-Lira, Xue Liu, Jimin Huang, Sophia Ananiadou
| Challenge: | Existing research on LLM biases has focused on direct questioning or general-purpose settings . pronounced behavioral biase despite their growing deployment in financial analysis, forecasting, and decision support. |
| Approach: | They propose a benchmark to evaluate behavioral biases of large language models in MFMD . they use a multilingual financial misinformation dataset to integrate these with misinformation claims . |
| Outcome: | The proposed benchmark evaluates behavioral biases of large language models across economic scenarios. |
CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics (2026.acl-long)
Copied to clipboard
| Challenge: | Using a semantic memory, we score each utterance along three interpretable dimensions: Novelty, Relevance, and Implication Scope. |
| Approach: | They propose a framework for Conversational Information Gain that evaluates each utterance in terms of how it advances collective understanding of the target topic. |
| Outcome: | The proposed framework evaluates each utterance in terms of how it advances collective understanding of the target topic. |
WHoW: A Cross-domain Approach for Analysing Conversation Moderation (2025.naacl-long)
Copied to clipboard
| Challenge: | Using this framework, we annotated 5,657 sentences with human judges and 15,494 sentences with GPT-4o from two domains: TV debates and radio panel discussions. |
| Approach: | They propose an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who). |
| Outcome: | The framework is generalisable across domains and reveals distinct modes of moderation: debate moderators emphasise coordination and facilitate interaction through questions and instructions, panel discussion moderator prioritize information provision and actively participate in discussions. |
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism (2024.acl-long)
Copied to clipboard
Miao Li, Ming-Bin Chen, Bo Tang, ShengbinHou ShengbinHou, Pengyu Wang, Haiying Deng, Zhiyu Li, Feiyu Xiong, Keming Mao, Cheng Peng, Yi Luo
| Challenge: | a novel evaluation framework assesses the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. |
| Approach: | They propose to use a benchmark dataset to assess the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. |
| Outcome: | The proposed evaluation framework is based on a dataset of 1,267 test samples in 24 news domains. |
Moderation Matters: Measuring Conversational Moderation Impact in English as a Second Language Group Discussion (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing tools for ESL assessment focus on writing skills and lack in support for dynamic spoken interactions. |
| Approach: | They propose an approach that integrates automatic ESL dialogue assessment and a framework that categorizes moderation strategies to assess conversational engagement and moderation effectiveness. |
| Outcome: | The proposed approach integrates automatic ESL dialogue assessment and categorizes moderation strategies. |