Papers by Bolei Ma
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Using an annotation instrument, the design of the annotation instrument and the instructions given to annotators can impact training data. |
| Approach: | They investigate the impact of an annotation instrument on training data . they collect hate speech and offensive language annotations in a tweet corpus . |
| Outcome: | The proposed model performs better on holdout conditions than on the standard model. |
“My Answer is C”: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models (2024.findings-acl)
Copied to clipboard
Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Röttger, Frauke Kreuter, Dirk Hovy, Barbara Plank
| Challenge: | Multiple choice questions are one of the most popular evaluation formats for understanding the capabilities of autoregressive large language models (LLMs). |
| Approach: | They evaluated how aligned first-token evaluation is with the text output along several dimensions, namely final option choice, refusal rate, choice distribution and robustness under prompt perturbation. |
| Outcome: | The proposed evaluation methods are misaligned on all dimensions, reaching mismatch rates over 60%. |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |
FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to correct factually inaccurate outputs are lacking the semantic richness needed to properly understand its internal states of trustworthiness and honesty. |
| Approach: | They propose a framework for factuality alignment that integrates natural-language uncertainty signals with external knowledge and computes confidence scores and semantic entropy from LLM outputs. |
| Outcome: | Extensive experiments on four knowledge-intensive benchmarks show that FAITH improves the factual accuracy and truthfulness of Large Language Models (LLMs). |
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation (2026.eacl-long)
Copied to clipboard
Bolei Ma, Yong Cao, Indira Sen, Anna-Carolina Haensch, Frauke Kreuter, Barbara Plank, Daniel Hershcovich
| Challenge: | Large Language Models (LLMs) are increasingly used to simulate public opinion and other social phenomena. |
| Approach: | They argue that open-endedness is essential for realistic social simulations . they argue that it captures expressiveness and individuality . |
| Outcome: | The proposed frameworks can improve measurement and design, support exploration of unanticipated views, and reduce researcher-imposed directive bias. |
M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysis (2025.emnlp-main)
Copied to clipboard
ChengYan Wu, Bolei Ma, Yihong Liu, Zheyu Zhang, Ningyuan Deng, Yanshu Li, Baolan Chen, Yi Zhang, Yun Xue, Barbara Plank
| Challenge: | Existing studies focus on English-centric aspects of sentiment analysis, limiting scope for multilingual evaluation and research. |
| Approach: | They propose to use a multilingual dataset to analyze aspects with associated sentiment elements in text. |
| Outcome: | The proposed dataset is the most extensive multilingual parallel dataset for ABSA to date. |
Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation (2026.acl-long)
Copied to clipboard
| Challenge: | Table Question Answering (TQA) aims to answer natural language questions using tabular data. |
| Approach: | They propose a systematic overview of TQA research using large language models and summarize available benchmarks based on task features. |
| Outcome: | The proposed framework provides a comprehensive overview of the current state of the art in the field of Table Question Answering. |
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models (2024.findings-emnlp)
Copied to clipboard
Bolei Ma, Xinpeng Wang, Tiancheng Hu, Anna-Carolina Haensch, Michael Hedderich, Barbara Plank, Frauke Kreuter
| Challenge: | Recent advances in Large Language Models have sparked interest in validating human-like cognitive-behavioral traits. |
| Approach: | They examine whether LLM outputs reflect human-like cognitive-behavioral traits . they find that measuring AOVs embedded within LLMs remains opaque . |
| Outcome: | The proposed model can be used to evaluate human-like cognitive-behavioral traits . the proposed model could be used in writing assistants and other applications . |
Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study (2025.acl-long)
Copied to clipboard
Bolei Ma, Berk Yoztyurk, Anna-Carolina Haensch, Xinpeng Wang, Markus Herklotz, Frauke Kreuter, Barbara Plank, Matthias Aßenmacher
| Challenge: | Recent advances in large language models have generated significant interest in their potential for synthetic data generation across various domains. |
| Approach: | They use open-ended survey data from the German Longitudinal Election Studies to prompt different LLMs to generate synthetic public opinions reflective of German subpopulations by incorporating demographic features into the persona prompts. |
| Outcome: | The LLM performs better for supporters of left-leaning parties like The Greens and The Left compared to other parties, and matches the least with the right-party AfD. |
Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Multimodal Emotion Recognition in Conversations (MERC) is a new way to enhance human-computer interaction. |
| Approach: | This survey offers a systematic overview of Multimodal Emotion Recognition in Conversations . it examines motivations, core tasks, representative methods, and evaluation strategies . |
| Outcome: | The survey examines the effectiveness of MERC and its evaluation strategies. |
Can Large Language Models Advance Crosswalks? The Case of Danish Occupation Codes (2025.naacl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are used to map classification systems to each other . however, their use is labor-intensive and requires domain expertise . |
| Approach: | They propose a prompt-based framework where LLMs perform similarity assessments between classification codes and identify final mappings through a guided decision process. |
| Outcome: | The proposed framework shows that LLMs perform better than the embedding-based framework in creating crosswalks. |
The Imperfective Paradox in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing models rely on surface-level probabilistic heuristics to grasp compositional semantics of events . authors: current open-weight models operate as predictive narrative engines rather than faithful reasoners . |
| Approach: | They propose a diagnostic dataset to probe the imperfective paradox . they uncover a pervasive Teleological Bias in open-weight models . |
| Outcome: | The proposed dataset reveals a pervasive Teleological Bias in open-weight models . the findings suggest that these models operate as predictive narrative engines rather than faithful reasoners . |
MSMO-ABSA: Multi-Scale and Multi-Objective Optimization for Cross-Lingual Aspect-Based Sentiment Analysis (2026.acl-long)
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) has seen success with English texts, but real-world social media interactions often involve multiple languages. |
| Approach: | They propose a framework for cross-lingual ABSA that incorporates code-switched bilingual sentences into the language discriminator and consistency training modules to enhance cross-linguistic alignment. |
| Outcome: | The proposed framework achieves cross-lingual sentence-level and aspect-level alignment, aligning features of aspect terms in different contextual environments. |
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks (2024.eacl-long)
Copied to clipboard
| Challenge: | Prompt-based methods have been successfully applied to multilingual pretrained language models for zero-shot cross-lingual understanding. |
| Approach: | They propose a prompt-based method for token-level sequence labeling tasks . they propose to decompose an input sentence into single tokens and apply one prompt template to each token. |
| Outcome: | The proposed method outperforms Vanilla fine-tuning and Prompt-Tuning in zero-shot cross-lingual transfer . the method also attains state-of-the-art performance when employed with the mT5 model . |
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly applied to creative domains, yet performance in classical Chinese poetry generation and evaluation remains poorly understood. |
| Approach: | They propose a framework that combines computational metrics, LLM-as-a-judge assessment, and human expert validation to evaluate large language models. |
| Outcome: | The proposed framework evaluates state-of-the-art LLMs across multiple dimensions of poetic quality in Tang poetry generation. |