Papers by Bolei Ma

15 papers
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance (2023.findings-emnlp)

Copied to clipboard

Challenge: Using an annotation instrument, the design of the annotation instrument and the instructions given to annotators can impact training data.
Approach: They investigate the impact of an annotation instrument on training data . they collect hate speech and offensive language annotations in a tweet corpus .
Outcome: The proposed model performs better on holdout conditions than on the standard model.
“My Answer is C”: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Multiple choice questions are one of the most popular evaluation formats for understanding the capabilities of autoregressive large language models (LLMs).
Approach: They evaluated how aligned first-token evaluation is with the text output along several dimensions, namely final option choice, refusal rate, choice distribution and robustness under prompt perturbation.
Outcome: The proposed evaluation methods are misaligned on all dimensions, reaching mismatch rates over 60%.
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)

Copied to clipboard

Challenge: linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions.
Approach: They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications.
Outcome: The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models .
FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to correct factually inaccurate outputs are lacking the semantic richness needed to properly understand its internal states of trustworthiness and honesty.
Approach: They propose a framework for factuality alignment that integrates natural-language uncertainty signals with external knowledge and computes confidence scores and semantic entropy from LLM outputs.
Outcome: Extensive experiments on four knowledge-intensive benchmarks show that FAITH improves the factual accuracy and truthfulness of Large Language Models (LLMs).
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to simulate public opinion and other social phenomena.
Approach: They argue that open-endedness is essential for realistic social simulations . they argue that it captures expressiveness and individuality .
Outcome: The proposed frameworks can improve measurement and design, support exploration of unanticipated views, and reduce researcher-imposed directive bias.
M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysis (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on English-centric aspects of sentiment analysis, limiting scope for multilingual evaluation and research.
Approach: They propose to use a multilingual dataset to analyze aspects with associated sentiment elements in text.
Outcome: The proposed dataset is the most extensive multilingual parallel dataset for ABSA to date.
Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Table Question Answering (TQA) aims to answer natural language questions using tabular data.
Approach: They propose a systematic overview of TQA research using large language models and summarize available benchmarks based on task features.
Outcome: The proposed framework provides a comprehensive overview of the current state of the art in the field of Table Question Answering.
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models have sparked interest in validating human-like cognitive-behavioral traits.
Approach: They examine whether LLM outputs reflect human-like cognitive-behavioral traits . they find that measuring AOVs embedded within LLMs remains opaque .
Outcome: The proposed model can be used to evaluate human-like cognitive-behavioral traits . the proposed model could be used in writing assistants and other applications .
Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have generated significant interest in their potential for synthetic data generation across various domains.
Approach: They use open-ended survey data from the German Longitudinal Election Studies to prompt different LLMs to generate synthetic public opinions reflective of German subpopulations by incorporating demographic features into the persona prompts.
Outcome: The LLM performs better for supporters of left-leaning parties like The Greens and The Left compared to other parties, and matches the least with the right-party AfD.
Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal Emotion Recognition in Conversations (MERC) is a new way to enhance human-computer interaction.
Approach: This survey offers a systematic overview of Multimodal Emotion Recognition in Conversations . it examines motivations, core tasks, representative methods, and evaluation strategies .
Outcome: The survey examines the effectiveness of MERC and its evaluation strategies.
Can Large Language Models Advance Crosswalks? The Case of Danish Occupation Codes (2025.naacl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used to map classification systems to each other . however, their use is labor-intensive and requires domain expertise .
Approach: They propose a prompt-based framework where LLMs perform similarity assessments between classification codes and identify final mappings through a guided decision process.
Outcome: The proposed framework shows that LLMs perform better than the embedding-based framework in creating crosswalks.
The Imperfective Paradox in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing models rely on surface-level probabilistic heuristics to grasp compositional semantics of events . authors: current open-weight models operate as predictive narrative engines rather than faithful reasoners .
Approach: They propose a diagnostic dataset to probe the imperfective paradox . they uncover a pervasive Teleological Bias in open-weight models .
Outcome: The proposed dataset reveals a pervasive Teleological Bias in open-weight models . the findings suggest that these models operate as predictive narrative engines rather than faithful reasoners .
MSMO-ABSA: Multi-Scale and Multi-Objective Optimization for Cross-Lingual Aspect-Based Sentiment Analysis (2026.acl-long)

Copied to clipboard

Challenge: Aspect-based sentiment analysis (ABSA) has seen success with English texts, but real-world social media interactions often involve multiple languages.
Approach: They propose a framework for cross-lingual ABSA that incorporates code-switched bilingual sentences into the language discriminator and consistency training modules to enhance cross-linguistic alignment.
Outcome: The proposed framework achieves cross-lingual sentence-level and aspect-level alignment, aligning features of aspect terms in different contextual environments.
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks (2024.eacl-long)

Copied to clipboard

Challenge: Prompt-based methods have been successfully applied to multilingual pretrained language models for zero-shot cross-lingual understanding.
Approach: They propose a prompt-based method for token-level sequence labeling tasks . they propose to decompose an input sentence into single tokens and apply one prompt template to each token.
Outcome: The proposed method outperforms Vanilla fine-tuning and Prompt-Tuning in zero-shot cross-lingual transfer . the method also attains state-of-the-art performance when employed with the mT5 model .
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly applied to creative domains, yet performance in classical Chinese poetry generation and evaluation remains poorly understood.
Approach: They propose a framework that combines computational metrics, LLM-as-a-judge assessment, and human expert validation to evaluate large language models.
Outcome: The proposed framework evaluates state-of-the-art LLMs across multiple dimensions of poetic quality in Tang poetry generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations