PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational Summarization (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models exhibit positional bias, a problem that can undermine the completeness of conversation summarizations. |
| Approach: | They propose a semantic similarity-based sentence-level metric to quantify positional bias in conversational summaries. |
| Outcome: | The proposed benchmark provides the first systematic evaluation of positional bias in conversational summarization across languages and contexts. |
Similar Papers
On Context Utilization in Summarization with Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Large language models excel in abstractive summarization tasks, delivering fluent and pertinent summaries. |
| Approach: | They conduct the first comprehensive study on context utilization and position bias in summarization. |
| Outcome: | The proposed benchmark compares two methods to alleviate position bias in summarization tasks. |
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias (2024.naacl-short)
Copied to clipboard
| Challenge: | Position bias is a tendency of a model unfairly prioritizing information from certain parts of the input text over others, leading to undesirable behavior. |
| Approach: | They propose to measure position bias in large language models for zero-shot summarization tasks by measuring position bias. |
| Outcome: | The proposed model performance and position biases lead to new insights and discussion on zero-shot summarization tasks. |
Characterizing Positional Bias in Large Language Models: A Multi-Model Evaluation of Prompt Order Effects (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models can be influenced by various forms of biases, says a new study . positional bias affects how LLMs interpret and weigh information, the authors say . |
| Approach: | a new study examines the impact of positional bias on large language models . positional biased models prioritize items based on their position rather than content or quality . |
| Outcome: | a new study shows that LLMs prioritize items based on their position rather than content or quality . the positional bias affects how LLM interpret and weigh information, the authors say . |
On Positional Bias of Faithfulness for Long-form Summarization (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models exhibit positional bias in long-context settings, under-attending to information in the middle. |
| Approach: | They compile eight human-annotated long-form summarization datasets to evaluate faithfulness . they find that LLMs faithfully summarize beginning and end of documents but neglect middle content . |
| Outcome: | The proposed methods show that LLMs under-attend to information in the middle of inputs. |
Spot the BlindSpots: Systematic Identification and Quantification of Fine-Grained LLM Biases in Contact Center Call Summarization (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Abstractive summarization is a core application in contact centers, where Large Language Models generate millions of summaries of call transcripts daily. |
| Approach: | They propose a framework that uses an LLM as a zero-shot classifier to derive categorical distributions for each bias dimension in a pair of transcripts and its summary. |
| Outcome: | The proposed framework identifies and quantifies 15 operational bias dimensions and measures them using two metrics: Fidelity Gap and Coverage. |
Not Lost After All: How Cross-Encoder Attribution Challenges Position Bias Assumptions in LLM Summarization (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Position bias is a key limitation in automatic summarization. |
| Approach: | They propose a cross-encoder-based alignment method that processes summary-source sentence pairs . |
| Outcome: | The proposed method allows better identification of semantic correspondences even when summaries substantially rewrite the source. |
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs (2025.naacl-long)
Copied to clipboard
Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan
| Challenge: | Using format-following capabilities, state-of-the-art large language models (LLMs) can be leveraged to tailor outputs to specific task formats. |
| Approach: | They propose to define a format bias evaluation metric and establish effective strategies to reduce it. |
| Outcome: | The proposed evaluation reduces the variance in ChatGPT’s performance among wrapping formats from 235.33 to 0.71 (%2) |
Bias in News Summarization: Measures, Pitfalls and Corpora (2024.findings-acl)
Copied to clipboard
| Challenge: | Pretrained large language models can reproduce harmful social biases in constrained settings, such as summarization. |
| Approach: | They propose a method to generate input documents with carefully controlled demographic attributes and then apply it to a controlled setting. |
| Outcome: | The proposed method allows to generate input documents with carefully controlled demographic attributes while working with real-world input documents. |
DateLogicQA: Benchmarking Temporal Biases in Large Language Models (2025.naacl-srw)
Copied to clipboard
| Challenge: | DateLogicQA examines temporal biases in Large Language Models (LLMs) 190 questions are curated by humans to examine temporal reasoning across date formats and contexts . |
| Approach: | They propose a human-curated benchmark of 190 questions specifically designed to understand temporal bias in Large Language Models. |
| Outcome: | The proposed dataset covers seven date formats across past, present, and future contexts . it examines four reasoning types: commonsense, factual, conceptual, and numerical . |
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models can adapt outputs to align with community-specific norms, perspectives and communication styles. |
| Approach: | They propose a benchmark to assess community-specific steering using contrasting reddit communities. |
| Outcome: | STEER-BENCH assesses how well large language models understand community-specific instructions, their resilience to adversarial steering attempts, and their ability to accurately represent cultural and ideological perspectives. |