Challenge: Large language models' confidence scores are degraded after fine-tuning with reinforcement learning from human feedback.
Approach: They propose a post-hoc calibration method that predicts a temperature scaling parameter for each token prediction.
Outcome: Adaptive temperature scaling improves calibration by over 10% compared to prior methods . RLHF fine-tuning improves model accuracy, but degradation is not significant .

Similar Papers

Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that unsupervised pre-training produces large language models whose conditional probabilities are remarkably well-calibrated.
Approach: They propose to use verbalized confidences to extract confidence from large language models with reinforcement learning from human feedback to improve their accuracy.
Outcome: The proposed methods reduce the expected calibration error by 50% for RLHF-LMs such as ChatGPT, GPT-4, and Claude.
A Survey of Post-Training Scaling in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated proficiency in understanding and generating human natural languages.
Approach: They propose a framework for scaling large language models using supervised fine-tuning, RLxF and test-time compute methodologies.
Outcome: The proposed model can be used to understand and generate human natural languages.
Balancing Classification and Calibration Performance in Decision-Making LLMs via Calibration Aware Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly deployed in decision-making tasks where accuracy and reliable confidence estimates are essential.
Approach: They propose a calibration-aware reinforcement learning formulation that directly adjusts decision-token probabilities.
Outcome: The proposed model preserves RLVR’s accuracy level while mitigating overconfidence, reducing ECE scores up to 9 points.
A Survey of Confidence Estimation and Calibration in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive capabilities across a wide range of tasks in various domains, but they can be unreliable due to factual errors in their generations.
Approach: They summarize recent advances in LLM confidence estimation and calibration and outline their main lessons learned.
Outcome: The proposed methods can be used to assess the reliability of models and to calibrate them across tasks.
Enhancing Language Model Alignment: A Confidence-Based Approach to Label Smoothing (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have remarkable capabilities across various domains . Reinforcement Learning with Human Feedback (RLHF) phase is crucial for training . label smoothing is a technique that replaces hard labels with soft labels .
Approach: They propose a method that iteratively updates the label smoothing parameter based on preference labels and model forecasts.
Outcome: The proposed method improves the performance of large language models on state-of-the-art alignment tasks.
A Close Look into the Calibration of Pre-trained Language Models (2023.acl-long)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) may fail in giving reliable estimates of their predictive uncertainty.
Approach: They conduct fine-grained control experiments to study the dynamic change in PLMs’ calibration performance in training.
Outcome: The proposed methods significantly reduce PLMs’ confidence in wrong predictions.
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning (2026.acl-long)

Copied to clipboard

Challenge: elucidating scaling laws for large language models (LLMs) during pre-training remains unexplored.
Approach: They characterize how model scale, data, and compute interact during pre-training . they find that large models consistently demonstrate superior compute and data efficiency .
Outcome: The proposed scaling laws offer practical guidance for scaling reasoning capabilities through reinforcement learning post-training.
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems (2025.emnlp-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable performance across a wide range of industrial applications.
Approach: They propose two techniques for training and deploying small language models that deliver high performance for a variety of industry use cases.
Outcome: The proposed techniques retain much of the quality of larger models while reducing training/serving costs and latency.
Calibration Across Layers: Understanding Calibration Evolution in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated inherent calibration capabilities, where predicted probabilities align well with correctness . previous studies have linked this behavior to specific components in the final layer, such as entropy neurons and the unembedding matrix’s null space.
Approach: They propose to examine how calibration evolves throughout the network's depth.
Outcome: The proposed calibration direction improves calibration metrics without harming accuracy.
LRQuant: Learnable and Robust Post-Training Quantization for Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for post-training quantization (PTQ) are limited by the complexity of the quantization parameter and performance degradations when tested on unseen datasets.
Approach: They propose a learnable smooth-based PTQ framework that allows for rapid adaptation during testing.
Outcome: The proposed framework improves performance on unseen datasets and reduces memory constraints.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations