Papers by Ziquan Liu
Confidence Should Be Calibrated More Than One Turn Deep (2026.acl-long)
Copied to clipboard
| Challenge: | Existing work on confidence estimation and calibration focuses on single-turn settings . existing work on multi-turn calibration ignores the risks and potential of multi-turned conversations . |
| Approach: | They propose a multi-turn calibration task that reframes calibration from a static property into a dynamic challenge central to reliable multi- turn conversations. |
| Outcome: | The proposed model minimizes ECE@T and leverages ConfChat to improve confidence . the proposed model preserves and even enhances model performance in multi-turn interactions. |
GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for assessing the reliability of Large Language Models (LLMs) by confidence elicitation require expensive computational overhead or suffer from poor calibration, making them unreliable for real-world deployment. |
| Approach: | They propose a Generative Approach to Confidence Elicitation that enables reliable confidence elicitation for Large Language Models. |
| Outcome: | The proposed method achieves the best discriminative capacity and calibration on open-ended tasks without resorting to additional sampling or an auxiliary model. |
Get Confused Cautiously: Textual Sequence Memorization Erasure with Selective Entropy Maximization (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for erasure of memorized text fail to unlearn large numbers of memorizable samples without jeopardizing model utility. |
| Approach: | They propose a method that allows LLMs to memorize and recite some training sequences verbatim . they propose an entropy-based loss method that is shown to be more stable . |
| Outcome: | The proposed method improves model utility and accuracy while preserving model ability in language generation and understanding. |
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede’s Cultural Dimensions (2025.coling-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are deployed in many countries, but they fail to account for cultural variances among their potential users. |
| Approach: | They propose to use Hofstede’s cultural dimension framework to quantify cultural alignment using latent variable analysis to evaluate large language models against cultural dimensions of regions like the United States, China, and Arab countries. |
| Outcome: | The proposed model is compared against LLMs in the United States, China, and Arab countries and demonstrates that all models struggle to grasp cultural values, while GPT-4 shows a unique capability to adapt to cultural nuances, particularly in Chinese settings. |