Papers with Ensuring
Investigating the Multilingual Calibration Effects of Language Model Instruction Tuning (2026.eacl-short)
Copied to clipboard
Jerry Huang, Peng Lu, Qiuhao Zeng, Yusuke Iwasawa, Yutaka Matsuo, Sarath Chandar, Edison Marrese-Taylor, Irene Li
| Challenge: | despite advances in foundation model research, the relationship between large language models and their calibration remains an open area of research. |
| Approach: | They examine a gap in the calibration of large language models within multilingual settings to better understand how data scarcity can potentially lead to different calibration effects. |
| Outcome: | The proposed calibration gap is found in two multilingual benchmarks over 29 and 42 languages. |
Studying and Mitigating Biases in Sign Language Understanding Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using crowd-sourced sign language datasets to reduce performance disparities is critical to addressing potential biases and inequities. |
| Approach: | They use demographic information to study biases that may result from models trained on crowd-sourced sign datasets. |
| Outcome: | The proposed approach reduces performance disparities without decreasing accuracy. |
Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety (2026.acl-long)
Copied to clipboard
Can Jin, Rui Wu, Tong Che, Qixin Zhang, Hongwu Peng, Jiahui Zhao, Zhenting Wang, Wenqi Wei, Ligong Han, Zhao Zhang, Yuan Cao, Ruixiang Tang, Dimitris N. Metaxas
| Challenge: | OpenAI introduces deliberative alignment (DA) to enhance safety of its o-series models, but effectiveness of this approach in open-source LLMs is understudied. |
| Approach: | They propose a case-augmented deliberative alignment method for large language models . they propose to use reinforcement learning on self-generated safety reasoning chains . |
| Outcome: | The proposed method avoids narrowly enumerated rules and allows broader adaptability. |
Adaptive Helpfulness–Harmlessness Alignment with Preference Vectors (2026.eacl-long)
Copied to clipboard
Ren-Wei Liang, Chin Ting Hsu, Chan-Hung Yu, Saransh Agrawal, Shih-Cheng Huang, Chieh-Yen Lin, Shang-Tse Chen, Kuan-Hao Huang, Shao-Hua Sun
| Challenge: | Existing approaches to balancing helpfulness and harmlessness suffer from performance conflicts, limited controllability, and poor extendability. |
| Approach: | They propose a framework that allows users to control their own preferences and dynamically merge them at test time. |
| Outcome: | The proposed framework improves helpfulness without conservatism and smooth control over preference trade-offs. |
LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine Feedback (2024.acl-long)
Copied to clipboard
| Challenge: | Creating trusted and safe online spaces for people with different backgrounds and opinions is a challenge for social media platforms. |
| Approach: | They propose a reinforcement learning-based rewriting approach that balances content preservation and appropriateness based on existing classifiers. |
| Outcome: | The proposed approach significantly outperforms baselines including few-shot learning, prompting, and humans. |
An Active Learning Framework for Inclusive Generation by Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit bias toward underrepresented groups, despite advances in active learning. |
| Approach: | They propose a clustering-based active learning framework enhanced with knowledge distillation that transforms the intermediate outputs of the learner model to yield more representative models without prior knowledge of underlying data distribution. |
| Outcome: | The proposed framework improves performance across data subgroups and lexical diversity, underscoring the model’s resilience to skewness in available data. |
Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalization (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Prior research has shown that LLMs fail to perform satisfactorily on moral cognizance tasks . |
| Approach: | They propose to use curated datasets to improve LLMs' moral cognizance . they find pragmatic dilemma constrains generalization ability of current learning paradigms . |
| Outcome: | The proposed learning paradigms fail to perform on moral cognizance tasks, the authors show . they show that the pragmatic dilemma is the primary bottleneck for moral reasoning acquisition . |