Papers by Luciano Corro
Automatic Pair Construction for Contrastive Post-training (2024.findings-naacl)
Copied to clipboard
Canwen Xu, Corby Rosset, Ethan Chau, Luciano Corro, Shweti Mahajan, Julian McAuley, Jennifer Neville, Ahmed Awadallah, Nikhil Rao
| Challenge: | Large language models (LLMs) have unprecedented proficiency in a wide array of tasks. |
| Approach: | They propose a way to construct contrastive data using preference pairs from multiple models of varying strengths using SLiC and DPO. |
| Outcome: | The proposed method outperforms existing models like Orca in the comparison of SLiC and DPO with SFT baselines. |
The Greatest Good Benchmark: Measuring LLMs’ Alignment with Utilitarian Moral Dilemmas (2024.emnlp-main)
Copied to clipboard
| Challenge: | Our analysis across 15 diverse LLMs reveals consistently encoded moral preferences that diverge from established moral theories and lay population moral standards. |
| Approach: | They propose to evaluate the moral judgments of large language models using utilitarian dilemmas to determine their moral alignment. |
| Outcome: | The findings highlight the ‘artificial moral compass’ of Large Language Models, offering insights into their moral alignment. |