Papers by Natasha Jaques
Human-centric dialog training via offline reinforcement learning (2020.emnlp-main)
Copied to clipboard
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, Rosalind Picard
| Challenge: | a novel offline RL method can train dialog models to produce better conversations without the risk of humans teaching it harmful chat behaviors. |
| Approach: | They develop offline reinforcement learning algorithms that use human feedback to train dialog models . they use language similarity, laughter, sentiment, and more to identify positive feedback . |
| Outcome: | The proposed method improves on existing methods with 80 users in an open-domain setting. |
Moral Foundations of Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Moral foundations theory (MFT) is a psychological assessment tool that decomposes human moral reasoning into five factors, including care/harm, liberty/oppression, and sanctity/degradation. |
| Approach: | They propose to use moral foundations theory to analyze whether popular LLMs have acquired a bias towards a particular set of moral values. |
| Outcome: | The proposed model can be adversarially selected to exhibit a particular moral foundations and can affect downstream tasks. |