Papers by Marc Bätje
Analyzing Effects of Learning Downstream Tasks on Moral Bias in Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning large language models replicate and perpetuate social biases . pre-existing moral bias may be mitigated or amplified even when presented with opposing views . |
| Approach: | They develop methods to assess the agreement of LMs to explicit codified norms . they find that introducing downstream tasks may lead to unexpected inconsistencies . |
| Outcome: | The proposed model can be used to improve morality in data-scarce tasks. |