Papers by Victoria Li
ChatGPT Doesn’t Trust Chargers Fans: Guardrail Sensitivity in Context (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing work addresses the limitations of chatbot guardrails, which limit responses to uncertain or sensitive questions. |
| Approach: | They generate user biographies that offer ideological and demographic information about the user. |
| Outcome: | The proposed model can infer a likely political ideology and modify guardrail behavior accordingly. |