Papers by Sheela Agarwal
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings (2026.eacl-industry)
Copied to clipboard
Jean-Philippe Corbeil, Minseon Kim, Maxime Griot, Sheela Agarwal, Alessandro Sordoni, Francois Beaulieu, Paul Vozila
| Challenge: | Existing risk evaluations focused on general safety benchmarks, resulting in role-dependent vulnerabilities in real-world medical and clinical deployments. |
| Approach: | They propose a patient-oriented dataset called PatientSafetyBench that evaluates a variety of open- and closed-source LLMs. |
| Outcome: | The proposed benchmark examines medical risks from 466 open- and closed-source LLMs across 5 risk categories. |