Papers by Rahul Madhavan
CARMO: Dynamic Criteria Generation for Context Aware Reward Modelling (2025.findings-acl)
Copied to clipboard
Taneesh Gupta, Shivam Shandilya, Xuchao Zhang, Rahul Madhavan, Supriyo Ghosh, Chetan Bansal, Huaxiu Yao, Saravan Rajmohan
| Challenge: | Reward modeling in large language models is susceptible to reward hacking . flawed reward signals often lead to outputs that optimize for spurious correlates . |
| Approach: | They propose a new approach that generates dynamic, context-relevant criteria to ground the reward model prior to producing reward scores. |
| Outcome: | The proposed approach generates dynamic, context-relevant criteria to ground the model prior to producing reward scores. |
CFL: Causally Fair Language Models Through Token-level Attribute Controlled Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to control attributes of Language Models (LMs) for text generation are not safe, as toxicity and bias goals are opposed to each other. |
| Approach: | They propose a method to control the attributes of Language Models (LMs) for the text generation task using Causal Average Treatment Effect (ATE) scores and counterfactual augmentation. |
| Outcome: | The proposed architecture achieves state of the art performance for toxic degeneration, which are computed using Real Toxicity Prompts. |