Papers by Fereshte Khani
Targeted Data Generation: Finding and Fixing Model Weaknesses (2023.acl-long)
Copied to clipboard
| Challenge: | Existing models fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust. |
| Approach: | They propose a framework that automatically identifies challenging subgroups and generates new data for those subgroup using large language models with a human in the loop. |
| Outcome: | The proposed framework improves accuracy on challenging subgroups while improving overall test accuracy. |
Prompt Engineering a Prompt Engineer (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent studies indicate that large language models can be meta-prompted to perform automatic prompt engineering, but their potential is limited due to insufficient guidance for complex reasoning in the meta-prompt. |
| Approach: | They propose to infuse three key components into a meta-prompt to guide reasoning . they find prompts that outperform “let’s think step by step” by 6.3% on MultiArith and 3.1% on GSM8K . |
| Outcome: | The proposed method outperforms “let’s think step by step” by 6.3% on MultiArith and 3.1% on GSM8K and outperfies baselines on counterfactual tasks by 6.9%. |