Papers with QM
DuQM: A Chinese Dataset of Linguistically Perturbed Natural Questions for Evaluating the Robustness of Question Matching Models (2022.emnlp-main)
Copied to clipboard
| Challenge: | a comprehensive evaluation of QM models should be conducted on natural texts, not on artificial adversarial examples . ral models are often not robust to adversarials, which means they predict unexpected outputs . |
| Approach: | They use a Chinese dataset to evaluate the robustness of QM models . they show that the effect of artificial adversarial examples does not work on natural texts . |
| Outcome: | The proposed model is more robust than other models on natural questions with 32 linguistic perturbations. |
Leveraging In-Context Learning for Political Bias Testing of LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Existing probing methods for evaluating LLMs with political questions have limited stability and are unreliable. |
| Approach: | They propose to use human survey data as in-context examples to query LLMs with political questions to evaluate their potential biases. |
| Outcome: | The proposed task improves the stability of question-based bias evaluation and may be used to compare instruction-tuned models to their base versions. |