Papers by Divyansh Kaushik
Dynabench: Rethinking Benchmarking in NLP (2021.naacl-main)
Copied to clipboard
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams
| Challenge: | Dynabench is an open-source platform for dynamic dataset creation and model benchmarking. |
| Approach: | They propose an open-source platform for dynamic dataset creation and model benchmarking. |
| Outcome: | The proposed platform can be used to create models that fail on simple challenges and falter in real-world scenarios. |
How Much Reading Does Reading Comprehension Require? A Critical Investigation of Popular Benchmarks (D18-1)
Copied to clipboard
| Challenge: | Recent research addresses reading comprehension, where examples consist of (question, passage, answer) tuples. |
| Approach: | They establish sensible baselines for bAbI, SQuAD, CBT, CNN and Who-did-What datasets and compare them to their previous work. |
| Outcome: | The proposed models perform on 14 out of 20 bAbI, SQuAD, CBT, CNN and Who-did-What datasets. |
On the Efficacy of Adversarial Data Collection for Question Answering: Results from a Large-Scale Randomized Study (2021.acl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that adversarial data collection (ADC) models perform better on other adversarially collected data but are liable under plausible domain shifts. |
| Approach: | They conduct a large-scale controlled study on question answering by assigning workers at random to compose questions either adversarially (with a model in the loop) or in the standard fashion (without a modeling). |
| Outcome: | The proposed model performs better on other adversarial datasets but worse on diverse collection of out-of-domain evaluation sets. |