Papers by Tristan Williams
Dynatask: A Framework for Creating Dynamic AI Benchmark Tasks (2022.acl-demo)
Copied to clipboard
Tristan Thrush, Kushal Tirumala, Anmol Gupta, Max Bartolo, Pedro Rodriguez, Tariq Kane, William Gaviria Rojas, Peter Mattson, Adina Williams, Douwe Kiela
| Challenge: | Open source system for setting up custom NLP tasks aims to lower technical knowledge and effort required for hosting and evaluating state-of-the-art models. |
| Approach: | They propose to integrate Dynatask with Dynabench to simplify benchmarking . they use a dataset to collect and clean data and train and evaluate models . |
| Outcome: | Dynatask is an open source system for setting up custom NLP tasks . it is integrated with Dynabench, a research platform for rethinking benchmarking in AI . |
Dynabench: Rethinking Benchmarking in NLP (2021.naacl-main)
Copied to clipboard
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams
| Challenge: | Dynabench is an open-source platform for dynamic dataset creation and model benchmarking. |
| Approach: | They propose an open-source platform for dynamic dataset creation and model benchmarking. |
| Outcome: | The proposed platform can be used to create models that fail on simple challenges and falter in real-world scenarios. |
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing work on marginal distributions and model steering fails to account for deeper latent structures that characterise real populations. |
| Approach: | They propose a framework for evaluating the representativeness of aligned models through multivariate correlation patterns in addition to marginal distributions. |
| Outcome: | The proposed framework compares two model steering techniques against human responses from the World Values Survey. |