Papers by Prisha Samdarshi
Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game (2024.emnlp-main)
Copied to clipboard
Prisha Samdarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf, Tuhin Chakrabarty, Smaranda Muresan
| Challenge: | We evaluate the performance of large language models (LLMs) against expert and novice human players. |
| Approach: | They propose to use the New York Times Connections game as a test bed to evaluate the abstract reasoning capabilities of large language models (LLMs) they propose to test the ability of large-language models to be able to cluster and categorize words using semantic relations. |
| Outcome: | The proposed game is a test bed for evaluating abstract reasoning capabilities in humans and AI systems. |