Papers by Arda Uzunoğlu
PARADISE: Evaluating Implicit Planning Skills of Language Models with Procedural Warnings and Tips Dataset (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have focused on whether large language models are capable of planning or executing plans. |
| Approach: | They propose an abductive reasoning task using wikiHow to test the effectiveness of small models over large models. |
| Outcome: | The proposed task demonstrates the effectiveness of small models over large models in most scenarios. |