Papers with DEFREASING
Evaluating Defeasible Reasoning in LLMs with DEFREASING (2025.naacl-long)
Copied to clipboard
| Challenge: | Defeasible inferences are highly plausible but can be impacted by new information. |
| Approach: | They construct a dataset to evaluate defeasible reasoning about property inheritance . they use generics to represent the inheritance rules because their semantics include exceptions . |
| Outcome: | The proposed model performs poorly across all pattern types and achieves 0.64 F 1 . the best performing model only achieves F 1 and the model is not well tuned . |