Papers by Chad DeLuca
STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks often treat complex tasks as monolithic, resulting in inconsistent performance and inconsistent explanations. |
| Approach: | They propose a framework for creating controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner. |
| Outcome: | The proposed framework enables systematic probing of model behavior by identifying the specific reasoning skill compositions they lack. |
Don’t be my Doctor! Recognizing Healthcare Advice in Large Language Models (2024.emnlp-industry)
Copied to clipboard
| Challenge: | Large language models (LLMs) are becoming increasingly popular in everyday use, especially in highly regulated domains such as healthcare, where misleading advice may influence users to commit malpractice. |
| Approach: | They present a large-scale health-advice benchmark dataset that evaluates large language models' ability to recognize health-related advice in industrial settings. |
| Outcome: | The proposed model can be misinterpreted as direct advice in highly regulated domains such as healthcare, but the results are not enough to protect them from misinterpreting them as medical advice. |