Papers by Shane Storks
In-Context Analogical Reasoning with Pre-Trained Language Models (2023.acl-long)
Copied to clipboard
| Challenge: | Analogical reasoning is a fundamental capacity of human cognition that allows us to reason abstractly about novel situations by relating them to past experiences. |
| Approach: | They apply large pre-trained language models to visual Raven’s Progressive Matrices (RPM) and use language-based abstractions to support analogy in AI systems. |
| Outcome: | The proposed language-based abstractions outperform human models on Raven’s Progressive Matrices and supervised vision-based methods. |
DANLI: Deliberative Agent for Following Natural Language Instructions (2022.emnlp-main)
Copied to clipboard
Yichi Zhang, Jianing Yang, Jiayi Pan, Shane Storks, Nikhil Devraj, Ziqiao Ma, Keunwoo Yu, Yuwei Bao, Joyce Chai
| Challenge: | Recent work on embodied AI agents that can perform tasks by following human language instructions is limited by reactive methods, which are insufficient for long-horizon complex tasks. |
| Approach: | They propose a neuro-symbolic deliberative agent that, while following language instructions, proactively applies reasoning and planning based on its neural and symbolic representations acquired from past experience. |
| Outcome: | The proposed agent achieves greater than 70% improvement over reactive baselines on the challenging TEACh benchmark. |
From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense Reasoning (2023.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained language models have shown impressive performance in various language tasks, but are prone to spurious correlations and illusory information. |
| Approach: | They propose to use pre-trained language models to justify decisions with formalized, coherent reasoning chains. |
| Outcome: | The proposed strategies improve coherence of rationalizations yielding state-of-the-art results on Tiered Reasoning for Intuitive Physics (TRIP). |
Mind the Gap: How BabyLMs Learn Filler-Gap Dependencies (2025.emnlp-main)
Copied to clipboard
| Challenge: | a long-standing debate concerns whether the linguistic input children receive is sufficient to explain the grammatical knowledge they develop. |
| Approach: | They evaluate baby language models trained on child-oriented input from the BabyLM Challenge and two base models trained in 10M and 100M tokens. |
| Outcome: | The proposed models acquire filler-gap dependencies but fail to generalize or fully capture island constraints. |
Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale, pre-trained language models (LMs) have achieved human-level performance on a breadth of language understanding tasks. |
| Approach: | They propose a commonsense reasoning dataset with dense annotations that allows multi-tiered evaluation of machines’ reasoning process. |
| Outcome: | The proposed model can achieve high end performance but struggle to support predictions with valid supporting evidence. |
Can Foundation Models Watch, Talk and Guide You Step by Step to Make a Cake? (2023.findings-emnlp)
Copied to clipboard
Yuwei Bao, Keunwoo Yu, Yichi Zhang, Shane Storks, Itamar Bar-Yossef, Alex de la Iglesia, Megan Su, Xiao Zheng, Joyce Chai
| Challenge: | despite advances in AI, it remains a challenge to develop interactive task guidance systems that can offer situated, personalized guidance and assist humans in various tasks. |
| Approach: | They propose to use a multimodal benchmark dataset to study whether interactive task guidance systems can be quickly adapted to perceptually enabled tasks. |
| Outcome: | The proposed models demonstrate fair performances in some cases with no training . the results will provide a stepping stone for future work on situated task guidance . |
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties (2024.emnlp-main)
Copied to clipboard
| Challenge: | Emergent In-context Learning on Videos induces in-contact learning over video and text . eILeV-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions. |
| Approach: | They implement Emergent In-context Learning on Videos (EILeV) that induces in-contact learning over video and text by capturing key properties of pre-training data. |
| Outcome: | The proposed training paradigm outperforms off-the-shelf VLMs in few-shot video narration for novel, rare actions. |
Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models (2026.acl-long)
Copied to clipboard
Ruixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Mike Angstadt, Chandra Sripada, Joyce Chai
| Challenge: | Recent work has focused on layerwise interpretations, lacking fine-grained interpretation of specific features and their interaction. |
| Approach: | They identify semantically coherent, context-consistent network components in large language models . they use sparse autoencoders to coactivate sparsity features from a handful of prompts . |
| Outcome: | The proposed model can capture concepts and relations more comprehensively than individual features while maintaining specificity. |
SafetyALFRED: Evaluating Safety-Conscious Planning of Vision Language Models (2026.findings-acl)
Copied to clipboard
Josue Torres-Fonseca, Naihao Deng, Yinpei Dai, Shane Storks, Yichi Zhang, Rada Mihalcea, Casey Kennington, Joyce Chai
| Challenge: | Existing safety evaluations focus on hazard recognition through disembodied question answering (QA) settings, but lack a critical gap in evaluating an agent. |
| Approach: | They evaluate multimodal large language models with six categories of kitchen hazards . they propose a safety-based approach that prioritizes multi-step corrective actions . |
| Outcome: | The proposed model can recognize hazards in QA settings, but average mitigation success rates are low . the proposed model is based on the embodied agent benchmark ALFRED . |
Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, but statistical bias in benchmark data and probing studies has recently called into question their true capabilities. |
| Approach: | They propose to evaluate systems through a measure of prediction coherence by using two existing language understanding benchmarks with different properties to demonstrate its versatility. |
| Outcome: | The proposed evaluation framework is quick, effective, and versatile to provide insight into the coherence of machines’ predictions. |
The Double Bind: Revisiting Preprinting and Peer Review Two Years After the Removal of the ACL Anonymity Period (2026.findings-acl)
Copied to clipboard
| Challenge: | ACL removed the anonymity period for conference submissions in February 2024 . |
| Approach: | They track preprinting trends for 47k publications and analyze 1.9k peer reviews . they suggest improving visibility and investing in diversity initiatives . |
| Outcome: | The proposed anonymity period was removed in 2024, but it was ineffective for underrepresented researchers . the authors suggest addressing D&I issues rather than implementing anonymity policies. |
NLP Reproducibility For All: Understanding Experiences of Beginners (2023.acl-long)
Copied to clipboard
| Challenge: | a study with 93 students in an introductory NLP class shows that beginners' programming skill and comprehension of research papers have a limited impact on their effort spent completing the exercise. |
| Approach: | a study conducted with 93 students in an introductory NLP course questioned them on their programming background and programming background. |
| Outcome: | The results show that beginners' programming skill and comprehension of research papers have a limited impact on their effort spent on the exercise. |
Discovering Properties of Inflectional Morphology in Neural Emergent Communication (2026.acl-long)
Copied to clipboard
| Challenge: | Emergent communication studies protocols developed between two or more deep neural network-based agents . common evaluation metrics for large-vocabulary setting are overly simplified . |
| Approach: | They propose to reinterpret an EmCom setting by imposing a small-vocabulary constraint to simulate double articulation and formulating a novel setting analogous to naturalistic inflectional morphology. |
| Outcome: | The proposed model favors protocols that represent attributes with unique characters and compose them syntactically. |
Transparent and Coherent Procedural Mistake Detection (2025.emnlp-main)
Copied to clipboard
| Challenge: | Procedural mistake detection (PMD) is a problem of classifying whether a human user has successfully executed a task. |
| Approach: | They extend PMD to require generating visual self-dialog rationales to inform decisions . they leverage a natural language inference model to formulate two automated metrics for coherence of generated rationale. |
| Outcome: | The proposed model improves on a reframed task with a natural language inference model and a multi-faceted metrics visualization of common outcomes. |