Papers by Shane Storks

14 papers
In-Context Analogical Reasoning with Pre-Trained Language Models (2023.acl-long)

Copied to clipboard

Challenge: Analogical reasoning is a fundamental capacity of human cognition that allows us to reason abstractly about novel situations by relating them to past experiences.
Approach: They apply large pre-trained language models to visual Raven’s Progressive Matrices (RPM) and use language-based abstractions to support analogy in AI systems.
Outcome: The proposed language-based abstractions outperform human models on Raven’s Progressive Matrices and supervised vision-based methods.
DANLI: Deliberative Agent for Following Natural Language Instructions (2022.emnlp-main)

Copied to clipboard

Challenge: Recent work on embodied AI agents that can perform tasks by following human language instructions is limited by reactive methods, which are insufficient for long-horizon complex tasks.
Approach: They propose a neuro-symbolic deliberative agent that, while following language instructions, proactively applies reasoning and planning based on its neural and symbolic representations acquired from past experience.
Outcome: The proposed agent achieves greater than 70% improvement over reactive baselines on the challenging TEACh benchmark.
From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense Reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have shown impressive performance in various language tasks, but are prone to spurious correlations and illusory information.
Approach: They propose to use pre-trained language models to justify decisions with formalized, coherent reasoning chains.
Outcome: The proposed strategies improve coherence of rationalizations yielding state-of-the-art results on Tiered Reasoning for Intuitive Physics (TRIP).
Mind the Gap: How BabyLMs Learn Filler-Gap Dependencies (2025.emnlp-main)

Copied to clipboard

Challenge: a long-standing debate concerns whether the linguistic input children receive is sufficient to explain the grammatical knowledge they develop.
Approach: They evaluate baby language models trained on child-oriented input from the BabyLM Challenge and two base models trained in 10M and 100M tokens.
Outcome: The proposed models acquire filler-gap dependencies but fail to generalize or fully capture island constraints.
Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding (2021.findings-emnlp)

Copied to clipboard

Challenge: Large-scale, pre-trained language models (LMs) have achieved human-level performance on a breadth of language understanding tasks.
Approach: They propose a commonsense reasoning dataset with dense annotations that allows multi-tiered evaluation of machines’ reasoning process.
Outcome: The proposed model can achieve high end performance but struggle to support predictions with valid supporting evidence.
Can Foundation Models Watch, Talk and Guide You Step by Step to Make a Cake? (2023.findings-emnlp)

Copied to clipboard

Challenge: despite advances in AI, it remains a challenge to develop interactive task guidance systems that can offer situated, personalized guidance and assist humans in various tasks.
Approach: They propose to use a multimodal benchmark dataset to study whether interactive task guidance systems can be quickly adapted to perceptually enabled tasks.
Outcome: The proposed models demonstrate fair performances in some cases with no training . the results will provide a stepping stone for future work on situated task guidance .
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties (2024.emnlp-main)

Copied to clipboard

Challenge: Emergent In-context Learning on Videos induces in-contact learning over video and text . eILeV-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions.
Approach: They implement Emergent In-context Learning on Videos (EILeV) that induces in-contact learning over video and text by capturing key properties of pre-training data.
Outcome: The proposed training paradigm outperforms off-the-shelf VLMs in few-shot video narration for novel, rare actions.
Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent work has focused on layerwise interpretations, lacking fine-grained interpretation of specific features and their interaction.
Approach: They identify semantically coherent, context-consistent network components in large language models . they use sparse autoencoders to coactivate sparsity features from a handful of prompts .
Outcome: The proposed model can capture concepts and relations more comprehensively than individual features while maintaining specificity.
SafetyALFRED: Evaluating Safety-Conscious Planning of Vision Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety evaluations focus on hazard recognition through disembodied question answering (QA) settings, but lack a critical gap in evaluating an agent.
Approach: They evaluate multimodal large language models with six categories of kitchen hazards . they propose a safety-based approach that prioritizes multi-step corrective actions .
Outcome: The proposed model can recognize hazards in QA settings, but average mitigation success rates are low . the proposed model is based on the embodied agent benchmark ALFRED .
Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers (2021.findings-emnlp)

Copied to clipboard

Challenge: Large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, but statistical bias in benchmark data and probing studies has recently called into question their true capabilities.
Approach: They propose to evaluate systems through a measure of prediction coherence by using two existing language understanding benchmarks with different properties to demonstrate its versatility.
Outcome: The proposed evaluation framework is quick, effective, and versatile to provide insight into the coherence of machines’ predictions.
The Double Bind: Revisiting Preprinting and Peer Review Two Years After the Removal of the ACL Anonymity Period (2026.findings-acl)

Copied to clipboard

Challenge: ACL removed the anonymity period for conference submissions in February 2024 .
Approach: They track preprinting trends for 47k publications and analyze 1.9k peer reviews . they suggest improving visibility and investing in diversity initiatives .
Outcome: The proposed anonymity period was removed in 2024, but it was ineffective for underrepresented researchers . the authors suggest addressing D&I issues rather than implementing anonymity policies.
NLP Reproducibility For All: Understanding Experiences of Beginners (2023.acl-long)

Copied to clipboard

Challenge: a study with 93 students in an introductory NLP class shows that beginners' programming skill and comprehension of research papers have a limited impact on their effort spent completing the exercise.
Approach: a study conducted with 93 students in an introductory NLP course questioned them on their programming background and programming background.
Outcome: The results show that beginners' programming skill and comprehension of research papers have a limited impact on their effort spent on the exercise.
Discovering Properties of Inflectional Morphology in Neural Emergent Communication (2026.acl-long)

Copied to clipboard

Challenge: Emergent communication studies protocols developed between two or more deep neural network-based agents . common evaluation metrics for large-vocabulary setting are overly simplified .
Approach: They propose to reinterpret an EmCom setting by imposing a small-vocabulary constraint to simulate double articulation and formulating a novel setting analogous to naturalistic inflectional morphology.
Outcome: The proposed model favors protocols that represent attributes with unique characters and compose them syntactically.
Transparent and Coherent Procedural Mistake Detection (2025.emnlp-main)

Copied to clipboard

Challenge: Procedural mistake detection (PMD) is a problem of classifying whether a human user has successfully executed a task.
Approach: They extend PMD to require generating visual self-dialog rationales to inform decisions . they leverage a natural language inference model to formulate two automated metrics for coherence of generated rationale.
Outcome: The proposed model improves on a reframed task with a natural language inference model and a multi-faceted metrics visualization of common outcomes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations