Challenge: Existing zero-shot (ZS) approaches emphasize human motion while underutilizing contextual information, particularly human–object interactions.
Approach: They propose a framework for ZS recognition and zero-to-few-shot adaptation that leverages instance-level language descriptions.
Outcome: The proposed framework outperforms keypoint-based ZS methods while remaining data-efficient and robust.

Similar Papers

Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspective (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to cross-lingual Named Entity Recognition focus on Latin script language (LSL) for non-Latin script language, performance often degrades due to deep structural differences.
Approach: They propose an entity-aligned translation approach to align entities between NSL and English .
Outcome: The proposed approach aims to transfer knowledge from high-resource languages to low-resourced languages.
Decoupling Structure and Lexicon for Zero-Shot Semantic Parsing (D18-1)

Copied to clipboard

Challenge: Existing methods for training semantic parsers in new domains require expensive supervision and lack the ability to generalize to new domain.
Approach: They propose a zero-shot approach to parsing utterances in unseen domains . they map an utterant to an abstract, domain independent, logical form and replace slots with KB constants based on lexical alignment scores and global inference .
Outcome: The proposed model achieves 53.4% accuracy on 7 domains in the OVERNIGHT dataset, significantly better than other zero-shot baselines and performs as good as a parser trained on over 30% of the target domain examples.
OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation Models (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for object navigation are limited to household datasets with close-set objects, and they lack the ability to generalize to new environments in a zero-shot manner.
Approach: They propose a framework that leverages reasoning abilities of large vision language models to extract proposed objects from natural language instructions that meet the user’s demand.
Outcome: The proposed framework surpasses baselines on all metrics and can be used in a HM3D ObjectNav benchmark.
AlignRE: An Encoding and Semantic Alignment Approach for Zero-Shot Relation Extraction (2024.findings-acl)

Copied to clipboard

Challenge: Existing prototype-based methods for ZSRE ignore abundant side information and suffer from a significant encoding gap between prototypes and sentences.
Approach: They propose a framework to encode schema alignment to enhance prototype-based ZSRE methods.
Outcome: The proposed method outperforms existing methods on FewRel and Wiki-ZSL datasets and exhibits substantially faster performance and reduces the need for extensive manual labor in prototype construction.
ZEROTOP: Zero-Shot Task-Oriented Semantic Parsing using Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing LLMs cannot generalize to domain-specific parsing tasks in a zero-shot setting.
Approach: They propose a task-oriented parsing method that decomposes parse problem into abstractive and extractive question-answering problems.
Outcome: The proposed method decomposes a parsing problem into abstractive and extractive question-answering (QA) problems.
Text2Model: Text-based Model Induction for Zero-shot Image Classification (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to zero-shot learning are limited in two ways: Query-dependence and richness of language description.
Approach: They propose a task-agnostic approach to image classification using only text descriptions . they train a hypernetwork that receives class descriptions and outputs a multi-class model .
Outcome: The proposed approach generates non-linear classifiers, handles rich textual descriptions, and may be adapted to produce lightweight models efficient enough for on-device applications.
Z-GMOT: Zero-shot Generic Multiple Object Tracking (2024.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to Multi-Object Tracking (MOT) rely on initial bounding boxes and struggle with unseen objects.
Approach: They propose a cutting-edge multi-object tracking solution that can track unseen objects . they propose iGLIP and MA-SORT, which integrate motion and appearance matching strategies .
Outcome: The proposed solution can track objects from never-seen categories without initial bounding boxes or predefined categories.
LLM-Driven Implicit Target Augmentation and Fine-Grained Contextual Modeling for Zero-Shot and Few-Shot Stance Detection (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies on zero-shot and few-shot stance detection neglect implicit yet semantically important targets.
Approach: They propose a framework that uses Large Language Models to annotate implicit targets . they also propose 'DyMCA' to dynamically adjust text-target contributions based on context .
Outcome: The proposed framework achieves state-of-the-art on a benchmark dataset.
Look-up and Adapt: A One-shot Semantic Parser (D19-1)

Copied to clipboard

Challenge: Current conversational agents such as Siri, Alexa or Google Assistant do not cater to the specific phrasing of a user or the specific action.
Approach: They propose a semantic parser that generalizes to out-of-domain examples by adapting the logical forms of seen utterances to fit an unseen utterant.
Outcome: The proposed parser improves on one-shot parsing by 68.8% compared to baselines . it adapts the logical forms of seen utterances to fit the unseen utterant .
Zero-Shot Semantic Parsing for Instructions (P19-1)

Copied to clipboard

Challenge: Recent years have seen an increasing number of applications that have a natural language interface, such as chatbots or "intelligent personal assistants"
Approach: They propose a new training algorithm that trains a semantic parser on examples from a set of source domains and augment it with features and a logical form candidate filtering logic to support zero-shot adaptation.
Outcome: The proposed framework performs better than a non-adapted parser with features and logical form candidate filtering logic.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations