Challenge: Existing approaches to extract spatial knowledge focus on extracting locations of events, someone or something.
Approach: They propose a method to annotate temporally-anchored spatial knowledge on top of OntoNotes by crowdsourcing annotations.
Outcome: The proposed method can be automated and validated using syntactic dependencies and crowdsourced annotations.

Similar Papers

Spatial and Temporal Language Understanding: Representation, Reasoning, and Grounding (2024.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
Approach: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
Outcome: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
Annotating Temporal Dependency Graphs via Crowdsourcing (2020.emnlp-main)

Copied to clipboard

Challenge: Existing temporal annotation schemes have been limited due to the complexity of temporal relations between events.
Approach: They propose to build a corpus of Wikinews articles annotated with temporal dependency graphs . they also propose a crowdsourcing strategy to annotate TDGs based on the corpus .
Outcome: The proposed method achieves a good trade-off between completeness and practicality in temporal annotation.
Spatial AMR: Expanded Spatial Annotation in the Context of a Grounded Minecraft Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing annotation tools for spatial relations capture fine-grained semantics and pragmatics derived from spatial information.
Approach: They propose an extension to the Abstract Meaning Representation annotation schema that captures fine-grained spatial information in grounded corpora.
Outcome: The proposed tool can handle fine-grained spatial relationships grounded in quantized space.
NarrativeTime: Dense Temporal Annotation on a Timeline (2024.lrec-main)

Copied to clipboard

Challenge: e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text.
Approach: They propose a timeline-based framework that achieves full coverage of all possible TLINKs.
Outcome: The proposed framework achieves full coverage of all possible TLINKs in a text.
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work on temporal sentence grounding rely on expensive video-query paired annotations . despite this, there are no ground-truth annotations in the current work .
Approach: They propose to use paired video-query and segment boundary annotations to generate temporal sentence grounding without training.
Outcome: The proposed model outperforms existing unsupervised methods and beats supervised ones on two challenging datasets.
Robust and Interpretable Grounding of Spatial References with Relation Networks (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models for understanding spatial references in text are vulnerable to noise in input text or state observations.
Approach: They propose a text-conditioned relation network with a cross-modal attention module to capture fine-grained spatial relations between entities and a model that is robust and interpretable.
Outcome: The proposed model improves performance on three tasks with a 17% improvement in predicting goal locations and a 15% improvement in robustness compared to state-of-the-art systems.
BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded Dialogues (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to video-grounded dialogues focus on superficial temporal-level visual cues, but neglect more fine-grained spatial signals from videos.
Approach: They propose a vision-language neural framework for high-resolution queries in videos based on textual cues that exploits both spatial and temporal-level information.
Outcome: The proposed approach outperforms previous approaches on the TGIF-QA benchmark and significantly outperformed previous approaches.
Into the Unknown: Generating Geospatial Descriptions for New Environments (2024.findings-acl)

Copied to clipboard

Challenge: Similar to vision-and-language navigation tasks, the Rendezvous (RVS) task requires reasoning over allocentric spatial relationships using non-sequential navigation instructions and maps.
Approach: They propose a large-scale augmentation method for generating high-quality synthetic data for new environments using readily available geospatial data.
Outcome: The proposed method improves accuracy on unseen and seen environments by 45.83% on the Rendezvous (RVS) task.
Comprehensive Annotation of Various Types of Temporal Information on the Time Axis (L18-1)

Copied to clipboard

Challenge: Existing studies linking event and time information have been conducted to train and evaluate models.
Approach: They propose an annotation scheme that anchors expressions in text to the time axis comprehensively.
Outcome: The proposed scheme can be utilized for integrated information analysis of events, entities and time.
STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing models focus on predictive accuracy over reasoning, a gap exists . time series data are ubiquitous in real-world systems and exhibit complex spatio-temporal structures.
Approach: They propose a time series reasoning model that integrates time series, graph structure, and text for explicit reasoning.
Outcome: The proposed model achieves average accuracy gains between 17% and 135% at 0.004x the cost of proprietary models and generalizes robustly to real-world data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations