Papers by Unnat Jain
Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments (2021.emnlp-main)
Copied to clipboard
| Challenge: | Prior work on Vision-and-Language Navigation (VLN) tasks do not measure how much of a language instruction the agent is able to follow. |
| Approach: | They propose a language-aligned supervision scheme that measures the number of sub-instructions the agent has completed during navigation. |
| Outcome: | The proposed method is based on the previous work on the Vision-and-Language Navigation task, which assumes a discrete navigation graph (navgraph) but not on the current work. |
Compositional Reasoning via Joint Image and Language Decomposition (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing approaches typically decompose only language queries, treating images as monolithic inputs. |
| Approach: | They propose a framework that decomposes both images and questions into visual sub-domains with corresponding sub-questions. |
| Outcome: | REDI achieves absolute accuracy improvements of 8.9%, 8.2%, and 16.0% over existing models. |