Papers by Yatai Ji
CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space (2025.emnlp-main)
Copied to clipboard
Yong Zhao, Kai Xu, Zhengqiu Zhu, Yue Hu, Zhiheng Zheng, Yingfeng Chen, Yatai Ji, Chen Gao, Yong Li, Jincai Huang
| Challenge: | Embodied Question Answering (EQA) tasks are primarily focused on indoor environments, leaving the complexities of urban settings unexplored. |
| Approach: | They propose a task where an embodied agent answers open-vocabulary questions in dynamic city spaces. |
| Outcome: | The proposed agent achieves 60.7% of human-level answering accuracy compared to baselines . the proposed agent outperforms existing agents in open-ended city spaces . |
MIRTT: Learning Multimodal Interaction Representations from Trilinear Transformers for Visual Question Answering (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing bilinear methods focus on inter-modality information between images and questions . existing models focus on the interaction between images, questions, and images . |
| Approach: | They propose a trilinear interaction framework that incorporates attention mechanisms for capturing inter-modality and intra-modal relationships. |
| Outcome: | The proposed model outperforms bilinear models on the Visual7W Telling task and VQA-1.0 Multiple Choice task and outperformed baselines on the VQA, TDIUC and GQA datasets. |