Papers by Ke-Jyun Wang
OCID-Ref: A 3D Robotic Dataset With Embodied Language For Clutter Scene Grounding (2021.naacl-main)
Copied to clipboard
| Challenge: | Visual grounding (VG) is a crucial task in natural language processing, computer vision, and robotics. |
| Approach: | They propose a visual grounding task with referring expressions of occluded objects in a OCID-Ref dataset with 2,300 scenes and a point cloud input. |
| Outcome: | The proposed dataset shows that it can handle 2D and 3D signals but referring to occluded objects remains challenging for the modern visual grounding systems. |