Papers by Ryota Tanaka
How Well Do Vision Models Encode Diagram Attributes? (2024.acl-srw)
Copied to clipboard
Haruto Yoshida, Keito Kudo, Yoichi Aoki, Ryota Tanaka, Itsumi Saito, Keisuke Sakaguchi, Kentaro Inui
| Challenge: | Experimental results show vision models struggle to identify diagram attributes such as node colors and shapes, along with edge colors and connection patterns. |
| Approach: | They evaluated vision models and retrieving diagrams using text queries to determine how well they recognize diagram attributes and edge connection patterns. |
| Outcome: | The models can recognize node colors, shapes, and edge colors, but struggle to identify differences in edge connection patterns that play a pivotal role in the semantics of diagrams. |