Papers with FTC
Find-the-Common: A Benchmark for Explaining Visual Patterns from Images (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in Instruction-fine-tuned Vision and Language Models (IVLMs) have prompted some studies to analyze the reasoning capabilities of IVLMs. |
| Approach: | They introduce a vision and language task for Inductive Visual Reasoning that uses common attributes across visual scenes to find common answers. |
| Outcome: | The proposed model can archive with 48% accuracy on the FTC, compared with state-of-the-art models. |