Papers by Daniel Flaherty
Re-evaluating the Need for Visual Signals in Unsupervised Grammar Induction (2024.findings-naacl)
Copied to clipboard
Boyi Li, Rodolfo Corona, Karttikeya Mangalam, Catherine Chen, Daniel Flaherty, Serge Belongie, Kilian Weinberger, Jitendra Malik, Trevor Darrell, Dan Klein
| Challenge: | Recent studies show multimodal inputs can improve grammar induction, but weak textual baselines are needed for training. |
| Approach: | They use a fixed grammar family to compare multimodal grammar induction methods . they find multimodal inputs can improve grammar induction by grounding textual inputs to the visual world . |
| Outcome: | The proposed model outperforms weaker baselines on four benchmark datasets. |