Papers by Paola Cascante-Bonilla
Chat-crowd: A Dialog-based Platform for Visual Layout Composition (N19-4)
Copied to clipboard
| Challenge: | We present Chat-crowd, an interactive environment for visual layout composition via conversational interactions . system can be integrated with crowdsourcing platforms for both synchronous and asynchronous data collection . |
| Approach: | They introduce an interactive environment for visual layout composition via conversational interactions that supports multiple agents with two conversational roles. |
| Outcome: | The proposed system can be integrated with crowdsourcing platforms for both synchronous and asynchronous data collection and has quality controls on the performance of both types of agents. |
Can Hallucination Correction Improve Video-Language Alignment? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing work on hallucination correction for large vision-language models focuses on mitigating hallucisations, but a new approach is needed to improve video-language alignment. |
| Approach: | They propose a self-training framework learning to correct hallucinations in descriptions that do not align with the video content. |
| Outcome: | The proposed framework improves video-language alignment by identifying and correcting inconsistencies in descriptions that do not align with the video content. |
PropTest: Automatic Property Testing for Improved Visual Programming (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Visual Programming is an alternative to end-to-end black-box visual reasoning models. |
| Approach: | They propose a visual programming strategy that leverages Large Language Models to generate the logic of a program in the form of its source code. |
| Outcome: | The proposed method improves ViperGPT on visual question answering and referring expression comprehension with an LLM. |