| Challenge: | Recent advances in language and vision have made incredible progress in describing images and interacting with visual content in a physical or embodied environment. |
| Approach: | This tutorial will provide an overview of the growing number of multimodal tasks and datasets that combine textual and visual understanding. |
| Outcome: | This tutorial will review the state-of-the-art approaches to selected tasks such as image captioning, visual question answering and visual dialog. |
Similar Papers
Embodied Language Learning: Opportunities, Challenges, and Future Directions (2024.findings-acl)
Copied to clipboard
| Challenge: | embodied language learning is a form of language understanding where the language learner is situated in the world, perceives it, and interacts with it. |
| Approach: | They propose to use a concept of World Scopes to measure progress in language understanding research. |
| Outcome: | The proposed framework identifies gaps and suggests future directions for language understanding research. |
Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions (2022.acl-long)
Copied to clipboard
| Challenge: | Vision-and-Language Navigation (VLN) is a research topic that is gaining attention in the field of artificial intelligence. |
| Approach: | They propose to build an embodied agent that can communicate with humans in natural language and navigate in real 3D environments. |
| Outcome: | This paper reviews current studies in the emerging field of vision-and-language navigation . it highlights limitations and opportunities for future work . |
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility. |
| Approach: | They propose to redefine the design of vision-language models by identifying key components and creating efficient models with constrained inference costs. |
| Outcome: | The proposed models achieve significant improvements in inference throughput while maintaining high performance. |
ViLMedic: a framework for research at the intersection of vision and language in medical AI (2022.acl-demo)
Copied to clipboard
Jean-benoit Delbrouck, Khaled Saab, Maya Varma, Sabri Eyuboglu, Pierre Chambon, Jared Dunnmon, Juan Zambrano, Akshay Chaudhari, Curtis Langlotz
| Challenge: | Multimodal medical AI is a growing field of interest, especially for tasks that involve multimodal data. |
| Approach: | They propose a vision-and-language medical library to improve multimodal medical predictions and enable new applications. |
| Outcome: | The vision-and-language medical library aims to improve reproducibility and speed up progress across medical AI . it contains a dozen implementations replicating state-of-the-art results on medical datasets . the library is extensible by researchers but also simple for practitioners . |
Pushing the Limits of Radiology with Joint Modeling of Visual and Textual Information (P18-3)
Copied to clipboard
| Challenge: | Recent research has focused on the intersection of computer vision and natural language processing, but its adaption to the medical domain is not fully explored. |
| Approach: | They aim to develop machine learning models that can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
| Outcome: | The proposed models can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives (2024.findings-acl)
Copied to clipboard
Thong Nguyen, Yi Bin, Junbin Xiao, Leigang Qu, Yicong Li, Jay Zhangjie Wu, Cong-Duy Nguyen, See-Kiong Ng, Anh Tuan Luu
| Challenge: | Existing video-language understanding systems with human-like senses can mimic both our linguistic medium and visual environment with temporal dynamics. |
| Approach: | They propose to develop video-language understanding systems with human-like senses . they summarize their methods and highlight challenges associated with them . |
| Outcome: | The proposed models perform well in a variety of tasks and domains. |
Language Agents: Foundations, Prospects, and Risks (2024.emnlp-tutorials)
Copied to clipboard
| Challenge: | Language agents are autonomous agents that can follow language instructions to perform diverse tasks in real-world or simulated environments. |
| Approach: | They propose to provide a conceptual framework for language agents and a comprehensive discussion on key topics. |
| Outcome: | The proposed tutorial provides a conceptual framework of language agents and comprehensive discussion on important topic areas. |
Vision-Language Pretraining: Current Trends and the Future (2022.acl-tutorials)
Copied to clipboard
| Challenge: | Recent vision-language models are being used for downstream tasks that require large datasets and supervised datasets. |
| Approach: | They focus on recent vision-language pretraining paradigms and their strengths and shortcomings . they compare the different family of models used for vision- language pretraining . |
| Outcome: | This paper provides the background on image–language datasets, benchmarks, and modeling innovations before the multimodal pretraining area. |
The Why and The How: A Survey on Natural Language Interaction in Visualization (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent research shows that different forms of natural language-based interaction prove suitable to support users in accomplishing various visualization tasks. |
| Approach: | They propose a taxonomy of visualization tasks and a classification system to illustrate the state-of-the-art of natural language-based interaction in visualization. |
| Outcome: | The proposed model can support annotations, recommendations, explanations, and documentation tasks. |
Human-AI Interaction in the Age of LLMs (2024.naacl-tutorials)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized the capabilities of AI systems. |
| Approach: | This tutorial will provide an overview of the interaction between humans and Large Language Models (LLMs) it will start with a review of the types of AI models we interact with and walkthrough of the core concepts in Human-AI Interaction. |
| Outcome: | This tutorial will provide an overview of the interaction between humans and LLMs, exploring the challenges, opportunities, and ethical considerations that arise in this dynamic landscape. |