Connecting Language and Vision to Actions (P18-5)

Copied to clipboard

Challenge: Recent advances in language and vision have made incredible progress in describing images and interacting with visual content in a physical or embodied environment.
Approach: This tutorial will provide an overview of the growing number of multimodal tasks and datasets that combine textual and visual understanding.
Outcome: This tutorial will review the state-of-the-art approaches to selected tasks such as image captioning, visual question answering and visual dialog.

Similar Papers

Embodied Language Learning: Opportunities, Challenges, and Future Directions (2024.findings-acl)

Copied to clipboard

Challenge: embodied language learning is a form of language understanding where the language learner is situated in the world, perceives it, and interacts with it.
Approach: They propose to use a concept of World Scopes to measure progress in language understanding research.
Outcome: The proposed framework identifies gaps and suggests future directions for language understanding research.
Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions (2022.acl-long)

Copied to clipboard

Challenge: Vision-and-Language Navigation (VLN) is a research topic that is gaining attention in the field of artificial intelligence.
Approach: They propose to build an embodied agent that can communicate with humans in natural language and navigate in real 3D environments.
Outcome: This paper reviews current studies in the emerging field of vision-and-language navigation . it highlights limitations and opportunities for future work .
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility.
Approach: They propose to redefine the design of vision-language models by identifying key components and creating efficient models with constrained inference costs.
Outcome: The proposed models achieve significant improvements in inference throughput while maintaining high performance.
ViLMedic: a framework for research at the intersection of vision and language in medical AI (2022.acl-demo)

Copied to clipboard

Challenge: Multimodal medical AI is a growing field of interest, especially for tasks that involve multimodal data.
Approach: They propose a vision-and-language medical library to improve multimodal medical predictions and enable new applications.
Outcome: The vision-and-language medical library aims to improve reproducibility and speed up progress across medical AI . it contains a dozen implementations replicating state-of-the-art results on medical datasets . the library is extensible by researchers but also simple for practitioners .
Pushing the Limits of Radiology with Joint Modeling of Visual and Textual Information (P18-3)

Copied to clipboard

Challenge: Recent research has focused on the intersection of computer vision and natural language processing, but its adaption to the medical domain is not fully explored.
Approach: They aim to develop machine learning models that can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images.
Outcome: The proposed models can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images.
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives (2024.findings-acl)

Copied to clipboard

Challenge: Existing video-language understanding systems with human-like senses can mimic both our linguistic medium and visual environment with temporal dynamics.
Approach: They propose to develop video-language understanding systems with human-like senses . they summarize their methods and highlight challenges associated with them .
Outcome: The proposed models perform well in a variety of tasks and domains.
Language Agents: Foundations, Prospects, and Risks (2024.emnlp-tutorials)

Copied to clipboard

Challenge: Language agents are autonomous agents that can follow language instructions to perform diverse tasks in real-world or simulated environments.
Approach: They propose to provide a conceptual framework for language agents and a comprehensive discussion on key topics.
Outcome: The proposed tutorial provides a conceptual framework of language agents and comprehensive discussion on important topic areas.
Vision-Language Pretraining: Current Trends and the Future (2022.acl-tutorials)

Copied to clipboard

Challenge: Recent vision-language models are being used for downstream tasks that require large datasets and supervised datasets.
Approach: They focus on recent vision-language pretraining paradigms and their strengths and shortcomings . they compare the different family of models used for vision- language pretraining .
Outcome: This paper provides the background on image–language datasets, benchmarks, and modeling innovations before the multimodal pretraining area.
The Why and The How: A Survey on Natural Language Interaction in Visualization (2022.naacl-main)

Copied to clipboard

Challenge: Recent research shows that different forms of natural language-based interaction prove suitable to support users in accomplishing various visualization tasks.
Approach: They propose a taxonomy of visualization tasks and a classification system to illustrate the state-of-the-art of natural language-based interaction in visualization.
Outcome: The proposed model can support annotations, recommendations, explanations, and documentation tasks.
Human-AI Interaction in the Age of LLMs (2024.naacl-tutorials)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized the capabilities of AI systems.
Approach: This tutorial will provide an overview of the interaction between humans and Large Language Models (LLMs) it will start with a review of the types of AI models we interact with and walkthrough of the core concepts in Human-AI Interaction.
Outcome: This tutorial will provide an overview of the interaction between humans and LLMs, exploring the challenges, opportunities, and ethical considerations that arise in this dynamic landscape.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations