Papers by Alexander Hauptmann

11 papers
Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models (2021.naacl-main)

Copied to clipboard

Challenge: a new study examines zero-shot cross-lingual transfer of vision-language models . we study multilingual text-to-video search in non-English languages without annotations .
Approach: They propose a Transformer-based model that learns contextual multilingual multimodal embeddings . they propose 'zero-shot cross-lingual transfer' to improve multilingual search .
Outcome: The proposed model outperforms baselines on multilingual text-to-video search and multilingual image search on VTT and VATEX.
Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations (D19-1)

Copied to clipboard

Challenge: Recent studies have advanced learning VSE under the monolingual setup.
Approach: They propose a model with diverse multi-head attention to learn grounded multilingual multimodal representations by leveraging visual object detection.
Outcome: The proposed model performs well in German-Image and English-Image matching tasks and in the Semantic Textual Similarity task with English descriptions of visual content.
DocumentNet: Bridging the Data Gap in Document Pre-training (2023.emnlp-industry)

Copied to clipboard

Challenge: Document understanding tasks are a tedious task that requires extensive training and privacy constraints.
Approach: They propose a method to collect weakly labeled data from the web to benefit VDER training . the collected dataset does not depend on specific document types or entity sets .
Outcome: The proposed method does not depend on specific document types or entity sets, making it universally applicable to all VDER tasks.
Transitive Consistency Constrained Learning for Entity-to-Entity Stance Detection (2024.acl-long)

Copied to clipboard

Challenge: Entity-to-entity stance detection is a streamlined task without the complex dependency structure for structural sentiment analysis.
Approach: They propose a method that models transitive consistency constraints during training to help train entity-to-entity stance detection models.
Outcome: The proposed method improves both classification-based and generation-based models while large language models struggle with predicting link direction and neutral labels.
KAT: A Knowledge Augmented Transformer for Vision-and-Language (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for knowledge retrieval and answer prediction have left open questions about the quality and relevance of the retrieved knowledge and how the reasoning processes over implicit and explicit knowledge should be integrated.
Approach: They propose a Knowledge Augmented Transformer which integrates both implicit and explicit knowledge in an encoder-decoder architecture while simultaneously reasoning over both knowledge sources during answer generation.
Outcome: The proposed model achieves a strong state-of-the-art (+6% absolute) on the open-domain multimodal task of OK-VQA.
Unsupervised Multimodal Neural Machine Translation with Pseudo Visual Pivoting (2020.acl-main)

Copied to clipboard

Challenge: Unsupervised machine translation (MT) has recently achieved impressive results with monolingual corpora.
Approach: They propose to utilize visual content for disambiguation and promoting latent space alignment in unsupervised machine translation by using multimodal back-translation and pseudo visual pivoting.
Outcome: The proposed model improves over state-of-the-art methods and generalizes well when images are not available at the testing time.
SHIELD: LLM-Driven Schema Induction for Predictive Analytics in EV Battery Supply Chain Disruptions (2024.emnlp-industry)

Copied to clipboard

Challenge: EV battery supply chain is vulnerable to disruptions caused by natural disasters and geopolitical tensions.
Approach: They propose a system integrating Large Language Models with domain expertise for EV supply chain risk assessment.
Outcome: Evaluated on 12,070 paragraphs from 365 sources (2022-2023), SHIELD outperforms baseline GCNs and LLM+prompt methods in disruption prediction.
ExCL: Extractive Clip Localization Using Natural Language Descriptions (N19-1)

Copied to clipboard

Challenge: Prior approaches to retrieving clips within videos based on a given query are inefficient and text-clip similarity driven ranking-based approaches are far more complicated.
Approach: They propose an extractive approach that extracts the start and end frames by leveraging cross-modal interactions between the text and video to generate a joint representation.
Outcome: The proposed approach significantly outperforms state-of-the-art on two datasets and has comparable performance on a third.
Event-Related Bias Removal for Real-time Disaster Events (2020.findings-emnlp)

Copied to clipboard

Challenge: Social media has become an important tool to share information about crisis events such as natural disasters and mass attacks.
Approach: They propose to train an adversarial neural model to remove latent event-specific biases and improve the performance on tweet importance classification.
Outcome: The proposed model removes event-specific biases and improves on tweet importance classification.
Zero-Shot and Few-Shot Stance Detection on Varied Topics via Conditional Generation (2023.acl-short)

Copied to clipboard

Challenge: Existing work on stance detection focuses on in-domain or leave-out targets with only a few target choices.
Approach: They propose to use a conditional generation framework to denoise from partially-filled templates to better utilize the semantics among input, label, and target texts.
Outcome: The proposed method significantly outperforms strong baselines on VAST and achieves new state-of-the-art performance.
Towards Open-Domain Twitter User Profile Inference (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to user profile inference focus on limited attributes and can reveal users' private information.
Approach: They propose a prompt-based generation method which can infer values implicitly mentioned in Twitter user profiles.
Outcome: The proposed method can infer more comprehensive user profiles than baseline extraction-based methods, but limitations remain to be applied for real-world use.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations