Papers by Amit Parekh

4 papers
Voices in a Crowd: Searching for clusters of unique perspectives (2024.emnlp-main)

Copied to clipboard

Challenge: Proposed solutions aim to capture minority perspectives by either modelling annotator disagreements or grouping annotators based on shared metadata.
Approach: They propose a framework that trains models without encoding annotator metadata and creates clusters of similar opinions, that are called voices.
Outcome: The proposed framework captures minority perspectives based on demographic factors in two distinct datasets while also capturing majority perspectives.
Multitask Multimodal Prompted Training for Interactive Embodied Task Completion (2023.emnlp-main)

Copied to clipboard

Challenge: Embodied MultiModal Agent (EMMA) is a unified encoder-decoder model that reasons over images and trajectories and casts action prediction as multimodal text generation.
Approach: They propose an Embodied MultiModal Agent (EMMA) that uses a unified encoder-decoder model that reasons over images and trajectories and casts action prediction as multimodal text.
Outcome: The proposed model performs on par with similar models on several VL benchmarks and sets a new state-of-the-art success rate on the Dialog-guided Task Completion (DTC) benchmark.
Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks (2024.emnlp-main)

Copied to clipboard

Challenge: Evaluating generalisation capabilities of multimodal models based solely on performance on out-of-distribution data fails to capture their true robustness . proposed framework examines the role of instructions and inputs in generalisation abilities of such models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity.
Approach: They propose a framework that examines the role of instructions and inputs in the generalisation abilities of multimodal models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity.
Outcome: The proposed framework examines the role of instructions and inputs in the generalisation abilities of multimodal models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity.
FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks (2025.findings-emnlp)

Copied to clipboard

Challenge: Current approaches to embodied AI tend to learn policies from expert demonstrations, but without a mechanism to evaluate the quality of demonstrated actions, they are limited to learning from optimal behaviour or risk replicating errors and inefficiencies.
Approach: They propose to embed language feedback into a Transformer-based policy and optionally complement the traditional next action prediction objective with auxiliary self-supervised learning objectives for feedback prediction.
Outcome: The proposed method improves agents’ compositional generalisation abilities and robustness on a range of embodied Vision-and-Language tasks in a custom babyAI-XGen environment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations