Papers by Anushka Anushka

9 papers
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: SteerVLM is a lightweight steering module designed to guide Vision-Language Models (VLMs) towards outputs that better adhere to desired instructions.
Approach: They propose a lightweight steering module that learns from latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activations connecting the language modality with image context.
Outcome: The proposed steering module outperforms existing intervention techniques on steering and hallucination mitigation benchmarks for VLMs.
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models’ Understanding on Indian Culture (2025.emnlp-main)

Copied to clipboard

Challenge: DRISHTIKON is a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture.
Approach: They evaluate a wide range of vision-language models across zero-shot and chain-of-thought settings and use them to evaluate cultural understanding of generative AI systems.
Outcome: The DRISHTIKON dataset covers 15 languages, all states and union territories, and incorporating over 64,000 aligned text-image pairs.
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages? (2024.acl-short)

Copied to clipboard

Challenge: a recent study focused on machine translation evaluation for low-resource languages . linguistic aspects that vary across languages are factors that will exacerbate the problem in low-source languages due to the reliance on extensive data resources.
Approach: They propose to use multi-dimensional quality metrics and DA annotations to meta-evaluate MT evaluation metrics for low-resource languages.
Outcome: The proposed evaluation metrics are based on human scores on the candidate translations of assamese, maithili, and Punjabi.
LLM4RE: A Data-centric Feasibility Study for Relation Extraction (2025.coling-main)

Copied to clipboard

Challenge: Relation Extraction (RE) is a critical step in information extraction due to its wide-scale applicability for downstream applications such as Knowledge Base creation and Question Answering (QA).
Approach: They propose to conduct the first feasibility analysis to explore the viability of Large Language Models for RE by investigating their robustness to various RE scenarios stemming from data-specific characteristics.
Outcome: The proposed models are robust to various RE scenarios stemming from data-specific characteristics, but their performance is not yet fully understood.
SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models’ Knowledge of Indian Culture (2025.findings-acl)

Copied to clipboard

Challenge: Language models excel in syntactic and semantic analysis, while small language models struggle in region-specific contexts.
Approach: They evaluate SANSKRITI on leading Large Language Models, Indic Language Model, and Small Language Model (SLM) it covers 16 key attributes of Indian culture including rituals and ceremonies, history, tourism, cuisine, dance and music, costume, language, art, festivals, religion, medicine, transport, sports, nightlife and personalities.
Outcome: The SANSKRITI dataset covers 16 attributes of Indian culture . it reveals that many models struggle in region-specific contexts .
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to cross-modal image-text retrieval struggle with nuanced cross-modal relationships.
Approach: They propose a set-based approach that represents each sample with multiple embeddings to capture nuanced and diverse relationships.
Outcome: The proposed method achieves state-of-the-art performance on MS-COCO and Flickr30k without external data.
Subjective Behaviors and Preferences in LLM: Language of Browsing (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) fuel expectations that a single trained model can effectively align with preferences of myriad users for a given task within a domain.
Approach: They introduce clusterwise LM training, HeTLM, appropriate for subjective behaviors . authors say small LM outperforms large pretrained LMs; heterogeneous cluster specific set of parameters outperformed single LM .
Outcome: The proposed model outperforms large pretrained or finetuned models in the domain of subjective behavior and preferences.
Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game (2024.emnlp-main)

Copied to clipboard

Challenge: We evaluate the performance of large language models (LLMs) against expert and novice human players.
Approach: They propose to use the New York Times Connections game as a test bed to evaluate the abstract reasoning capabilities of large language models (LLMs) they propose to test the ability of large-language models to be able to cluster and categorize words using semantic relations.
Outcome: The proposed game is a test bed for evaluating abstract reasoning capabilities in humans and AI systems.
Flexible-length Text Infilling for Discrete Diffusion Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing discrete diffusion models lack flexibility for text infilling without ground-truth positional data.
Approach: They propose a discrete diffusion model that jointly denoises token values and token positions using a novel sample-level Optimal Transport coupling.
Outcome: The proposed method outperforms existing methods on infilling benchmarks such as One-Billion-Word and Yelp.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations