Papers by Ravi Shekhar

9 papers
Ask No More: Deciding when to guess in referential visual dialogue (C18-1)

Copied to clipboard

Challenge: Using a task-oriented visual dialogue model, we add a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess.
Approach: They augment a task-oriented visual dialogue model with a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess.
Outcome: The proposed model can be enhanced with a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess.
CWID-hi: A Dataset for Complex Word Identification in Hindi Text (2022.lrec-1)

Copied to clipboard

Challenge: Text simplification is a method for improving the accessibility of text by converting complex sentences into simple sentences.
Approach: They propose to use Hindi knowledge annotators to capture the annotator’s language knowledge to build an automatic complex word classifier using a soft voting approach.
Outcome: The proposed dataset shows that native and non-native annotators perceive complex words differently depending on their language knowledge.
Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat (N19-1)

Copied to clipboard

Challenge: Existing systems that address the abilities that need to be put to work during conversations are lacking in terms of visual grounding.
Approach: They propose a visually-grounded dialogue state encoder which integrates visual grounding with dialogue system components.
Outcome: The proposed system improves the GuessWhat?! game by combining guessing and asking questions with multi-task learning.
Language Family Matters: Evaluating SpeechLLMs Across Linguistic Boundaries (2026.eacl-short)

Copied to clipboard

Challenge: Existing approaches to integrate speech encoders with large language models (LLMs) have limited resources and lack linguistic relatedness.
Approach: They propose a connector-sharing strategy based on linguistic family membership that allows one connector per family to share a frozen speech encoder with a pretrained LLM.
Outcome: The proposed system reduces parameter count while improving generalization across domains, compared with existing connectors.
CoRAL: a Context-aware Croatian Abusive Language Dataset (2022.findings-aacl)

Copied to clipboard

Challenge: Semi-automated comment moderation systems can greatly aid human moderators by either automatically classifying the examples or allowing the moderator to prioritize which comments to consider first.
Approach: They propose to use a language and culturally aware Croatian Abusive dataset to analyze inappropriate comments in a context-based manner.
Outcome: The proposed dataset shows that current models degrade when comments are not explicit and further degrades when language skill and context knowledge are required to interpret the comment.
LEDA: a Large-Organization Email-Based Decision-Dialogue-Act Analysis Dataset (2023.findings-acl)

Copied to clipboard

Challenge: Using dialog acts to study decision-making in large distributed organizations is challenging due to the size and distributed nature of such groups.
Approach: They propose a set of dialog acts for the study of decision-making mechanisms in large distributed organizations.
Outcome: The proposed dataset can be used to better understand decision-making in large distributed organizations.
Tracing Linguistic Markers of Influence in a Large Online Organisation (2023.acl-short)

Copied to clipboard

Challenge: Social science and psycholinguistic research have shown that power and status affect how people use language in a range of domains.
Approach: They propose to use lexical categories and BERT to predict levels of influence in an online community and identify key linguistic differences between people before and after becoming influential.
Outcome: The results show that participants' levels of influence can be predicted from their email text, and identify key differences in language use for the same person before and after becoming influential.
Denoising Labeled Data for Comment Moderation Using Active Learning (2024.lrec-main)

Copied to clipboard

Challenge: Large contextualized language models (LLMs) are becoming ubiquitous in natural language processing due to their performance and adaptability to diverse tasks.
Approach: They propose to use active learning methods to denoise textual data for model training by sampling the most informative examples with noisy labels with active learning.
Outcome: The proposed method reduces the cost of reannotation by reducing noise in noisy examples.
Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing MCQA benchmarks fail to capture the full reasoning capabilities of video language models due to selection bias.
Approach: They propose a method to reduce selection bias in video-to-text LLMs by suppressing "blind guessing" they propose 'bold' calibration technique to balance selection bias.
Outcome: The proposed method reduces selection bias and improves model performance compared to existing methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations