Papers by Khyati Mahajan

5 papers
A Case Study of Analysis of Construals in Language on Social Media Surrounding a Crisis Event (2021.acl-srw)

Copied to clipboard

Challenge: construal level theory (CLT) uses concreteness as covariate to analyze language around political import events.
Approach: They propose to include psycholinguistic measures of concreteness as covariates in topic models to analyze the language around an event of political import.
Outcome: The proposed model incorporates measures of concreteness as covariates to inform the analysis of language around the 2017 rally.
Persona-aware Multi-party Conversation Response Generation (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in natural language generation have addressed multi-turn dialogues . interactions with more than 2 participants pose new and interesting challenges for MPC modeling .
Approach: They propose to include persona attributes of speaker and addressee relevant to each utterance in a multi-party conversation dataset and a persona-aware heterogeneous graph transformer response generation model.
Outcome: The proposed model includes persona attributes of speaker and addressee relevant to each utterance.
Prompting with Phonemes: Enhancing LLMs’ Multilinguality for Non-Latin Script Languages (2025.naacl-long)

Copied to clipboard

Challenge: Multilingual LLMs have achieved remarkable benchmark performance, but continue to underperform on non-Latin script languages.
Approach: They propose to integrate phonemic transcriptions as complementary signals to induce script-invariant representations by integrating phonemic and orthographic transcriptions.
Outcome: The proposed approach improves performance for Latin and non-Latin script languages, with 12.6% performance improvement and 15.1% performance improvement compared to randomized ICL retrieval.
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to collect instruction fine-tuning data are limited due to their toxicity, privacy and toxicity concerns.
Approach: They propose to use a two-step taxonomy to transform a small set of human written instructions into complex and challenging conversations.
Outcome: M2Lingual has 175K conversations across 70 languages with a balanced mix of high, low and mid-resourced languages.
Controllable Clustering with LLM-driven Embeddings (2025.emnlp-industry)

Copied to clipboard

Challenge: Unsupervised text clustering is unlikely to produce groupings that work across use cases . authors present techniques to effectively control text embeddings with minimal human input .
Approach: They propose techniques to control text embeddings with minimal human input . they evaluate clustering performance for datasets with multiple independent labels .
Outcome: The proposed techniques improve clustering for one perspective or use case, but at a tradeoff in performance for another use case.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations