Papers by Khyati Mahajan
A Case Study of Analysis of Construals in Language on Social Media Surrounding a Crisis Event (2021.acl-srw)
Copied to clipboard
| Challenge: | construal level theory (CLT) uses concreteness as covariate to analyze language around political import events. |
| Approach: | They propose to include psycholinguistic measures of concreteness as covariates in topic models to analyze the language around an event of political import. |
| Outcome: | The proposed model incorporates measures of concreteness as covariates to inform the analysis of language around the 2017 rally. |
Persona-aware Multi-party Conversation Response Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in natural language generation have addressed multi-turn dialogues . interactions with more than 2 participants pose new and interesting challenges for MPC modeling . |
| Approach: | They propose to include persona attributes of speaker and addressee relevant to each utterance in a multi-party conversation dataset and a persona-aware heterogeneous graph transformer response generation model. |
| Outcome: | The proposed model includes persona attributes of speaker and addressee relevant to each utterance. |
Prompting with Phonemes: Enhancing LLMs’ Multilinguality for Non-Latin Script Languages (2025.naacl-long)
Copied to clipboard
Hoang H Nguyen, Khyati Mahajan, Vikas Yadav, Julian Salazar, Philip S. Yu, Masoud Hashemi, Rishabh Maheshwary
| Challenge: | Multilingual LLMs have achieved remarkable benchmark performance, but continue to underperform on non-Latin script languages. |
| Approach: | They propose to integrate phonemic transcriptions as complementary signals to induce script-invariant representations by integrating phonemic and orthographic transcriptions. |
| Outcome: | The proposed approach improves performance for Latin and non-Latin script languages, with 12.6% performance improvement and 15.1% performance improvement compared to randomized ICL retrieval. |
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing approaches to collect instruction fine-tuning data are limited due to their toxicity, privacy and toxicity concerns. |
| Approach: | They propose to use a two-step taxonomy to transform a small set of human written instructions into complex and challenging conversations. |
| Outcome: | M2Lingual has 175K conversations across 70 languages with a balanced mix of high, low and mid-resourced languages. |
Controllable Clustering with LLM-driven Embeddings (2025.emnlp-industry)
Copied to clipboard
Kerria Pang-Naylor, Shivani Manivasagan, Aitong Zhong, Mehak Garg, Nicholas Mondello, Blake Buckner, Jonathan P. Chang, Khyati Mahajan, Masoud Hashemi, Fabio Casati
| Challenge: | Unsupervised text clustering is unlikely to produce groupings that work across use cases . authors present techniques to effectively control text embeddings with minimal human input . |
| Approach: | They propose techniques to control text embeddings with minimal human input . they evaluate clustering performance for datasets with multiple independent labels . |
| Outcome: | The proposed techniques improve clustering for one perspective or use case, but at a tradeoff in performance for another use case. |