Papers by Mayank Gupta
Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code (2025.coling-industry)
Copied to clipboard
Taishi Nakamura, Mayank Mishra, Simone Tedeschi, Yekun Chai, Jason T. Stillerman, Felix Friedrich, Prateek Yadav, Tanmay Laud, Vu Minh Chien, Terry Yue Zhuo, Diganta Misra, Ben Bogin, Xuan-Son Vu, Marzena Karpinska, Arnav Varma Dantuluri, Wojciech Kusa, Tommaso Furlanello, Rio Yokota, Niklas Muennighoff, Suhas Pai, Tosin Adewumi, Veronika Laippala, Xiaozhe Yao, Adalberto Barbosa Junior, Aleksandr Drozd, Jordan Clive, Kshitij Gupta, Liangyu Chen, Qi Sun, Ken Tsui, Nour Moustafa-Fahmy, Nicolo Monti, Tai Dang, Ziyang Luo, Tien-Tung Bui, Roberto Navigli, Virendra Mehta, Matthew Blumberg, Victor May, Hiep Nguyen, Sampo Pyysalo
| Challenge: | Pretrained language models are integral part of AI applications, but their high computational cost limits accessibility. |
| Approach: | They evaluate Aurora-M, a 15B parameter multilingual open-source model trained on English, Finnish, Hindi, Japanese, Vietnamese, and code. |
| Outcome: | The proposed model outperforms existing models on English, Finnish, Hindi, Japanese, Vietnamese, and code. |
MUTANT: A Multi-sentential Code-mixed Hinglish Dataset (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods to identify code-mixed text are difficult to scale effectively and efficiently on multi-sentential data. |
| Approach: | They propose to identify multi-sentential code-mixed text (MCT) from multilingual articles using a token-level language-aware pipeline. |
| Outcome: | The proposed dataset includes 67k articles with 85k identified Hinglish MCTs. |
Improving Human-Labeled Data through Dynamic Automatic Conflict Resolution (2020.coling-main)
Copied to clipboard
| Challenge: | a scalable method for estimating the noisiness of labels produced by crowdsourcing annotation tasks is developed. |
| Approach: | They propose a scalable method for estimating the noisiness of labels produced by crowdsourcing semantic annotation tasks and reducing the resulting error by 20-30%. |
| Outcome: | The proposed method reduces the error of the labeling process by 20-30% compared to other common labeling strategies. |
JobMatchAI - An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI (2026.acl-demo)
Copied to clipboard
| Challenge: | Recruiters and job seekers rely on search systems to navigate labor markets . many systems fail to handle skill synonyms and nonlinear careers . |
| Approach: | They propose a production-ready system that integrates Transformer embeddings, skill knowledge graphs, and interpretable reranking. |
| Outcome: | The proposed system optimizes utility across skill fit, experience, location, salary, and company preferences. |