Papers by Rajvee Sheth

3 papers
Commentator: A Code-mixed Multilingual Text Annotation Framework (2024.emnlp-demo)

Copied to clipboard

Challenge: Existing annotation tools fail to address multilingual datasets efficiently.
Approach: They introduce a code-mixed multilingual text annotation framework, COMMENTATOR . they perform robust qualitative human-based evaluations to showcase its effectiveness .
Outcome: The proposed framework performs faster than baseline annotations in Hinglish and Hindi.
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing (2025.findings-emnlp)

Copied to clipboard

Challenge: COMI-LINGUA is the largest manually annotated Hindi-English code-mixed dataset . 125K+ high-quality instances across five core NLP tasks are annotating by three bilingual annotators .
Approach: COMI-LINGUA is the largest manually annotated Hindi-English code-mixed dataset . 125K+ high-quality instances are annotating by three bilingual annotators .
Outcome: The dataset covers five core NLP tasks, including Token-level Language Identification, Matrix Language Identification and Named Entity Recognition.
Beyond Monolingual Assumptions: A Survey on Code-Switched NLP in the Era of Large Language Models across Modalities (2026.acl-long)

Copied to clipboard

Challenge: Amidst the rapid advances of large language models, most LLMs struggle with mixed-language inputs, limited Code-switching datasets, and evaluation biases.
Approach: They propose a roadmap for inclusive datasets, fair evaluation, and linguistically grounded models to achieve truly multilingual intelligence.
Outcome: The proposed frameworks are based on 327 studies spanning five research areas, 15+ NLP tasks, 30+ datasets, and 80+ languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations