Papers by Arnav Goel
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation (2026.acl-long)
Copied to clipboard
| Challenge: | In morphologically rich languages, gender influences verb conjugation, pronouns, and even first-person constructions with explicit and implicit mentions of gender. |
| Approach: | They propose a morphologically grounded large-scale benchmark dataset for evaluating gender-aware generation in three typologically diverse grammatically gendered languages: French, Arabic, and Hindi. |
| Outcome: | The proposed dataset compares 15 popular multilingual large language models on their ability to handle morphological gender and morphology agreement. |
HORIZON: A Benchmark for In-the-wild User Behaviour Modeling (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing user modeling benchmarks focus on short sessions and next-item prediction within a single domain. |
| Approach: | They propose a benchmark that reformulates user modeling along three axes . it covers 54M users and 35M items, enabling pretraining and evaluation . they propose tasks and evaluation setups that better reflect real-world deployment scenarios . |
| Outcome: | The proposed benchmark covers 54M users and 35M items, and is based on Amazon Reviews. |
HLDC: Hindi Legal Documents Corpus (2022.findings-acl)
Copied to clipboard
Arnav Kapoor, Mudit Dhawan, Anmol Goel, Arjun T H, Akshala Bhatnagar, Vibhu Agrawal, Amul Agrawal, Arnab Bhattacharya, Ponnurangam Kumaraguru, Ashutosh Modi
| Challenge: | Existing systems that process legal documents are lacking high-quality corpora in low resource languages such as Hindi. |
| Approach: | They propose a Hindi Legal Documents Corpus (HLDC) that contains 900K legal documents in Hindi. |
| Outcome: | The proposed model is based on a corpus of more than 900K legal documents in Hindi. |
Mind the Gap: Multilingual Divide in LLM Bias Detection and Reasoning (2026.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly deployed in multilingual settings . but most bias evaluation remains English-centric and ignores how bias manifests within reasoning . |
| Approach: | They evaluate large language models with supervised fine-tuning and preference optimization . they find that bias varies substantially across languages, with consistent degradation in non-English settings . |
| Outcome: | The proposed model improves in English, Dutch, Spanish, and Turkish using the MBBQ benchmark. |