Papers by Kalika Bali

23 papers
Language Modeling for Code-Mixing: The Role of Linguistic Theory based Synthetic Data (P18-1)

Copied to clipboard

Challenge: Code-mixed (CM) language training is a difficult problem because of lack of data and the increased confusability due to the presence of more than one language.
Approach: They propose a computational technique for creating grammatically valid artificial CM data based on the Equivalence Constraint Theory.
Outcome: The proposed method reduces the perplexity of the model and does not reduce the perceptibility of the models.
Global Readiness of Language Technology for Healthcare: What Would It Take to Combat the Next Pandemic? (2022.coling-1)

Copied to clipboard

Challenge: Language Technology (LT) has been used in the COVID-19 pandemic, but only in a handful of languages.
Approach: They propose to use conversational agents for information dissemination and basic diagnosis in 15 Asian and African languages with varying resource-availability to test their knowledge of LT.
Outcome: The proposed research confirms the pitiful state of LT even for languages with large speaker bases, such as Sinhala and Hausa, and identifies the gaps that could help prioritize research and investment strategies in LT for healthcare.
Discovering Canonical Indian English Accents: A Crowdsourcing-based Approach (L18-1)

Copied to clipboard

Challenge: Automated Speech Recognition systems degrade in performance when recognizing accents that are different from the ones in training data.
Approach: They propose to adapt Acoustic Models that are trained on one accent to a target accent by using a small amount of speech data in the target accent.
Outcome: The proposed model can be used to identify accents in Indian English and other languages.
Crowdsourcing Speech Data for Low-Resource Languages from Low-Income Workers (2020.lrec-1)

Copied to clipboard

Challenge: Existing platforms collect labelled speech data from urban speakers whose dialects are often very different from low-income users.
Approach: They propose to collect labelled speech data directly from low-income workers . they collect 109 hours of data from 36 participants in the Marathi language .
Outcome: The proposed approach can provide valuable supplemental earning opportunities to low-income rural and urban workers.
Language Patterns and Behaviour of the Peer Supporters in Multilingual Healthcare Conversational Forums (2022.lrec-1)

Copied to clipboard

Challenge: a quantitative linguistic analysis of multilingual peer supporters in health-focused WhatsApp forums in Kenya is needed.
Approach: They conduct a quantitative linguistic analysis of the language usage patterns of multilingual peer supporters in two health-focused WhatsApp forums in Kenya.
Outcome: The proposed language analyzer can be used to analyze language usage patterns in two health-focused WhatsApp forums in Kenya.
X-RiSAWOZ: High-Quality End-to-End Multilingual Dialogue Datasets and Few-shot Agents (2023.findings-acl)

Copied to clipboard

Challenge: X-RiSAWOZ dataset has more than 18,000 human-verified dialogue utterances for each language . Xiaoping and Xinhui are the main challenges for task-oriented dialogue research .
Approach: They develop a toolkit to accelerate the post-editing of a new language dataset after translation . their dataset, code, and toolkit are released open-source .
Outcome: The proposed toolkit accelerates the post-editing of a new language dataset after translation.
METAL: Towards Multilingual Meta-Evaluation (2024.findings-naacl)

Copied to clipboard

Challenge: Recent studies show that Large Language Models excel on many standard NLP benchmarks.
Approach: They propose a framework for end-to-end evaluation of Large Language Models as evaluators in multilingual scenarios.
Outcome: The proposed framework evaluates LLMs as evaluators in multilingual scenarios.
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? (2024.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations.
Approach: They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Outcome: The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
The State and Fate of Linguistic Diversity and Inclusion in the NLP World (2020.acl-main)

Copied to clipboard

Challenge: a small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications.
Approach: They examine the relationship between types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time.
Outcome: The proposed model will help to bridge the gap between languages and their resources and convince the ACL community to prioritise the resolution of the predicaments highlighted.
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks (2024.naacl-long)

Copied to clipboard

Challenge: Several new LLMs have been introduced necessitating their evaluation on non-English languages.
Approach: They perform a thorough evaluation of the non-English capabilities of SoTA LLMs by comparing them on the same set of multilingual datasets.
Outcome: The proposed model outperforms models on multilingual datasets on 22 languages including low-resource African languages.
An Integrated Representation of Linguistic and Social Functions of Code-Switching (L18-1)

Copied to clipboard

Challenge: Linguistic studies on code-switching focus on the "how" and "why" of CS . a new model aims to derive CS functions from local and global properties of the code-witched discourse .
Approach: They propose a model that integrates CS phenomena and modalities into a representation that includes local and global properties of the code-switched discourse.
Outcome: The proposed model simplifies the analysis of English/Hindi CS datasets and provides a flexible framework for further studies.
MEGA: Multilingual Evaluation of Generative AI (2023.emnlp-main)

Copied to clipboard

Challenge: Large Large Models (LLMs) have shown impressive performance on many natural language processing tasks such as language understanding, reasoning, and language generation.
Approach: They present a framework for evaluating generative LLMs in the multilingual setting and provide directions for future progress in the field.
Outcome: The proposed framework evaluates generative models on 16 NLP datasets across 70 typologically diverse languages and compares them to state-of-the-art non-autoregressive models.
INMT: Interactive Neural Machine Translation Prediction (D19-3)

Copied to clipboard

Challenge: Existing MT systems are only useful for information assimilation, and require substantial manual post processing.
Approach: They propose an Interactive Machine Translation interface that assists human translators with on-the-fly hints and suggestions.
Outcome: The proposed interface makes the end-to-end translation process faster, more efficient and creates high-quality translations.
INMT-Lite: Accelerating Low-Resource Language Data Collection via Offline Interactive Neural Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Interactive Neural Machine Translation (INMT) systems can be used to promote data collection in several under-resourced languages, but are often not adapted to the deployment constraints native language speakers operate in.
Approach: They propose to use interactive neural machine translation systems to promote data collection in several under-resourced languages by integrating three different modes of Internet-independent deployment and four assistive interfaces suitable for data-sparse languages.
Outcome: The proposed model improves the data generation experience of community members along multiple axes without compromising on the quality of the generated translations.
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) excel in diverse applications but still struggle with non-Latin scripts and low-resource languages.
Approach: They propose a dynamic learning approach that optimizes prompt strategy, embedding model, and LLM per query at runtime.
Outcome: The proposed approach achieves 10-15% improvements in multilingual performance over pre-trained models and 4x gains compared to fine-tuned, language-specific models.
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models (2025.naacl-long)

Copied to clipboard

Challenge: Neural Machine Translation models traditionally use Sinusoidal Positional Embeddings . retraining with newer methods like ROPE or ALIBI is computationally expensive .
Approach: They propose to transition NMT models from Sinusoidal to Relative PEs without compromising performance.
Outcome: The proposed approach outperforms models trained with Sinusoidal PEs on document-level benchmarks . the results show that parameter-efficient fine-tuning can facilitate the transition .
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being integrated with social applications . large data sets are limited in their representation of information and do not capture knowledge from the Web .
Approach: They propose a gamified framework that uses collective sensemaking to collect artifacts from 19 different Indian geographic subcultures and benchmark four popular LLMs.
Outcome: The proposed framework is based on 260 participants from 19 different Indian geographic subcultures and shows that it can be used across regional sub-cultures.
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting (2024.emnlp-main)

Copied to clipboard

Challenge: Socio-demographic prompting is a commonly employed approach to study cultural biases in LLMs as well as for aligning models to certain cultures.
Approach: They propose to use socio-demographic prompting to probe four LLMs with culturally sensitive and non-sensitive cues on datasets that are supposed to be culturally neutral or sensitive.
Outcome: The proposed model shows significant differences in responses on both kinds of datasets, casting doubt on its robustness.
“Fifty Shades of Bias”: Normative Ratings of Gender Bias in GPT Generated English Text (2023.emnlp-main)

Copied to clipboard

Challenge: Prior work treats gender bias as a binary classification task, but a comparative annotation framework can be used to assess the impact of biases.
Approach: They propose to generate a dataset with normative ratings of gender bias in English text with a comparative annotation framework.
Outcome: The first dataset of GPT-generated English text with normative ratings of gender bias is analyzed using Best–Worst Scaling .
Everything you need to know about Multilingual LLMs: Towards fair, performant and reliable models for languages of the world (2023.acl-tutorials)

Copied to clipboard

Challenge: Responsible AI issues such as fairness, bias and toxicity will be discussed in this tutorial .
Approach: This tutorial will describe various aspects of scaling up language technologies to many of the world’s languages by describing the latest research in Massively Multilingual Language Models (MMLMs).
Outcome: This tutorial will cover various aspects of scaling up language technologies to many of the world's languages by describing the latest research in multilingual models.
Learnings from Technological Interventions in a Low Resource Language: A Case-Study on Gondi (2020.lrec-1)

Copied to clipboard

Challenge: 40% of all the languages in the world face the danger of extinction in the near future . when a language dies out, future generations lose a vital part of the culture that is necessary to completely understand it.
Approach: They propose to use 4 technology-driven methods of data collection to collect data on Gondi, a low-resource vulnerable language spoken by 2.3 million tribal people in south and central India.
Outcome: The proposed methods collected 12,000 translated words and/or sentences and identified more than 650 community members whose help can be solicited for future translation efforts.
An Interdisciplinary Approach to Human-Centered Machine Translation (2025.emnlp-main)

Copied to clipboard

Challenge: Despite progress in MT, a gap persists between how the technology is developed and how it is used in real-world contexts.
Approach: They propose a human-centered approach to machine translation (MT) they argue that MT should be evaluated with diverse goals and contexts of use .
Outcome: The proposed approach emphasizes alignment of evaluation and design with diverse communicative goals and contexts of use.
“#DisabledOnIndianTwitter” : A Dataset towards Understanding the Expression of People with Disabilities on Indian Twitter (2022.findings-aacl)

Copied to clipboard

Challenge: a majority of disabled Indians exist at the margins of society with little to no access to social media . as access to ICTs and high-speed internet grows, Indian Twitter's user base is expanding to include disability influencers, activists, and everyday disabled users.
Approach: They propose a hierarchical annotation taxonomy to classify tweets into various themes including discrimination, advocacy, and self-identification.
Outcome: The proposed taxonomy classifies 2,384 tweets into various themes including discrimination, advocacy, and self-identification.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations