Papers by Monojit Choudhury
Copied to clipboard
| Challenge: | Recent studies show that Large Language Models are biased towards a Western and Anglo-centric worldview. |
| Approach: | They propose to extend the Octopus test to measure "cultural awareness" they argue that cultural awareness is needed for AI systems to be useful across cultures . |
| Outcome: | The proposed method argues that cultural awareness is not cultural knowledge, but meta-cultural competence . the proposed method is based on the octopus test, which shows it is impossible to learn meaning from real-world concepts without knowing intent and meaning . |
Copied to clipboard
| Challenge: | Massively Multilingual Transformer based Language Models have been shown to be effective on zero-shot transfer across languages, though performance varies from language to language depending on pivot language(s) used for fine-tuning. |
| Approach: | They propose to combine multi-task learning problems with multi-lingual Transformers to model zero-shot transfer across languages. |
| Outcome: | The proposed model can predict zero-shot transfer across languages with a multi-task learning problem with pretraining data in very few languages. |
Copied to clipboard
| Challenge: | Existing approaches for few-shot transfer show significant gain over zero-shot transfers . language resource distribution is skewed across the world's languages . proposed methods use multiple measures such as data entropy and gradient embedding . |
| Approach: | They propose a loss embedding method for sequence labeling tasks that induces diversity and uncertainty sampling similar to gradient embeddment. |
| Outcome: | The proposed methods outperform baseline methods for POS tagging, NER, and NLI tasks for up to 20 languages. |
Copied to clipboard
| Challenge: | Ethical reasoning is a crucial skill for Large Language Models (LLMs). However, moral values are not universal, but rather influenced by language and culture. |
| Approach: | They extend the study of ethical reasoning of LLMs by (CITATION) to a multilingual setup using six languages: English, Spanish, Russian, Chinese, Hindi, and Swahili. |
| Outcome: | The proposed model is based on a multilingual setup in English, Spanish, Russian, Chinese, Hindi, and Swahili. |
Copied to clipboard
| Challenge: | a paper by a team of researchers proposes that large language models should be morally aligned to ethical principles . a moral compass is a model that integrates moral dilemmas with moral principles pertaining to different foramlisms of normative ethics . |
| Approach: | They propose to infuse generic ethical reasoning capabilities into large-scale models . they argue that LLMs should take a moral stance on value pluralism . |
| Outcome: | a new ethical reasoning framework integrates moral dilemmas with moral principles . the framework is based on the results of a hypothetical case study on a large-scale model . |
Copied to clipboard
| Challenge: | Zero-shot cross-lingual transfer has been shown to be sub-optimal across low-resource languages due to the skew in resource distribution in languages. |
| Approach: | They propose to jointly reduce feature incongruity between the source and target language and increase generalization capabilities of pre-trained multilingual transformers. |
| Outcome: | Empirical results show that the proposed approach outperforms the standard zero-shot fine-tuning method on multiple datasets across all languages using only unlabeled instances in the target language. |
Copied to clipboard
| Challenge: | Code-mixed (CM) language training is a difficult problem because of lack of data and the increased confusability due to the presence of more than one language. |
| Approach: | They propose a computational technique for creating grammatically valid artificial CM data based on the Equivalence Constraint Theory. |
| Outcome: | The proposed method reduces the perplexity of the model and does not reduce the perceptibility of the models. |
Copied to clipboard
| Challenge: | Language Technology (LT) has been used in the COVID-19 pandemic, but only in a handful of languages. |
| Approach: | They propose to use conversational agents for information dissemination and basic diagnosis in 15 Asian and African languages with varying resource-availability to test their knowledge of LT. |
| Outcome: | The proposed research confirms the pitiful state of LT even for languages with large speaker bases, such as Sinhala and Hausa, and identifies the gaps that could help prioritize research and investment strategies in LT for healthcare. |
Copied to clipboard
| Challenge: | a framework to evaluate the performance and cost trade-offs between machine-translated and manually-created labelled data is presented. |
| Approach: | They propose a framework to evaluate the performance and cost trade-offs between machine-translated and manually-created labelled data for task-specific fine-tuning of massively multilingual language models. |
| Outcome: | The proposed framework can be used to evaluate cost trade-offs between machine-translated and manually-created labelled data for task-specific fine-tuning of massively multilingual models. |
Copied to clipboard
| Challenge: | Automated Speech Recognition systems degrade in performance when recognizing accents that are different from the ones in training data. |
| Approach: | They propose to adapt Acoustic Models that are trained on one accent to a target accent by using a small amount of speech data in the target accent. |
| Outcome: | The proposed model can be used to identify accents in Indian English and other languages. |
Copied to clipboard
| Challenge: | Existing platforms collect labelled speech data from urban speakers whose dialects are often very different from low-income users. |
| Approach: | They propose to collect labelled speech data directly from low-income workers . they collect 109 hours of data from 36 participants in the Marathi language . |
| Outcome: | The proposed approach can provide valuable supplemental earning opportunities to low-income rural and urban workers. |
Copied to clipboard
| Challenge: | a quantitative linguistic analysis of multilingual peer supporters in health-focused WhatsApp forums in Kenya is needed. |
| Approach: | They conduct a quantitative linguistic analysis of the language usage patterns of multilingual peer supporters in two health-focused WhatsApp forums in Kenya. |
| Outcome: | The proposed language analyzer can be used to analyze language usage patterns in two health-focused WhatsApp forums in Kenya. |
Copied to clipboard
| Challenge: | Massively Multilingual Language Models (MMLMs) have gained popularity due to their effectiveness in cross-lingual transfer. |
| Approach: | They investigate how well calibrated MMLMs are with respect to confidence . they find that calibration methods like temperature scaling and label smoothing improve calibration . |
| Outcome: | The proposed models are able to generalize in languages unseen during fine-tuning, but they are not reliable across languages. |
Copied to clipboard
| Challenge: | Topic-sensitive query set expansion is crucial for queries related to sensitive and emerging topics. |
| Approach: | They propose a method for topic-sensitive query set expansion using vector space interpolation. |
| Outcome: | The proposed method generates new queries about the sensitive topic by incorporating set diversity, which is not captured by traditional sentence-level augmentation methods such as paraphrasing or back-translation. |
Copied to clipboard
| Challenge: | X-RiSAWOZ dataset has more than 18,000 human-verified dialogue utterances for each language . Xiaoping and Xinhui are the main challenges for task-oriented dialogue research . |
| Approach: | They develop a toolkit to accelerate the post-editing of a new language dataset after translation . their dataset, code, and toolkit are released open-source . |
| Outcome: | The proposed toolkit accelerates the post-editing of a new language dataset after translation. |
Copied to clipboard
| Challenge: | a new study examines the extent and patterns of gaps in understandability of book reviews . 83% of the reviews had at least one culturally-specific difficult-to-understand element . |
| Approach: | They examine extent and patterns of gaps in understandability of book reviews . 83% of reviews had at least one culturally-specific difficult-to-understand element . authors say they have a significant scope for improvement . |
| Outcome: | The proposed approach improves the understanding of book reviews from different cultures . 83% of the reviews had at least one culturally-specific difficult-to-understand element . |
Copied to clipboard
| Challenge: | 'low resource' languages are understudied by the NLP community, while 'high resource' is referred to as 'achieved', while high-resource languages are referred . |
| Approach: | They qualitatively analyzed 150 papers from the ACL Anthology and popular speech-processing conferences that mention the keyword ‘low-resource. |
| Outcome: | The proposed analysis reveals that several interacting axes contribute to ‘low-resourceness’ of a language and why that makes it difficult to track progress for each individual language. |
Copied to clipboard
| Challenge: | Prior studies on LLM harms focus on Western concepts like race and gender, overlooking cultural concepts from other parts of the world. |
| Approach: | They propose a set of seven metrics to examine the presence of covert harms in LLM-generated conversations. |
| Outcome: | The proposed model detects that seven out of eight LLMs generated conversations riddled with CHAST, characterized by malign views expressed in seemingly neutral language, compared to Western ones such as race. |
Copied to clipboard
| Challenge: | Existing models are biased towards Western, Anglocentric or American cultures, a problem that is arguably detrimental to the performance of LLMs. |
| Approach: | They analyze more than 90 recent papers that aim to study cultural representation and inclusion in large language models. |
| Outcome: | The proposed models are biased towards Western, Anglocentric or American cultures, despite their diversity and their robustness. |
Copied to clipboard
| Challenge: | Recent work on code mixing in computational settings has leveraged social media code mixed texts to train NLP models. |
| Approach: | They propose to use language ID tags to measure syntactic variety in code-mixed text and their relationship with computational model performance. |
| Outcome: | The proposed measure can be applied to English(en)-hindi(hi) code-mixed datasets and compares them with other measures. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations. |
| Approach: | They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages. |
| Outcome: | The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages. |
Copied to clipboard
| Challenge: | Multilingual evaluation benchmarks usually contain limited high-resource languages and do not test models for specific linguistic capabilities. |
| Approach: | They propose an algorithm for automatically extracting target language CheckList templates from machine translated instances of a source language templates. |
| Outcome: | The proposed algorithm compares with CheckLists created with human verification in Hindi and 9 other languages. |
Copied to clipboard
| Challenge: | a small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications. |
| Approach: | They examine the relationship between types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time. |
| Outcome: | The proposed model will help to bridge the gap between languages and their resources and convince the ACL community to prioritise the resolution of the predicaments highlighted. |
Copied to clipboard
| Challenge: | Linguistic studies on code-switching focus on the "how" and "why" of CS . a new model aims to derive CS functions from local and global properties of the code-witched discourse . |
| Approach: | They propose a model that integrates CS phenomena and modalities into a representation that includes local and global properties of the code-switched discourse. |
| Outcome: | The proposed model simplifies the analysis of English/Hindi CS datasets and provides a flexible framework for further studies. |
Copied to clipboard
| Challenge: | Existing models for translation have not been systematically examined for their default ethical tendencies or their ability to employ and prioritize specified ethical approaches in conflicted translation situations. |
| Approach: | They propose a framework for examining ethical reasoning and implementation in large language models (LLMs) that systematically examines default ethical tendencies and their ability to employ and prioritize specified ethical approaches in conflicted translation situations. |
| Outcome: | The proposed framework examines the ethical reasoning and implementation of large language models in translation tasks. |
Copied to clipboard
| Challenge: | Existing large language models (LLMs) do not reliably classify verb language preferences to match native speaker judgments. |
| Approach: | They investigate whether large language models (LLMs) model linguistic variation by comparing Hindi-English verb code-mixing with English verb karna. |
| Outcome: | The proposed models do not reliably classify verb language preferences to match native speaker judgments, but with specific supervision, some models do predict human preference to an extent. |
Copied to clipboard
| Challenge: | Recent studies show multilingual contextual embedding models perform better on cross-lingual and multilingual tasks. |
| Approach: | They propose to evaluate multilingual contextual embedding models on multilingual data . they use language identification from text, POS tagging, Named Entity Recognition and Question Answering . |
| Outcome: | The proposed benchmark evaluates models on language identification from text, POS tagging, Named Entity Recognition, Question Answering and a new task for code-switching, Natural Language Inference. |
Copied to clipboard
| Challenge: | Existing MT systems are only useful for information assimilation, and require substantial manual post processing. |
| Approach: | They propose an Interactive Machine Translation interface that assists human translators with on-the-fly hints and suggestions. |
| Outcome: | The proposed interface makes the end-to-end translation process faster, more efficient and creates high-quality translations. |
Copied to clipboard
| Challenge: | Interactive Neural Machine Translation (INMT) systems can be used to promote data collection in several under-resourced languages, but are often not adapted to the deployment constraints native language speakers operate in. |
| Approach: | They propose to use interactive neural machine translation systems to promote data collection in several under-resourced languages by integrating three different modes of Internet-independent deployment and four assistive interfaces suitable for data-sparse languages. |
| Outcome: | The proposed model improves the data generation experience of community members along multiple axes without compromising on the quality of the generated translations. |
Copied to clipboard
| Challenge: | We argue that knowledge-retrieval and reasoning tasks are not ideal for measuring generalization, as LLMs are not trained for specific tasks. |
| Approach: | They propose a statistically motivated framework using personalization to assess generalization in Large Language Models. |
| Outcome: | The proposed framework outperforms existing models on movie and music recommendation datasets, but all models have room for improvement, especially Llama. |
Copied to clipboard
| Challenge: | Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. |
| Approach: | They propose a large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria. |
| Outcome: | The proposed dataset includes 30,000 tweets and a significant fraction of code-mixed tweets. |
Copied to clipboard
| Challenge: | Ethical aspects of research in language technologies have received much attention recently . do we observe a rise in formal ethical reviews of NLP studies? |
| Approach: | They conduct a qualitative and quantitative analysis of the ethics of NLP research . they compare the ethical reviews of NLAs to those of related disciplines . |
| Outcome: | The results compare the ACL Anthology to other related disciplines in the field . the results show that there is a heightened awareness of ethical issues that was previously lacking . |
Copied to clipboard
| Challenge: | Culturally Yours (CY) is a cultural reading assistant that helps users from diverse cultural backgrounds understand content from online sources that are written by people from a different culture. |
| Approach: | They propose to use culturally sensitive language to personalize a cultural reading assistant tool that can identify cultural-specific items for users from varying cultural contexts. |
| Outcome: | The tool personalizes to the user’s preferences based on the interaction of the user with the tool. |
Copied to clipboard
| Challenge: | Honorifics encode nuanced socio-pragmatic cues such as power, age, gender, fame, and cultural distance. |
| Approach: | They propose to study third-person honorific usage across 10,000 Hindi and Bengali Wikipedia articles . honorifics are more prevalent in Bengali than in Hindi, while non-honorifics dominate . |
| Outcome: | The authors show that large language models internalize similar socio-pragmatic norms . their analysis shows that honorifics are more prevalent in Bengali than in Hindi . |
Copied to clipboard
| Challenge: | Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text . |
| Approach: | a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities. |
| Outcome: | a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing. |
Copied to clipboard
| Challenge: | DUBLIN is a pixel-based visual document understanding model that does not rely on OCR. |
| Approach: | They propose a pixel-based visual document understanding model that does not rely on OCR. |
| Outcome: | The proposed model performs on extractive tasks such as DocVQA, InfoVQA and AI2D, and strong performance on abstraction datasets such as VisualMRC and text captioning. |
Copied to clipboard
| Challenge: | Automatic speech recognition systems have seen remarkable improvements in recent years, but evaluation of performance remains dependent on word and character error rate (WER/CER). |
| Approach: | They investigate how distribution shifts, model size and model architecture influence hallucination error rate (HER) HER is a metric used to quantify hallucinosity in automatic speech recognition systems. |
| Outcome: | The proposed model can be used to measure hallucination errors in high-stakes domains such as healthcare, legal, and aviation. |
Copied to clipboard
| Challenge: | MT errors are more pronounced in low-resourced languages where human translators are scarce and MT tools perform poorly. |
| Approach: | They propose to use a publicly available machine translation system to analyze machine translation errors in healthcare domains. |
| Outcome: | The proposed system reduces errors in two low-resourced languages for healthcare. |
Copied to clipboard
| Challenge: | Existing studies have shown that moral judgment depends on the language in which the dilemma is presented. |
| Approach: | They extend the work of beyond English, to 5 new languages (Chinese, Hindi, Russian, Spanish and Swahili) and probe three LLMs that show substantial multilingual text processing and generation abilities. |
| Outcome: | The models show substantial multilingual text processing and generation abilities. |
Copied to clipboard
| Challenge: | Socio-demographic prompting is a commonly employed approach to study cultural biases in LLMs as well as for aligning models to certain cultures. |
| Approach: | They propose to use socio-demographic prompting to probe four LLMs with culturally sensitive and non-sensitive cues on datasets that are supposed to be culturally neutral or sensitive. |
| Outcome: | The proposed model shows significant differences in responses on both kinds of datasets, casting doubt on its robustness. |
Copied to clipboard
| Challenge: | Existing training data for multilingual commonsense reasoning datasets is limited. |
| Approach: | They propose to use large language models for data augmentation in multilingual datasets . they use Dolly-v2, StableVicuna, ChatGPT, and GPT-4 to augment three datasets. |
| Outcome: | The proposed model outperforms larger general-purpose, zero-shot models when training in smaller models. |
Copied to clipboard
| Challenge: | Language models are inequitable at encoding and re-presentation, but there is much to be studied and criticism for the existing research that remains to be addressed. |
| Approach: | They propose to survey fairness in multilingual and non-English contexts . they argue that it is infeasible to achieve comprehensive coverage in terms of fairness datasets based on English . |
| Outcome: | The proposed methods are infeasible to scale across languages and cultures, the authors argue . they argue that the current methods are too narrowly focused on specific dimensions and types of biases and cannot scale across cultures. |
Copied to clipboard
| Challenge: | Responsible AI issues such as fairness, bias and toxicity will be discussed in this tutorial . |
| Approach: | This tutorial will describe various aspects of scaling up language technologies to many of the world’s languages by describing the latest research in Massively Multilingual Language Models (MMLMs). |
| Outcome: | This tutorial will cover various aspects of scaling up language technologies to many of the world's languages by describing the latest research in multilingual models. |
Copied to clipboard
| Challenge: | Code-mixing is a spoken language phenomenon and is difficult to train in multilingual communities. |
| Approach: | They propose a tool that can automatically generate code-mixed data given parallel data in two languages. |
| Outcome: | The proposed tool can generate code-mixed data in two languages using two linguistic theories. |
Copied to clipboard
| Challenge: | Large language models remain predominantly English-centric, which limits their utility for underrepresented languages. |
| Approach: | They propose to extend Llama’s vocabulary with 20% Hindi-specific tokens, thus halving Hindi tokenization fertility while preserving English efficiency. |
| Outcome: | The proposed models outperform open-weight models of comparable size on a 65B-token corpus and bilingual instruction and safety alignment on . a culturally grounded dataset. |
Copied to clipboard
| Challenge: | Existing music generation models are limited in their coverage of the musical genres and cultures of the world. |
| Approach: | They propose to use parametric fine tuning techniques to mitigat the bias in existing music datasets. |
| Outcome: | The proposed models are able to perform well across genres and cultures. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) generate ethically inappropriate texts even for seemingly innocuous contexts. |
| Approach: | They propose to use large language models to detect and filter toxic content in text prediction tasks by evaluating their toxicity detection approaches against a manually crafted CheckList of harms. |
| Outcome: | The proposed methods are compared against a checklist of harms targeted at different groups and different levels of severity in English. |
Copied to clipboard
| Challenge: | Despite progress in MT, a gap persists between how the technology is developed and how it is used in real-world contexts. |
| Approach: | They propose a human-centered approach to machine translation (MT) they argue that MT should be evaluated with diverse goals and contexts of use . |
| Outcome: | The proposed approach emphasizes alignment of evaluation and design with diverse communicative goals and contexts of use. |
Copied to clipboard
| Challenge: | a scalable approach to classify text with sensitivity is costly because of exponential time complexity. |
| Approach: | They propose a framework for calculating word-level local and global sensitivities . they use a CHECKLIST-generated sentiment analysis dataset to test their approach . |
| Outcome: | The proposed framework can be used to calculate word-level local and global sensitivities . it improves attacks by 15.58%, while using sensitivity as an additional reward improves . |
Copied to clipboard
| Challenge: | Existing methods to jailbreak large language models have been poorly studied . a recent study showed that non-expert users can jailbreak LLMs by manipulating their prompts . |
| Approach: | They propose a formalism and a taxonomy of known (and possible) jailbreaks . they propose generating a dataset of model outputs across 3700 jailbreak prompts a 'prompt' attack is a new attack popularly categorized as "prompting injection attacks" |
| Outcome: | The proposed model exploits 3700 jailbreak prompts over 4 tasks to analyze their effectiveness . authors show that the model can learn to perform a new task on unseen examples . |
Copied to clipboard
| Challenge: | Existing bilingual word embedding techniques are not ideal for code-mixed text processing and there is a need for learning multilingual word embeds from code-mixed texts. |
| Approach: | They propose to use bilingual word embedding techniques to train skip-grams on synthetic code-mixed text generated through linguistic models of code- mixing to perform two tasks. |
| Outcome: | The proposed embedding technique performs better on semantic and syntactic tasks than the existing embeddable techniques on sentiment analysis and POS tagging tasks. |
Copied to clipboard
| Challenge: | a large number of health-related queries require factually accurate answers . however, the lack of accurate answers may cause real-world harm . |
| Approach: | They evaluate the factual accuracy and consistency of large language models using a dataset consisting of health-related queries with varying degrees of presuppositions. |
| Outcome: | The proposed model responses agree with 23-32% of existing false claims and 49-55% with novel fabricated claims. |