Papers with Bengali
Copied to clipboard
| Challenge: | Automated headline generation systems have the potential to assist editors in finding interesting headlines to attract visitors or readers. |
| Approach: | They propose to use Bengali news article-headline pairings with auxiliary data to better model headline generation using pre-trained language models. |
| Outcome: | The proposed model improves on a Bengali news headline generation dataset by 3 to 10 percentage points over baselines. |
Copied to clipboard
| Challenge: | Existing studies on the effectiveness of different architectures for modeling low resource languages are limited. |
| Approach: | They propose a trainable memory efficient CNN architecture for Bengali and Hindi . they propose two learnable convolutional sub-models that are end to end trainable . |
| Outcome: | The proposed model outperforms pretrained BERT models on Bengali and Hindi with 16X less parameters and achieves much better performance than SOTA LSTMs on multiple real-world datasets. |
Copied to clipboard
| Challenge: | TableQA is the task of answering questions over tables of structured information, returning individual cells or tables as output. |
| Approach: | They propose a fully automatic large-scale tableQA data generation process for low-resource languages with limited budget. |
| Outcome: | The proposed method outperforms state-of-the-art LLMs on two Indic languages with no tableQA datasets and models on different aspects including mathematical reasoning capabilities and zero-shot cross-lingual transfer. |
Copied to clipboard
| Challenge: | Existing methods to classify Bengali text into six basic emotions are infancy for resource-constrained languages like English, Arabic, Chinese and French. |
| Approach: | They propose a transformer-based technique to classify Bengali text into one of the six basic emotions: anger, fear, disgust, sadness, joy, and surprise. |
| Outcome: | The proposed technique outperforms all other techniques by achieving highest weighted f_1-score on the test data. |
Copied to clipboard
| Challenge: | Pretrained multilingual translation models with massive coverage are becoming of the backbone of many translation systems. |
| Approach: | They propose to use a gradient-based inference-time controller to control a pretrained multilingual model by using a model with attribute annotations. |
| Outcome: | The proposed model performs well on pretrained multilingual models and is attribute- rather than language-specific. |
Copied to clipboard
| Challenge: | Existing research on hate speech detection in English does not cover low-resource languages like Bengali. |
| Approach: | They develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets. |
| Outcome: | The proposed model outperforms other models on training actual and romanized datasets by interpreting the semantic expressions better. |
Copied to clipboard
| Challenge: | Open-Domain Generative Question Answering has achieved impressive performance in English . combining document-level retrieval with answer generation can generate complete sentences . |
| Approach: | They propose an open-domain approach that combines document retrieval with answer generation to generate complete sentences in English . they propose a cross-lingual generative model that exploits passages written in multiple languages . |
| Outcome: | The proposed model outperforms answer sentence selection baselines for all 5 languages and monolingual pipelines for three out of five languages. |
Copied to clipboard
| Challenge: | Neural machine translation systems are vulnerable when trained on limited data. |
| Approach: | They propose to add noise to the training phase to increase robustness of NMT systems trained on limited data. |
| Outcome: | The proposed training strategy overcomes noise and improves robustness for low-resource tasks for abugida glyphs. |
Copied to clipboard
| Challenge: | LLM judges are often used to score generated answers, but their decisions may be affected by surface style rather than semantic correctness. |
| Approach: | They propose a benchmark to examine multilingual hedging effects in LLM evaluation . they find that when two answers are equally correct, the judge prefers the assertive answer . |
| Outcome: | The proposed benchmark shows that judge preferences are based on style and style . the results highlight that benchmark auditing is a key requirement for judge-bias research . |
Copied to clipboard
| Challenge: | Text style transfer (TST) is a multidimensional task requiring the assessment of style transfer accuracy, content preservation, and naturalness. |
| Approach: | They propose to use text style transfer metrics to evaluate outputs of text editors . they also investigate the potential of large language models as tools for TST evaluation . |
| Outcome: | The proposed methods provide better insights than existing metrics, the authors show . their meta-evaluation through correlation with hu-man judgments shows they are effective . |
Copied to clipboard
| Challenge: | IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages. |
| Approach: | They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families. |
| Outcome: | Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages. |
Copied to clipboard
| Challenge: | Bengali is the sixth most spoken language in the world, but handwritten text recognition systems for the language are underdeveloped. |
| Approach: | They propose a Bengali handwritten text recognition system that uses a decoder-only transformer to address the unique challenges of Bengali script. |
| Outcome: | The proposed system significantly improves on existing tokenizers on Bengali script. |
Copied to clipboard
| Challenge: | a dataset of parallel Bengali and English exam questions is used to compare LLMs in low-resource languages. |
| Approach: | They introduce BEnQA, a dataset comprising parallel Bengali and English exam questions . they benchmark several Large Language Models with their parallel dataset and observe performance disparity . |
| Outcome: | The proposed dataset consists of 5K questions covering several subjects in science . the authors find that the models perform poorly in Bengali and English . |
Copied to clipboard
| Challenge: | a new generation of English-oriented Large Language Models significantly outperforms older LLMs on low-resource languages. |
| Approach: | They compare Bengali-oriented LLMs with open-weight and closed-source LLM models . they conclude that there is a need for a Bengali model, but lacks high-quality pretraining data . |
| Outcome: | The proposed model outperforms existing models on Bengali on low-resource languages . the results highlight biases in machine-translated datasets used for Bengali NLP tasks . |
Copied to clipboard
| Challenge: | Gesture typing is a method of typing words on a touch-based keyboard by drawing a continuous trace passing through the relevant keys. |
| Approach: | They propose a keyboard that supports gesture typing in Indic languages by drawing a continuous trace over the keyboard and the finger needs to be lifted only once a word is completed. |
| Outcome: | The proposed model performs path decoding, transliteration and transliterations correction. |
Copied to clipboard
| Challenge: | Large Language Models have been used for sentiment analysis, machine translation, and question answering, but their effectiveness in the multilingual financial domain remains unknown. |
| Approach: | They propose a fine-tuning approach that integrates positive and negative rationales alongside classification labels. |
| Outcome: | The proposed approach outperforms existing methods across English, Hindi, Bengali, and Telugu, and is suitable for industry applications. |
Copied to clipboard
| Challenge: | NLP is a technique that generates counterspeech that “counters” the vicious tone of online abuse and dilutes/ameliorates their rippling effect over the social network. |
| Approach: | They propose to use neural architectures to generate counterspeech that can "counter" the vicious tone of online abuse and dilute/ameliorate their rippling effect over the social network. |
| Outcome: | The proposed model can generate counterspeech in monolingual setups and is more transferable when languages belong to the same language family. |
Copied to clipboard
| Challenge: | a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages . |
| Approach: | They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation . |
| Outcome: | The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU . |
Copied to clipboard
| Challenge: | a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages. |
| Approach: | They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction . |
| Outcome: | The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms. |
Copied to clipboard
| Challenge: | a long line of work suggests that LLMs face issues along both dimensions . |
| Approach: | They investigate hallucination detectors' failure modes and their effects on the task accuracy of four LLMs and three halluciner detectors. |
| Outcome: | The models show impressive performance in high-resource languages like English but the performance degrades significantly in low-resourced languages like Bengali. |
Copied to clipboard
| Challenge: | Currently, there is a lack of data and technology for resource-poor languages in developing countries like India. |
| Approach: | They propose to use two different datasets to analyze query intents and entities in healthcare. |
| Outcome: | The proposed model is useful to identify query intents and entities in real-world scenarios. |
Copied to clipboard
| Challenge: | Recent research on memes’ detrimental facets is skewed towards high-resource languages, such as Bengali. |
| Approach: | They propose a dataset MIMOSA that annotates annotated memes across five aggression target categories in Bengali and propose 'Multimodal Attentive Fusion' to detect aggression targets. |
| Outcome: | The proposed method outperforms state-of-the-art methods in Bengali and in low-resource languages. |
Copied to clipboard
| Challenge: | Recent studies on sentiment analysis of memes have focused on English, but there is a significant barrier to performing multimodal sentiment analysis research in resource-constrained languages like Bengali. |
| Approach: | They propose to use a Bengali dataset to perform multimodal sentiment analysis in low resource languages. |
| Outcome: | The proposed dataset for Bengali contains 4417 memes with three annotated labels positive, negative, and neutral. |
Copied to clipboard
| Challenge: | Bangla is underrepresented in KGs due to lack of comprehensive datasets, encoders, NER models, part-of-speech taggers, and lemmatizers. |
| Approach: | Bangla is underrepresented in KGs due to lack of comprehensive datasets, encoders, NER models, part-of-speech taggers, and lemmatizers. authors propose a framework that can automatically construct Bengali KG from any Bangla text. |
| Outcome: | The proposed framework can automatically construct Bengali KGs from any Bangla text. |
Copied to clipboard
| Challenge: | despite being the seventh most widely spoken language, Bengali has received little attention in machine translation due to being low in resources. |
| Approach: | They propose a customized sentence segmenter for Bengali and two new methods for parallel corpus creation on low-resource setups. |
| Outcome: | The proposed method improves Bengali-English parallel corpus by 9 BLEU over previous approaches . the results will pave the way for future research on Bengali and other low-resource languages . |
Copied to clipboard
| Challenge: | a recent study has not examined their generalizability between formats, cultures, and genders. |
| Approach: | They evaluate large language models (LLMs) and small LLMs at clinical de-identification . they show that smaller models achieve comparable performance while substantially reducing inference cost . |
| Outcome: | The proposed models outperform larger models in de-identification tasks with limited data . the models can be fine-tuned with limited datasets to outperformed larger models . |
Copied to clipboard
| Challenge: | Abstractive summarization systems are difficult to perform due to the unavailability of the parallel data for low-resource languages like Bengali. |
| Approach: | They propose a graph-based unsupervised abstractive summarization system in Bengali text documents that requires only a Part-Of-Speech (POS) tagger and a pre-trained language model trained on Bengali texts. |
| Outcome: | The proposed system outperforms baselines without human-annotated reference summaries on a human-random dataset with Bengali text. |
Copied to clipboard
| Challenge: | Existing studies on media framing have focused on English only data, leaving a gap in research concerning multilingual contexts. |
| Approach: | They propose to use crowd-sourced datasets to automate framing analysis by automating translation and annotation. |
| Outcome: | The proposed system improves on existing models in Bengali and Portuguese . the proposed system can train on a crowd-sourced dataset in 12 languages . |
Copied to clipboard
| Challenge: | Existing QA benchmarks do not account for errors that speech recognition models might introduce . evaluating production-ready QA systems on data that is not representative of real-world inputs is problematic . |
| Approach: | They construct a multi-dialect, spoken QA benchmark on five languages with 68k audio prompts in 24 dialects from 255 speakers. |
| Outcome: | The proposed model is based on 68k audio prompts in 24 dialects from 255 speakers. |
Copied to clipboard
| Challenge: | Existing methods to fact-check content are not scaled well in non-English contexts. |
| Approach: | They propose to use a WhatsApp tipline and public group message dataset to find pairs of textual messages containing claims that can be served with one fact-check. |
| Outcome: | The proposed model outperforms existing models in English, Hindi, and Tamil in all settings. |
Copied to clipboard
| Challenge: | Existing methods for summarizing educational videos in Bengali are limited due to the rapid growth of educational video content. |
| Approach: | They propose an end-to-end pipeline for the abstractive summarization of Bengali videos . they fine-tuned the BanglaT5 model on a new benchmark dataset . |
| Outcome: | The proposed system preprocesses audio and converts speech to text using Google's Speech Recognition API. |
Copied to clipboard
| Challenge: | Existing approaches to generalization to resource-rich languages are difficult . a recent study shows that word representations can be useful in low resource languages . |
| Approach: | They propose two approaches for improving generalization to low-resource languages by adapting continuous word representations using linguistically motivated subword units. |
| Outcome: | The proposed method improves generalization to low resource languages . it requires neither parallel corpora nor bilingual dictionaries and requires no parallel training . |
Copied to clipboard
| Challenge: | Existing approaches to few-shot Question Generation (QG) are limited and require manual annotation. |
| Approach: | They propose to use multilingual BERT to perform few-shot question generation with cross-lingual transfer. |
| Outcome: | The proposed model improves in few-shot QG and human evaluation confirms it. |
Copied to clipboard
| Challenge: | Cognates are words that have a common etymological origin and can facilitate the Second Language Acquisition (SLA) however, they also pose a challenge to various NLP applications such as Machine Translation and Cross-lingual Sense Disambiguation. |
| Approach: | They create two cognate datasets for twelve Indian languages and use them to generate cognate sets. |
| Outcome: | The proposed datasets are curated using previously available baseline cognate detection approaches and evaluated with the help of lexicographers. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is used in a variety of downstream tasks in the biomedical domain, but is difficult when working with consumer health questions (CHQs). |
| Approach: | They propose to use a dataset to identify named entities in health-related texts in Bengali to address the scarcity of available data. |
| Outcome: | The proposed dataset captures the diverse range of linguistic styles and dialects used by native speakers from various regions in their day-to-day lives. |
Copied to clipboard
| Challenge: | Disfluency correction models can help alleviate this problem, but the unavailability of labeled data in low-resource languages impairs progress. |
| Approach: | They propose to use a pretrained multilingual model to detect zero-shot disfluency in Indian languages. |
| Outcome: | The proposed model achieves F1 scores of 75 and higher on five disfluency types across four languages. |
Copied to clipboard
| Challenge: | linguistics and morphology of resource-short code-mixed texts remain a key challenge in text processing. |
| Approach: | They propose a hierarchical transformer-based framework that captures the semantic relationship among words and hierarchically learns sentencelevel semantics using a fused attention mechanism. |
| Outcome: | The proposed framework improves on one European and five Indic languages on four NLP tasks on eleven datasets. |
Copied to clipboard
| Challenge: | Word sense disambiguation is a widely studied NLP task of identifying the meaning of a word in context. |
| Approach: | They propose a method to create parallel sense-annotated datasets in English . they use machine translation, word alignment, sense projection, and sense filtering to produce silver annotations . |
| Outcome: | The proposed method produces parallel sense-annotated datasets on Farsi, Chinese, and Bengali . the results are higher than those obtained with recent multilingual systems, the authors say . |
Copied to clipboard
| Challenge: | Evaluation of multilingual Large Language Models is challenging due to a variety of factors including the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and lack of local, cultural nuances in translated benchmarks. |
| Approach: | They evaluate 30 models across 10 Indic languages by conducting 90K human evaluations and 30K LLM-based evaluations. |
| Outcome: | The proposed models perform best in most Indic languages, while the agreement drops for direct assessment especially for Bengali and Odia. |
Copied to clipboard
| Challenge: | a growing body of research has focused on the negative aspects of memes in high-resource languages like Bengali . a new dataset for Bengali hateful memes is designed to detect their targeted entities . |
| Approach: | They propose a multimodal dataset that analyzes the modality of memes and compares them with other datasets. |
| Outcome: | The proposed dataset outperforms state-of-the-art datasets on Bengali hateful memes . the proposed dataset is generalizable on other low-resource hateful memes datasets compared with baselines based on the proposed model . |
Copied to clipboard
| Challenge: | Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency. |
| Approach: | They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages. |
| Outcome: | The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health). |
Copied to clipboard
| Challenge: | Several studies investigating methods to detect offensive content in social media use English data. |
| Approach: | They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources. |
| Outcome: | The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish. |
Copied to clipboard
| Challenge: | Existing research on LLM biases has focused on direct questioning or general-purpose settings . pronounced behavioral biase despite their growing deployment in financial analysis, forecasting, and decision support. |
| Approach: | They propose a benchmark to evaluate behavioral biases of large language models in MFMD . they use a multilingual financial misinformation dataset to integrate these with misinformation claims . |
| Outcome: | The proposed benchmark evaluates behavioral biases of large language models across economic scenarios. |
Copied to clipboard
| Challenge: | Large-scale multitask benchmarks have driven rapid progress in language modeling, yet most emphasize low-resource languages like English. |
| Approach: | They propose a benchmark for massive multitask language understanding in Bengali . they use a dataset that preserves mathematical content via MathML and a subset of questions most frequently missed by top systems to stress difficult cases. |
| Outcome: | The proposed benchmark covers 24 model variants across 11 LLM families. |
Copied to clipboard
| Challenge: | Existing open-domain dialogue systems suffer from data scarcity due to unavailability of high-quality datasets for low-resource languages like Bengali. |
| Approach: | They propose to prepare large-scale open-domain dialogue datasets from podcasts and talk-shows and label them based on weak-supervision techniques. |
| Outcome: | The proposed corpus improves performance of large language models in case of downstream classification tasks during fine-tuning. |
Copied to clipboard
| Challenge: | A study of multilingual fine-tuning yields better performance on downstream NLP applications . low resource languages such as Oriya and Punjabi are found to be the largest beneficiaries of multi-lingual fine tuning. |
| Approach: | They propose to leverage the relatedness of languages that belong to the same family in NLP models by multilingual fine-tuning. |
| Outcome: | The proposed approach improves performance on downstream NLP tasks by 15% compared to monolingual fine-tuning. |
Copied to clipboard
| Challenge: | Bengali and Hindi are low-resource languages, and the state-of-the-art tokenization methods fail to separate roots from affixes. |
| Approach: | They used BERT and Wordpiece tokenizers to train a wordpiece tokenization system for Bengali and Hindi to model fine-grained character-level information. |
| Outcome: | The proposed tokenizers outperform the state-of-the-art and Wordpiece tokenizer for modeling Bengali and Hindi. |
Copied to clipboard
| Challenge: | IndicFinNLP is a collection of 9 datasets relating to FinNLP for three Indian languages. |
| Approach: | They propose to use financial NLP to detect exaggerated numerals in financial texts written in Hindi, Bengali, and Telugu. |
| Outcome: | The proposed framework detects exaggerated numerals in financial texts written in Hindi, Bengali, and Telugu. |
Copied to clipboard
| Challenge: | India has 22 languages, each of them being spoken by over a million people . the current state of the art text-to-speech systems for Indian languages are lacking in the multimedia domain . |
| Approach: | They propose to train a state-of-the-art TTS system for Hindi, Malayalam and Bengali and publish the results. |
| Outcome: | The proposed system trains neural text-to-speech systems for Hindi, Malayalam and Bengali and makes them publicly available. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in general-purpose tasks but struggle with numerical reasoning, especially in low-resource languages like Bengali. |
| Approach: | They propose a benchmark to assess LLMs on numerical reasoning tasks in Bengali. |
| Outcome: | The proposed benchmark assesses LLMs on numerical reasoning tasks in Bengali. |
Copied to clipboard
| Challenge: | Existing approaches to improve the quality of responses generated by large language models (LLMs) however, these critique-refine steps require multiple expensive LLM calls. |
| Approach: | They propose to use critique distillation to train critic models that are trained on input-critique pairs generated by an LLM. |
| Outcome: | The proposed model trains two separate critics that focus on lexical and structure complexity, and is more effective than using an LLM directly as a critic in both 0-shot and few-shot settings. |
Copied to clipboard
| Challenge: | a number of studies have tried to detect and control the spread of such abusive memes on social media platforms. |
| Approach: | They build a Bengali meme dataset to test models for abusive memes . they find that multimodal models that use both textual and visual information outperform unimodal models . |
| Outcome: | The proposed model outperforms unimodal models in a Bengali meme dataset. |
Copied to clipboard
| Challenge: | Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content. |
| Approach: | They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set . |
| Outcome: | The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu. |
Copied to clipboard
| Challenge: | despite growing interest in GEC, most research has focused on English due to the lack of benchmark datasets for low-resource lan-guages. |
| Approach: | They propose a new approach to generate high-quality synthetic data for GEC using monolingual corpora. |
| Outcome: | The proposed framework outperforms other monolingual methods in English, Hindi, Bengali, Marathi, and Tamil. |
Copied to clipboard
| Challenge: | Existing benchmarks for numerical reasoning in multilingual Indic languages are inadequate . e.g., FinVQA is a framework for evaluating financial numerical reasoning . |
| Approach: | They propose a framework that combines supervised fine-tuning with constraint-aware decoding to promote faithful numerical reasoning. |
| Outcome: | The proposed framework spans English, Hindi, Bengali, Marathi, Gujarati, and Tamil . it combines supervised fine-tuning with constraint-aware decoding to promote faithful numerical reasoning . |
Copied to clipboard
| Challenge: | idioms provide a fascinating gateway to creativity, cultural values, historical context, and diverse perspectives inherent to diverse linguistic traditions. |
| Approach: | They propose a multimodal idiom corpus enriched with seven idiomatic tones . they propose idiomic hybridization framework that embeds multiple idiomatic expert opinions . |
| Outcome: | The proposed framework achieves 5–6% performance gains across advanced vision language models. |
Copied to clipboard
| Challenge: | Existing LLMs either reason in English and translate, or simply fail on multi-step Bengali math. |
| Approach: | They propose a Bengali mathematical reasoning model called GanitLLM with a difficulty-aware Bengali math corpus and a curriculum-based GRPO pipeline. |
| Outcome: | The proposed model improves on Bn-MGSM and Bn MSVAMP by +8 and +7 accuracy points while increasing the percentage of Bengali reasoning tokens from 14% to over 88% and reducing solution length from 943 to 193 words. |
Copied to clipboard
| Challenge: | Spontaneous speech is rarely fluent, and disfluencies can degrade readability and reliability . a sequence tagger first marks disfluent tokens, and these signals guide instruction fine-tuning . |
| Approach: | They propose a multilingual correction pipeline where a sequence tagger first marks disfluent tokens . they add a contrastive learning objective that penalizes the reproduction of disfluency tokens. |
| Outcome: | The proposed model improves readability and reliability of ASR transcripts in three languages . disfluencies can cause misinterpretations, incoherent responses, poor user experience . |