Challenge: Word problem solving is a challenging and interesting task in NLP.
Approach: They propose to use equations to solve Hindi arithmetic word problems . they propose to also use equation equivalence to evaluate word problem solvers .
Outcome: The proposed dataset is based on 2336 arithmetic word problems in Hindi . it also includes baseline systems and evaluation techniques .

Similar Papers

Are NLP Models really able to Solve Simple Math Word Problems? (2021.naacl-main)

Copied to clipboard

Challenge: Existing solvers for math word problems often achieve high performance on benchmark datasets . existing models rely on shallow heuristics to achieve high accuracy .
Approach: They restrict their attention to English MWPs taught in grades four and lower . they propose a challenge dataset to test the accuracy of MWp solvers .
Outcome: The proposed model can solve a large fraction of MWPs even with shallow heuristics . the proposed model is much lower on the challenge dataset SVAMP .
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources (2026.acl-long)

Copied to clipboard

Challenge: Existing reviews focus on a few high-resource languages or embed Indian languages within broad multilingual settings, limiting coverage of low-resourced and culturally diverse varieties.
Approach: They present a unified survey of Indian NLP resources, covering 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks.
Outcome: The proposed survey covers 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks.
Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challenges (2025.findings-emnlp)

Copied to clipboard

Challenge: a survey examines the current efforts and challenges of NLP models for South Asian languages . there are more than 650 languages in South Asia, but many have very limited computational resources or are missing from existing models.
Approach: a survey examines efforts and challenges of NLP for South Asian languages . they focus on transformer-based models such as BERT, T5, & GPT . findings highlight substantial issues, including missing data in critical domains .
Outcome: The findings highlight significant issues, including missing data in critical domains . the survey aims to raise awareness within the NLP community for more targeted data curation .
CWID-hi: A Dataset for Complex Word Identification in Hindi Text (2022.lrec-1)

Copied to clipboard

Challenge: Text simplification is a method for improving the accessibility of text by converting complex sentences into simple sentences.
Approach: They propose to use Hindi knowledge annotators to capture the annotator’s language knowledge to build an automatic complex word classifier using a soft voting approach.
Outcome: The proposed dataset shows that native and non-native annotators perceive complex words differently depending on their language knowledge.
IndicFinNLP: Financial Natural Language Processing for Indian Languages (2024.lrec-main)

Copied to clipboard

Challenge: IndicFinNLP is a collection of 9 datasets relating to FinNLP for three Indian languages.
Approach: They propose to use financial NLP to detect exaggerated numerals in financial texts written in Hindi, Bengali, and Telugu.
Outcome: The proposed framework detects exaggerated numerals in financial texts written in Hindi, Bengali, and Telugu.
A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers (2020.acl-main)

Copied to clipboard

Challenge: Existing MWP corpora are limited in language patterns and problem types . a new corpus of 2,305 MWps is proposed that is more diverse in terms of lexicon usage .
Approach: They propose to use ASDiv to measure lexicon usage diversity of a given MWP corpus.
Outcome: The proposed corpus covers more problem types and text patterns than existing corpora and reflects the true capability of solvers more faithfully.
ArMATH: a Dataset for Solving Arabic Math Word Problems (2022.lrec-1)

Copied to clipboard

Challenge: This paper is the first to use deep learning methods to solve Arabic MWPs . it is also the first study to use transfer learning to solve MWp across different languages .
Approach: They contribute to the first large-scale dataset for Arabic Math Word Problems . they use deep learning methods to solve Arabic MWPs and a transfer learning model to promote performance .
Outcome: The proposed model improves Arabic MWP solvers by 3% over the existing model.
A Platform for Event Extraction in Hindi (2020.lrec-1)

Copied to clipboard

Challenge: Event Extraction is an important task in the widespread field of NLP, but there is no benchmark setup in Hindi.
Approach: They propose an Event Extraction framework for Hindi language and develop deep learning based models to set as the baselines.
Outcome: The proposed framework crawls more than seventeen hundred disaster related Hindi news articles from various news sources.
WARM: A Weakly (+Semi) Supervised Math Word Problem Solver (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to solving math word problems require full supervision in the form of intermediate equations.
Approach: They propose a weakly supervised model that requires only the final answer as supervision to solve math word problems.
Outcome: The proposed model achieves accuracy gains of 4.5% and 32% over current weakly-supervised methods on standard Math23K and AllArith datasets.
Hi-GEC: Hindi Grammar Error Correction in Low Resource Scenario (2025.coling-main)

Copied to clipboard

Challenge: Automated Grammatical Error Correction (GEC) is a scarcely explored low-resource language . a recent study focused on English, but it focused on Hindi, which presents unique challenges due to its complex syntax and intricate morphology.
Approach: They propose to use a human-edited dataset to generate Hindi GEC data . they also investigate round trip translation using diverse languages for the technique .
Outcome: The proposed method outperforms other methods in Hindi, showing that it is highly efficient.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations