Challenge: a new study examines the efficacy of large language models (LLMs) for Persian . ChatGPT and LLMs have shown remarkable performance in English, but their efficiency for low-resource languages remains an open question.
Approach: They present a benchmarking study of large language models (LLMs) for Persian . they focus on GPT-3.5-turbo, but also GPT-4 and OpenChat-3.5 .
Outcome: The proposed model performs better in Persian than other low-resource languages . the study is the first comprehensive benchmarking of large language models .

Similar Papers

Advancing Persian LLM Evaluation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing evaluation approaches for large language models in low-resource languages like Persian lack comprehensive frameworks, limiting their ability to assess models’ performance over a wide range of tasks requiring considerable cultural and contextual knowledge.
Approach: They propose to provide two new benchmarks to assess models' performance over a wide range of tasks requiring considerable cultural and contextual knowledge.
Outcome: The proposed benchmarks challenge the current state-of-the-art models’ abilities in a variety of Persian language comprehension tasks while reducing data contamination while providing an accurate assessment of Persian LLMs.
LAraBench: Benchmarking Arabic AI with Large Language Models (2024.eacl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research.
Approach: They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets.
Outcome: The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions.
A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets (2023.findings-acl)

Copied to clipboard

Challenge: Currently, the evaluation of large language models (LLMs) such as ChatGPT in academic datasets is difficult due to the difficulty of evaluating the generative outputs produced by this model against the ground truth.
Approach: They evaluate ChatGPT across 140 tasks and analyze 255K responses it generates in academic datasets.
Outcome: The proposed model performs well on 140 tasks and generates 255K responses in these datasets.
TounsiBench: Benchmarking Large Language Models for Tunisian Arabic (2025.emnlp-main)

Copied to clipboard

Challenge: a dataset of Tunisian Arabic instructions and prompts is used to evaluate LLMs' ability to understand and generate responses in Tunisia . we assess the quality, correctness, relevance, and dialectal adherence of LLM responses .
Approach: They propose a benchmark for evaluating the capabilities of large language models in Tunisian Arabic . they use a dataset of Tunisia Arabic instructions and prompts to evaluate their models .
Outcome: The proposed model can judge quality, correctness, relevance, and dialectal adherence . the model can also generate a leaderboard for the Tunisian Arabic language .
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond (2024.findings-naacl)

Copied to clipboard

Challenge: Existing LLM leaderboards often reference scores reported in other papers without consistent settings and prompts, which may encourage cherry-picking favored settings and for better results.
Approach: They propose an open-source and reproducible LLM evaluation suite built on top of OpenAI Evals that systematically evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories.
Outcome: The evaluation suite is built on top of OpenAI Evals and evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories.
Can Large Language Models Automatically Score Proficiency of Written Essays? (2024.lrec-main)

Copied to clipboard

Challenge: Automated essay scoring (AES) is one of the earliest research problems in natural language processing.
Approach: They propose to use large language models to analyze and score written essays using four different prompts.
Outcome: The proposed models show comparable performance on four different prompts and a slight advantage over the state-of-the-art models.
IRUEX: A Study on Large Language Models Problem-Solving Skills in Iran’s University Entrance Exam (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have profound implications for education.
Approach: They present a novel multiple-choice educational resource specifically designed to evaluate the performance of Large Language Models (LLMs) they use a dataset that contains 868 questions and 36,485 additional questions .
Outcome: The IRUEX dataset contains 868 questions and 36,485 additional questions.
ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have led to significant breakthroughs in the field.
Approach: They evaluate ChatGPT over multiple tasks with diverse languages and large datasets to provide more comprehensive information for multilingual NLP applications.
Outcome: The proposed model can process and generate texts for multiple languages due to its multilingual training data.
An Empirical Study of Instruction-tuning Large Language Models in Chinese (2023.findings-emnlp)

Copied to clipboard

Challenge: emergence of ChatGPT validates the potential of large language models (LLMs) in artificial general intelligence (AGI) however, the closed source of LLMs coupled with the requirement for massive computing resources has deterred researchers from reaching the LLM training stage.
Approach: They propose to use Chinese instruction-tuning LLMs as a cookbook for customizing LLM models that can better respond to Chinese instructions.
Outcome: The proposed LLM can be used to customize Chinese LLMs that can better respond to Chinese instructions.
Generalists vs. Specialists: Evaluating Large Language Models for Urdu (2024.findings-emnlp)

Copied to clipboard

Challenge: Urdu is underrepresented in natural language processing, yet it is underserved.
Approach: They compare general-purpose models with special-purpose ones that have been fine-tuned on specific tasks.
Outcome: The proposed models outperform general-purpose models on seven classification and seven generation tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations