Matina: A Large-Scale 73B Token Persian Text Corpus (2025.naacl-long)

Copied to clipboard

Challenge: Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian.
Approach: They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality.
Outcome: The proposed model performs well on key Persian NLP tasks.

Similar Papers

Matina: A Culturally-Aligned Persian Language Model Using Multiple LoRA Experts (2025.findings-acl)

Copied to clipboard

Challenge: Existing Large language models fail to accurately model underrepresented languages and cultures, limiting their applicability and acceptance.
Approach: They develop a Persian-focused multi-expert model that incorporates Iranian cultural values and linguistic structures.
Outcome: The proposed model outperforms baseline models in task performance and user satisfaction.
MirasText: An Automatically Generated Text Corpus for Persian (L18-1)

Copied to clipboard

Challenge: Natural language processing is one of the most important fields of artificial intelligence.
Approach: They propose to use MirasText to generate Persian text corpus from Persian websites . MiraSText has over 2.8 million documents and over 1.4 billion tokens .
Outcome: The generated corpus has over 2.8 million documents and over 1.4 billion tokens . MirasText has over 800 billion token tokens and more than 300 thousand articles .
Parsivar: A Language Processing Toolkit for Persian (L18-1)

Copied to clipboard

Challenge: a preprocessing step is required to convert text into a standard format for NLP tasks.
Approach: They propose a Persian preprocessing toolkit that performs various kinds of activities . they use a plagiarism detection system to exploit the proposed toolkit .
Outcome: The proposed tool outperforms available Persian preprocessing tools by about 8 percent in terms of F1 . the proposed toolkit performs normalization, space correction, tokenization, stemming, parts of speech tagging and shallow parsing tasks.
FaMTEB: Massive Text Embedding Benchmark in Persian Language (2025.findings-emnlp)

Copied to clipboard

Challenge: a comprehensive benchmark for Persian text embeddings is built upon the Massive Text Embedding Benchmark (MTEB) 63 datasets are included in the benchmark, including a novel task of summary retrieval.
Approach: They propose a benchmark for Persian (Farsi) text embeddings built upon the Massive Text Embedding Benchmark.
Outcome: The proposed framework includes 63 datasets spanning seven different tasks . the evaluation datasets were rigorously evaluated by humans and automated systems .
Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) are mainly trained on English data and struggle with low-resource languages.
Approach: They propose to add a new language to Llama to improve classification accuracy for Persian tasks by aligning representations through bilingual pretraining and instruction datasets.
Outcome: The proposed model performs on generation and classification tasks with no adverse impact and sometimes even improvements on English tasks.
LSCP: Enhanced Large Scale Colloquial Persian Language Understanding (2020.lrec-1)

Copied to clipboard

Challenge: a gap exists in describing low-resource formal languages such as Persian . a large scale corpus of 120M sentences is proposed to fill this gap .
Approach: They propose to target a gap in describing the colloquial language for low-resource ones such as Persian . a large scale Persian corpus is hierarchically organized in a semantic taxonomy .
Outcome: The proposed corpus consists of 120M sentences from 27M tweets annotated with parsing tree, part-of-speech tags, sentiment polarity and translation in five different languages.
IEPile: Unearthing Large Scale Schema-Conditioned Information Extraction Corpus (2024.acl-short)

Copied to clipboard

Challenge: Large Language Models exhibit a significant performance gap in Information Extraction (IE) high-quality instruction data is the vital key for enhancing LLMs' specific capabilities .
Approach: They propose a bilingual (English and Chinese) IE instruction corpus that contains 0.32B tokens.
Outcome: The proposed model improves the performance of LLMs for IE with zero-shot generalization.
AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing (2023.findings-acl)

Copied to clipboard

Challenge: Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP).
Approach: They present AraMUS, the largest Arabic PLM with 11B parameters trained on 529GB of high-quality Arabic textual data.
Outcome: The proposed model achieves state-of-the-art performance on a diverse set of Arabic classification and generative tasks.
Transformers for Bridging Persian Dialects: Transliteration Model for Tajiki and Iranian Scripts (2024.lrec-main)

Copied to clipboard

Challenge: Despite its profound linguistic and cultural significance, Tajiki Persian remains a low-resource language with scant digitized datasets for computational applications.
Approach: They propose to use Shahnameh, a seminal Persian epic poem, to train and assess Tajiki Persian transliteration models using two prominent sequence-to-sequence architectures: GRU with attention and transformer.
Outcome: The proposed model outperforms pre-trained models with attention and transformer.
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT (2024.lrec-main)

Copied to clipboard

Challenge: a new study examines the efficacy of large language models (LLMs) for Persian . ChatGPT and LLMs have shown remarkable performance in English, but their efficiency for low-resource languages remains an open question.
Approach: They present a benchmarking study of large language models (LLMs) for Persian . they focus on GPT-3.5-turbo, but also GPT-4 and OpenChat-3.5 .
Outcome: The proposed model performs better in Persian than other low-resource languages . the study is the first comprehensive benchmarking of large language models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations