Papers by Fatemeh Taherinezhad

2 papers
Matina: A Large-Scale 73B Token Persian Text Corpus (2025.naacl-long)

Copied to clipboard

Challenge: Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian.
Approach: They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality.
Outcome: The proposed model performs well on key Persian NLP tasks.
Matina: A Culturally-Aligned Persian Language Model Using Multiple LoRA Experts (2025.findings-acl)

Copied to clipboard

Challenge: Existing Large language models fail to accurately model underrepresented languages and cultures, limiting their applicability and acceptance.
Approach: They develop a Persian-focused multi-expert model that incorporates Iranian cultural values and linguistic structures.
Outcome: The proposed model outperforms baseline models in task performance and user satisfaction.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations