Papers by Hamed Baghbani

    1 papers
    Matina: A Large-Scale 73B Token Persian Text Corpus (2025.naacl-long)

    Copied to clipboard

    Challenge: Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian.
    Approach: They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality.
    Outcome: The proposed model performs well on key Persian NLP tasks.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations