Papers by MohammadAli SadraeiJavaheri

3 papers
MorphBPE: Morphology-Aware Tokenization for Efficient LLM Training (2026.findings-acl)

Copied to clipboard

Challenge: Tokenization is a key design choice in modern NLP systems and a critical bottleneck for multilingual Large Language Models.
Approach: They propose a tokenization extension that constrains merge operations to respect morpheme boundaries while preserving inference.
Outcome: The proposed tokenization improves morphological coherence and language model cross-entropy in four languages.
Transformers for Bridging Persian Dialects: Transliteration Model for Tajiki and Iranian Scripts (2024.lrec-main)

Copied to clipboard

Challenge: Despite its profound linguistic and cultural significance, Tajiki Persian remains a low-resource language with scant digitized datasets for computational applications.
Approach: They propose to use Shahnameh, a seminal Persian epic poem, to train and assess Tajiki Persian transliteration models using two prominent sequence-to-sequence architectures: GRU with attention and transformer.
Outcome: The proposed model outperforms pre-trained models with attention and transformer.
The Touché23-ValueEval Dataset for Identifying Human Values behind Arguments (2024.lrec-main)

Copied to clipboard

Challenge: Cultural norms can influence the prioritization of values, leading to distinct perspectives on debatable topics.
Approach: They present a Touché23-ValueEval dataset that annotates 4780 new arguments and annotated 54 human values.
Outcome: The Touché23-ValueEval dataset doubles the original Webis-ArgValués-22 dataset to 9324 arguments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations