Papers by Abdullah Barayan

2 papers
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)

Copied to clipboard

Challenge: Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI.
Approach: They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages.
Outcome: The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment.
Analysing Zero-Shot Readability-Controlled Sentence Simplification (2025.coling-main)

Copied to clipboard

Challenge: Text simplification (RCTS) models often depend on parallel corpora with readability annotations on both source and target sides.
Approach: They propose to use instruction-tuned large language models for zero-shot RCTS to reduce reliance on parallel corpora with readability annotations on both source and target sides.
Outcome: The proposed model can generate sentences with the desired readability, but the model's limitations and characteristics of the source sentences impede it.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations