Hansi Hettiarachchi, Damith Premasiri, Lasitha Randunu Chandrakantha Uyangodage, Tharindu Ranasinghe
| Challenge: | introducing large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources. |
| Approach: | They propose a large news corpus for Sinhala with a set of NLP tasks for the language . NSina is the largest news corpuse for Sinha, available up to date . |
| Outcome: | The proposed model outperforms existing models in many benchmarks and outperformed previous models in high-resource languages. |
Similar Papers
Sinhala Encoder-only Language Models and Evaluation (2025.acl-long)
Copied to clipboard
Tharindu Ranasinghe, Hansi Hettiarachchi, Nadeesha Chathurangi Naradde Vidana Pathirana, Damith Premasiri, Lasitha Uyangodage, Isuri Nanomi Arachchige, Alistair Plum, Paul Rayson, Ruslan Mitkov
| Challenge: | Recent advances in language models (LMs) have produced excellent results in many NLP tasks, but their effectiveness is highly dependent on available pre-training resources. |
| Approach: | They propose to collect the largest monolingual corpus for Sinhala and compile a benchmark and evaluate LMs on it. |
| Outcome: | The proposed language models outperform the popular multilingual LMs in downstream NLP tasks. |
BERTifying Sinhala - A Comprehensive Analysis of Pre-trained Language Models for Sinhala Text Classification (2022.lrec-1)
Copied to clipboard
| Challenge: | Large-scale monolingual pre-trained language models have shown promising results for high-resource as well as lowresource languages, especially for text classification. |
| Approach: | They provide a set of recommendations for using pre-trained models for Sinhala text classification and introduce new annotated datasets useful for future research. |
| Outcome: | The proposed models are far superior to existing models for Sinhala and set a strong baseline for text classification when fine-tuned. |
A Systematic Approach to Derive a Refined Speech Corpus for Sinhala (2022.lrec-1)
Copied to clipboard
| Challenge: | Despite being large and generic, some languages such as Sinhala are left to underutilize the technology due to the lack of adequate resources. |
| Approach: | They propose to derive a corpus from a publicly available corpus for Sinhala speech recognition using crowdsourcing and web scraping techniques. |
| Outcome: | The proposed corpus reduces the Word-Error-Rate by 15.9%. |
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala (2025.emnlp-main)
Copied to clipboard
Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda, Chamila Liyanage, Yusuke Sakai, Randil Pushpananda, Ruvan Weerasinghe, Hidetaka Kamigaito, Taro Watanabe
| Challenge: | Large Language Models (LLMs) have been evaluated mostly on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content. |
| Approach: | They evaluate 26 Large Language Models using a multiple-choice question answering benchmark for Sinhala. |
| Outcome: | The new benchmarks show that Claude 3.5 sonnet and GPT-4o achieve the highest average accuracies, but overall performance remains limited. |
Matina: A Large-Scale 73B Token Persian Text Corpus (2025.naacl-long)
Copied to clipboard
Sara Bourbour Hosseinbeigi, Fatemeh Taherinezhad, Heshaam Faili, Hamed Baghbani, Fatemeh Nadi, Mostafa Amiri
| Challenge: | Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian. |
| Approach: | They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality. |
| Outcome: | The proposed model performs well on key Persian NLP tasks. |
LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content (2025.findings-naacl)
Copied to clipboard
Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Maram Hasanain, Sahinur Rahman Laskar, Naeemul Hassan, Firoj Alam
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields. |
| Approach: | They propose to develop a specialized LLM for analyzing news and social media content in a multilingual context. |
| Outcome: | The proposed model outperforms the current state-of-the-art on 23 testing sets and achieves comparable performance on 8 sets. |
INDUS: Effective and Efficient Language Models for Scientific Applications (2024.emnlp-industry)
Copied to clipboard
Bishwaranjan Bhattacharjee, Aashka Trivedi, Masayasu Muraoka, Muthukumaran Ramasubramanian, Takuma Udagawa, Iksha Gurung, Nishan Pantha, Rong Zhang, Bharath Dandala, Rahul Ramachandran, Manil Maskey, Kaylin Bugbee, Michael Little, Elizabeth Fancher, Irina Gerasimov, Armin Mehrabian, Lauren Sanders, Sylvain Costes, Sergi Blanco-Cuaresma, Kelly Lockhart, Thomas Allen, Felix Grezes, Megan Ansdell, Alberto Accomazzi, Yousef El-Kurdi, Davis Wertheimer, Birgit Pfitzmann, Cesar Berrospi Ramis, Michele Dolfi, Rafael Lima, Panagiotis Vagenas, S. Mukkavilli, Peter Staar, Sanaz Vahidinia, Ryan McGranaghan, Tsengdar Lee
| Challenge: | Large language models trained on general domain corpora showed remarkable results on natural language processing tasks. |
| Approach: | They develop a suite of large language models trained on general domain corpora that address NLP tasks and smaller versions of them created using knowledge distillation. |
| Outcome: | The proposed models outperform general-purpose and domain-specific encoders on new and existing tasks and in industrial settings. |
DEIE: Benchmarking Document-level Event Information Extraction with a Large-scale Chinese News Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing event-based datasets mainly target sentence-level tasks . current models struggle with "document" annotation, a key feature of the current model . |
| Approach: | They propose a large-scale document-level event information extraction dataset with over 56,000+ events and 242,000+ arguments. |
| Outcome: | The proposed dataset has over 56,000+ events and 242,000+ arguments. |
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment. |
| Approach: | They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge. |
| Outcome: | The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses. |
Corpora for Document-Level Neural Machine Translation (2020.lrec-1)
Copied to clipboard
| Challenge: | Document-level machine translation models translate sentences in isolation, but there are three main problems for document-level models. |
| Approach: | They propose to use document-level machine translation to capture discourse dependencies across sentences by considering a document as a whole. |
| Outcome: | The proposed method captures discourse dependencies across sentences by considering a document as a whole. |