NSina: A News Corpus for Sinhala (2024.lrec-main)

Copied to clipboard

Challenge: introducing large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources.
Approach: They propose a large news corpus for Sinhala with a set of NLP tasks for the language . NSina is the largest news corpuse for Sinha, available up to date .
Outcome: The proposed model outperforms existing models in many benchmarks and outperformed previous models in high-resource languages.

Similar Papers

Sinhala Encoder-only Language Models and Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in language models (LMs) have produced excellent results in many NLP tasks, but their effectiveness is highly dependent on available pre-training resources.
Approach: They propose to collect the largest monolingual corpus for Sinhala and compile a benchmark and evaluate LMs on it.
Outcome: The proposed language models outperform the popular multilingual LMs in downstream NLP tasks.
BERTifying Sinhala - A Comprehensive Analysis of Pre-trained Language Models for Sinhala Text Classification (2022.lrec-1)

Copied to clipboard

Challenge: Large-scale monolingual pre-trained language models have shown promising results for high-resource as well as lowresource languages, especially for text classification.
Approach: They provide a set of recommendations for using pre-trained models for Sinhala text classification and introduce new annotated datasets useful for future research.
Outcome: The proposed models are far superior to existing models for Sinhala and set a strong baseline for text classification when fine-tuned.
A Systematic Approach to Derive a Refined Speech Corpus for Sinhala (2022.lrec-1)

Copied to clipboard

Challenge: Despite being large and generic, some languages such as Sinhala are left to underutilize the technology due to the lack of adequate resources.
Approach: They propose to derive a corpus from a publicly available corpus for Sinhala speech recognition using crowdsourcing and web scraping techniques.
Outcome: The proposed corpus reduces the Word-Error-Rate by 15.9%.
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been evaluated mostly on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content.
Approach: They evaluate 26 Large Language Models using a multiple-choice question answering benchmark for Sinhala.
Outcome: The new benchmarks show that Claude 3.5 sonnet and GPT-4o achieve the highest average accuracies, but overall performance remains limited.
Matina: A Large-Scale 73B Token Persian Text Corpus (2025.naacl-long)

Copied to clipboard

Challenge: Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian.
Approach: They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality.
Outcome: The proposed model performs well on key Persian NLP tasks.
LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields.
Approach: They propose to develop a specialized LLM for analyzing news and social media content in a multilingual context.
Outcome: The proposed model outperforms the current state-of-the-art on 23 testing sets and achieves comparable performance on 8 sets.
DEIE: Benchmarking Document-level Event Information Extraction with a Large-scale Chinese News Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Existing event-based datasets mainly target sentence-level tasks . current models struggle with "document" annotation, a key feature of the current model .
Approach: They propose a large-scale document-level event information extraction dataset with over 56,000+ events and 242,000+ arguments.
Outcome: The proposed dataset has over 56,000+ events and 242,000+ arguments.
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment.
Approach: They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge.
Outcome: The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses.
Corpora for Document-Level Neural Machine Translation (2020.lrec-1)

Copied to clipboard

Challenge: Document-level machine translation models translate sentences in isolation, but there are three main problems for document-level models.
Approach: They propose to use document-level machine translation to capture discourse dependencies across sentences by considering a document as a whole.
Outcome: The proposed method captures discourse dependencies across sentences by considering a document as a whole.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations