Papers by Hansi Hettiarachchi

4 papers
NSina: A News Corpus for Sinhala (2024.lrec-main)

Copied to clipboard

Challenge: introducing large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources.
Approach: They propose a large news corpus for Sinhala with a set of NLP tasks for the language . NSina is the largest news corpuse for Sinha, available up to date .
Outcome: The proposed model outperforms existing models in many benchmarks and outperformed previous models in high-resource languages.
Sinhala Encoder-only Language Models and Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in language models (LMs) have produced excellent results in many NLP tasks, but their effectiveness is highly dependent on available pre-training resources.
Approach: They propose to collect the largest monolingual corpus for Sinhala and compile a benchmark and evaluate LMs on it.
Outcome: The proposed language models outperform the popular multilingual LMs in downstream NLP tasks.
MUSTS: MUltilingual Semantic Textual Similarity Benchmark (2025.acl-short)

Copied to clipboard

Challenge: Existing benchmarks for semantic textual similarity (STS) are limited to high-resource languages and do not include datasets annotated focusing on relatedness instead of similarity.
Approach: They propose to evaluate multilingual semantic textual similarity benchmarks which span 13 languages and annotated datasets to evaluate and compare them.
Outcome: The proposed method is the most comprehensive benchmark of multilingual STS methods.
The Causal News Corpus: Annotating Causal Relations in Event Sentences from News (2022.lrec-1)

Copied to clipboard

Challenge: Existing annotation guidelines for event causality focus on only explicit relations or clauses.
Approach: They propose an annotation schema for event causality that addresses these concerns . they annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not.
Outcome: The proposed annotation schema for event causality addresses these concerns . it performs well with 81.20% F1 score on test set and 83.46% in 5-folds cross-validation .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations