Papers by Van-Hien Tran

3 papers
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation (2025.coling-main)

Copied to clipboard

Challenge: Pre-trained sequence-to-sequence models are typically pretrained on extensive raw text corpora and fine-tuned on task-specific data.
Approach: They introduce a pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora.
Outcome: The proposed model outperforms existing models on three generative tasks and is data-efficient and effective in enhancing performance across various natural language generation tasks.
CovRelex: A COVID-19 Retrieval System with Relation Extraction (2021.eacl-demos)

Copied to clipboard

Challenge: Existing challenges to making the system more practical include dealing with newly created and unknown data, and solving the performance gap when utilizing present data.
Approach: They propose a scientific paper retrieval system targeting entities and relations via relation extraction on COVID-19 scientific papers.
Outcome: The proposed system can be accessed via https://www.jaist.ac.jp/is/labs/nguyen-lab/systems/covrelex/.
Relation Classification Using Segment-Level Attention-based CNN and Dependency-based RNN (N19-1)

Copied to clipboard

Challenge: Recent work on relation classification has gained much success by exploiting deep neural networks.
Approach: They propose a relation classification model using Segment-level Attention-based Convolutional Neural Networks and Dependency-based Recurrent Neural networks.
Outcome: The proposed model is comparable to the state-of-the-art without external lexical features on the SemEval-2010 dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations