Papers by Heuiyeen Yeen

3 papers
Ko-LongRAG: A Korean Long-Context RAG Benchmark Built with a Retrieval-Free Approach (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for long-context RAG focus primarily on English . low-resource languages lack comprehensive evaluation frameworks limiting their progress in retrieval-based tasks.
Approach: Ko-LongRAG is the first Korean long-context RAG benchmark . it adopts a retrieval-free approach designed around Specialized Content Knowledge (SCK) o1 model achieves the highest performance among proprietary models, while EXAONE 3.5 leads among open-sourced models .
Outcome: the benchmark is based on a Korean language model with a retrieval-free approach . o1 model achieves the highest performance among proprietary models, while EXAONE 3.5 leads among open-sourced models.
MANTA: A Scalable Pipeline for Transmuting Massive Web Corpora into Instruction Datasets (2025.findings-emnlp)

Copied to clipboard

Challenge: MANTA-1M generates high-quality large-scale instruction fine-tuning datasets from web corpora . scalability and diversity of the datasets are preserved, allowing expansion into domains requiring intensive knowledge.
Approach: a team of researchers introduce a pipeline that fine-tunes large-scale instruction datasets from web corpora with minimal human intervention.
Outcome: MANTA generates high-quality large-scale instruction fine-tuning datasets from web corpora . leveraging high-performance LLMs, MANTE outperforms other methods in knowledge-intensive tasks .
Towards Context-Based Violence Detection: A Korean Crime Dialogue Dataset (2024.findings-eacl)

Copied to clipboard

Challenge: Currently, there are three main branches of violence detection, including surveillance of potential threats in offline situation and automatic prevention of harmful media.
Approach: They propose to use the Korean Crime Dialogue Dataset to classify violence that occurs in offline scenarios.
Outcome: The proposed dataset shows that understanding varying relationships among interlocutors improves the performance of crime dialogue classification.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations