Papers by Heuiyeen Yeen
Ko-LongRAG: A Korean Long-Context RAG Benchmark Built with a Retrieval-Free Approach (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks for long-context RAG focus primarily on English . low-resource languages lack comprehensive evaluation frameworks limiting their progress in retrieval-based tasks. |
| Approach: | Ko-LongRAG is the first Korean long-context RAG benchmark . it adopts a retrieval-free approach designed around Specialized Content Knowledge (SCK) o1 model achieves the highest performance among proprietary models, while EXAONE 3.5 leads among open-sourced models . |
| Outcome: | the benchmark is based on a Korean language model with a retrieval-free approach . o1 model achieves the highest performance among proprietary models, while EXAONE 3.5 leads among open-sourced models. |
MANTA: A Scalable Pipeline for Transmuting Massive Web Corpora into Instruction Datasets (2025.findings-emnlp)
Copied to clipboard
| Challenge: | MANTA-1M generates high-quality large-scale instruction fine-tuning datasets from web corpora . scalability and diversity of the datasets are preserved, allowing expansion into domains requiring intensive knowledge. |
| Approach: | a team of researchers introduce a pipeline that fine-tunes large-scale instruction datasets from web corpora with minimal human intervention. |
| Outcome: | MANTA generates high-quality large-scale instruction fine-tuning datasets from web corpora . leveraging high-performance LLMs, MANTE outperforms other methods in knowledge-intensive tasks . |
Towards Context-Based Violence Detection: A Korean Crime Dialogue Dataset (2024.findings-eacl)
Copied to clipboard
| Challenge: | Currently, there are three main branches of violence detection, including surveillance of potential threats in offline situation and automatic prevention of harmful media. |
| Approach: | They propose to use the Korean Crime Dialogue Dataset to classify violence that occurs in offline scenarios. |
| Outcome: | The proposed dataset shows that understanding varying relationships among interlocutors improves the performance of crime dialogue classification. |