Papers by Jinkyung Jo
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)
Copied to clipboard
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo
| Challenge: | a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment. |
| Approach: | They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation . |
| Outcome: | The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks. |
An Integrated Search System for Korea Weather Data (2023.emnlp-industry)
Copied to clipboard
| Challenge: | a weather search system is used to retrieve weather data from a massive weather database . a lack of navigation and time-consuming navigation hinders accurate weather forecasting . |
| Approach: | They propose a weather search system that allows users to retrieve weather data from a massive weather database with simple queries. |
| Outcome: | The proposed system achieves an average MRR and Recall of 0.82 on 4 million data points . it is based on a weather database at the Korea Meteorological Administration . |