Papers by Sungjin Park

8 papers
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods for ensembling language models fail to address complex reasoning tasks.
Approach: They propose a framework for process-level ensembling of large language models using Monte Carlo tree search.
Outcome: The proposed framework outperforms both language model decoding and language model ensemble methods on five reasoning benchmarks.
Learning Slice-Aware Representations with Mixture of Attentions (2021.findings-acl)

Copied to clipboard

Challenge: Real-world machine learning systems are achieving excellent performance in terms of coarse-grained metrics like overall accuracy and F-1 score.
Approach: They extend slice-based learning (SBL) with a mixture of attentions to learn slice-aware dual attentive representations.
Outcome: The proposed approach outperforms the baseline method and the original SBL approach on monitored slices with two natural language understanding tasks.
A Scalable Framework for Learning From Implicit User Feedback to Improve Natural Language Understanding in Large-Scale Conversational AI Systems (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods to improve NLU are laborintensive and expensive.
Approach: They propose a scalable and automatic approach to improving NLU in a large-scale conversational AI system by leveraging implicit user feedback.
Outcome: The proposed framework improves NLU in a large-scale conversational AI system across 10 domains.
Do Language Models Understand Measurements? (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on numerical reasoning over text (NRoT) tests PLMs to understand numbers in contexts where numbers are an integral part of the context.
Approach: They propose a simple embedding strategy to better distinguish between numbers and units, which leads to a significant improvement in probing tasks.
Outcome: The proposed model distinguishes between numbers and units, which leads to significant improvement in probing tasks.
FactKG: Fact Verification via Reasoning on Knowledge Graphs (2023.acl-long)

Copied to clipboard

Challenge: knowledge graphs (KGs) have not been fully utilized as a knowledge source for fact verification.
Approach: They propose a dataset to enable the community to better use knowledge graphs . they propose 108k natural language claims with five types of reasoning .
Outcome: The proposed dataset consists of 108k natural language claims with five types of reasoning . authors believe the proposed method can advance reliability and practicality .
FreeTalky: Don’t Be Afraid! Conversations Made Easier by a Humanoid Robot using Persona-based Dialogue (2022.lrec-1)

Copied to clipboard

Challenge: FreeTalky is a deep learning-based foreign language learning platform for people who experience anxiety dealing with foreign languages.
Approach: They propose a deep learning-based foreign language learning platform called FreeTalky . it employs a humanoid robot NAO and various deep learning models .
Outcome: The proposed system provides personalized learning based on persona dialogue and grammar error correction, and also helps alleviate xenoglossophobia by replacing the real human in the conversation with a NAO robot, through human evaluation.
From Ambiguity to Accuracy: The Transformative Effect of Coreference Resolution on Retrieval-Augmented Generation systems (2025.acl-srw)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) is a key framework in natural language processing . however, the effectiveness of RAG is often hindered by coreferential complexity in retrieved documents .
Approach: They investigate how entity coreference affects document retrieval and generative performance in RAG-based systems.
Outcome: The proposed model improves QA performance and retrieval relevance and contextual understanding.
Open World Classification with Adaptive Negative Samples (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models with no effective open category data during training are limited by the lack of effective open categories data during the training stage.
Approach: They propose an approach to generate effective open category samples in the training stage and without requiring prior knowledge or external datasets.
Outcome: The proposed approach generates effective synthetic open category samples in the training stage and without requiring any prior knowledge or external datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations