Papers by Jinu Lee

7 papers
SymBa: Symbolic Backward Chaining for Structured Natural Language Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Among different methods for structured reasoning, we focus on backward chaining, where the goal is recursively decomposed into subgoals by searching and applying rules.
Approach: They propose a backward chaining system that integrates a symbolic solver and an LLM to improve the performance of LLM-based reasoning.
Outcome: The proposed system improves deductive, relational, and arithmetic reasoning benchmarks compared to baselines.
Entailment-Preserving First-order Logic Representations in Natural Language Entailment (2025.acl-long)

Copied to clipboard

Challenge: First-order logic (FOL) is often used to represent logical entailment, but determining natural language (NL) enanglement using FOL remains a challenge.
Approach: They propose an Entailment-Preserving FOL representations task and a method which trains an NL-to-FOL translator by using the natural language entailment labels as verifiable rewards.
Outcome: The proposed method achieves 1.8–2.7% improvement in EPR and 17.4–20.6% increase in E PR@16 compared to baselines in three datasets.
Scaling Evaluation-Time Compute with Reasoning Models as Evaluators (2026.findings-acl)

Copied to clipboard

Challenge: Language model (LM) evaluators that generate chain-of-thought reasoning are widely used for the assessment of LM responses.
Approach: They investigate whether increasing LMs' "thinking" time through scaling test-time compute can improve an LM's evaluation capability.
Outcome: The proposed reasoning models improve evaluation performance monotonically with the number of reasoning tokens generated, mirroring trends seen in LM reasoning.
Evaluating Step-by-step Reasoning Traces: A Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation practices are inconsistent, resulting in fragmented progress across evaluator design and benchmark development.
Approach: a survey provides a comprehensive overview of step-by-step reasoning evaluation . existing evaluation practices are inconsistent, resulting in fragmented progress .
Outcome: The proposed evaluation criteria are based on four top-level categories . the results are presented in a systematic review of the literature.
Learning to Rank Generation with Pairwise Partial Rewards (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for conditional text generation suffer from large action space and delayed reward, as the reward can be computed only after an entire sequence is generated.
Approach: They propose a method that provides partial rewards for intermediate actions taken on partial sequences to prioritize actions that lead to the generation of more desirable sequences.
Outcome: The proposed method overcomes the limitations of the prevalent supervised maximum likelihood estimation approach.
LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies on legal case retrieval have limited results . limited representations and legally irrelevant matches are often used .
Approach: They propose a large-scale Korean LCR benchmark and a retrieval model that performs legal element reasoning over the query case.
Outcome: a new model outperforms baseline models on a Korean LCR benchmark . it performs state-of-the-art on 411 diverse crime types in queries over 1.2M candidate cases . previous studies have shown that the model can generalize to out-of domain cases if it is trained on in-domain data .
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics (2026.acl-long)

Copied to clipboard

Challenge: Evaluating the quality of LLM-generated reasoning traces in expert domains is essential for ensuring credibility and explainability, yet remains challenging due to the inherent complexity of such reasoning tasks.
Approach: They propose a large-scale legal reasoning dataset with an emphasis on reasoning trace evaluation that converts court judgments into hierarchical trees of opposing parties’ arguments and the court’s conclusions.
Outcome: The proposed model improves the quality of LLM-generated reasoning traces in legal domains, whereas RL improves correctness albeit with reduced coverage.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations