Papers by Rongjunchen Zhang
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets with verifiable answers are limited in reliability, diversity, and scalability . a new approach to generate verifikatable data at scale is needed to improve models' performance . |
| Approach: | They propose a formal framework for synthesizing verifiable data at scale using a novel DSL-driven approach. |
| Outcome: | The proposed framework improves performance on a wide range of puzzles and logic benchmarks. |
CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to improve the reasoning performance of Large Language Models (LLMs) ignore annotated Chain-of-Thought (CoT) and incorporate unstable reasoning path sampling. |
| Approach: | They propose a Contrastive learning with annotated CoT-based Reinforced Fine-Tuning approach to enhance the reasoning performance of Large Language Models. |
| Outcome: | The proposed approach exploits annotated CoT and stabilizes the fine-tuning procedure by incorporating an additional unsupervised learning signal. |