Papers by Suren Gunturu

1 papers
Reflect, Rewrite, Repeat: How Simple Arithmetic Enables Advanced Reasoning in Small Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in language model reasoning require computationally intensive reinforcement learning and massive datasets.
Approach: They propose a framework that combines Direct Preference Optimization and Supervised Fine-Tuning with selective guidance from larger models and iteratively refining solutions through a "reflect, rewrite, repeat" cycle.
Outcome: The proposed framework shows significant performance improvements across arithmetic, symbolic and cognitive reasoning benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations