Papers by Meghana Sunil
IREASONER: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing self-evolving frameworks mainly reward final outcomes, leaving intermediate reasoning weakly constrained despite its importance for visually grounded decision making. |
| Approach: | They propose a framework that improves an LMM’s implicit reasoning by explicitly eliciting chain-of-thought (CoT) and rewarding its internal agreement. |
| Outcome: | The proposed framework yields +2.1 points across diverse multimodal reasoning benchmarks under fully unsupervised post-training. |