Papers by Chetan Bansal

5 papers
CARMO: Dynamic Criteria Generation for Context Aware Reward Modelling (2025.findings-acl)

Copied to clipboard

Challenge: Reward modeling in large language models is susceptible to reward hacking . flawed reward signals often lead to outputs that optimize for spurious correlates .
Approach: They propose a new approach that generates dynamic, context-relevant criteria to ground the reward model prior to producing reward scores.
Outcome: The proposed approach generates dynamic, context-relevant criteria to ground the model prior to producing reward scores.
Learning Optimal Message Representations for Agentic Communication (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches lack the intelligence necessary to understand, learn or apply optimal communication representations adaptively.
Approach: They propose to dynamically learn the optimal message representations to enhance agentic performance by using an Expanding Markov Decision Process.
Outcome: The proposed framework improves agentic performance while maintaining efficiency.
SynthAgent: Adapting Web Agents with Synthetic Supervision (2026.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on synthetic supervision but have encountered data quality issues.
Approach: They propose a fully synthetic supervision framework that aims at improving data quality via dual refinement of both tasks and trajectories.
Outcome: The proposed framework outperforms existing methods on standardized benchmarks and shows promising results on a standardized test.
Verifiable Format Control for Large Language Model Generations (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods focus on benchmarking general instruction following while overlooking how to improve specific format following ability for small LLMs.
Approach: They propose to synthesize massive datasets to improve LLMs' format following abilities by using a verifiable format following feature.
Outcome: The proposed method improves the format following ability of small LLMs with about 7B parameters.
Synergistic Weak-Strong Collaboration by Aligning Preferences (2025.acl-long)

Copied to clipboard

Challenge: Current Large Language Models excel in general reasoning yet struggle with specialized tasks requiring proprietary or domain-specific knowledge.
Approach: They propose a collaborative framework that pairs a specialized weak model with a general strong model to optimize collaboration.
Outcome: The proposed framework outperforms each model alone by leveraging complementary strengths.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations