Papers by Kinjal Basu
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls (2025.emnlp-main)
Copied to clipboard
Kinjal Basu, Ibrahim Abdelaziz, Kiran Kate, Mayank Agarwal, Maxwell Crouse, Yara Rizk, Kelsey Bradford, Asim Munawar, Sadhana Kumaravel, Saurabh Goyal, Xin Wang, Luis A. Lastras, Pavan Kapanipathi
| Challenge: | Existing benchmarks and datasets for tool calling have lagged behind . nested sequencing is a common problem in LLMs, but it is not enough to evaluate them. |
| Approach: | They propose a benchmark to evaluate LLMs on nested sequences of API calls, i.e. sequences where the output of one API call is passed as input to a subsequent call. |
| Outcome: | The proposed model achieves a full sequence match accuracy of 28% and a win-rate of 60% on nested sequences of API calls. |
EXPLORER: Exploration-guided Reasoning for Textual Reinforcement Learning (2024.eacl-long)
Copied to clipboard
Kinjal Basu, Keerthiram Murugesan, Subhajit Chaudhury, Murray Campbell, Kartik Talamadupula, Tim Klinger
| Challenge: | Text-based games (TBGs) combine natural language understanding with reasoning. |
| Approach: | They propose an exploration-guided reasoning agent for textual reinforcement learning that integrates natural language with reasoning. |
| Outcome: | The proposed agent outperforms baseline agents on TWG and TWC games. |
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks (2024.emnlp-industry)
Copied to clipboard
Ibrahim Abdelaziz, Kinjal Basu, Mayank Agarwal, Sadhana Kumaravel, Matthew Stallone, Rameswar Panda, Yara Rizk, G P Shrivatsa Bhargav, Maxwell Crouse, Chulaka Gunasekara, Shajith Ikbal, Sachindra Joshi, Hima Karanam, Vineet Kumar, Asim Munawar, Sumit Neelam, Dinesh Raghu, Udit Sharma, Adriana Soria, Dheeraj Sreedhar, Praveen Venkateswaran, Merve Unuvar, David Cox, Salim Roukos, Luis Lastras, Pavan Kapanipathi
| Challenge: | Existing research explores the use of Large Language Models (LLMs) as the backbone of agentic systems. |
| Approach: | They propose a model trained using a multi-task training approach on seven fundamental tasks encompassed in function calling that has better generalizability on multiple tasks across seven evaluation benchmarks. |
| Outcome: | The proposed model outperforms more than 15 other models on out-of-domain datasets and ranks among the top on the Berkeley Function Calling Leaderboard (BFCL). |
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling (2026.eacl-long)
Copied to clipboard
Benjamin Elder, Anupama Murthi, Jungkoo Kang, Ankita Rajaram Naik, Kiran Kate, Kinjal Basu, Danish Contractor
| Challenge: | Large language models rely on external tools and APIs to perform tasks specified in natural language. |
| Approach: | They propose a benchmark that transforms SQL queries from BIRD-SQL into executable API sequences. |
| Outcome: | The proposed benchmark evaluates 10 LLMs and 4 ReACT agents with low task completion rates and 50% task completion rate. |
R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for web agents struggle with efficient navigation and action execution due to limited visibility and understanding of web structures. |
| Approach: | They propose a framework that integrates memory-enhanced navigation and reflective learning to improve web agents' performance. |
| Outcome: | The proposed framework shows significant improvements over existing methods, including 50% reduction in navigation errors and threefold increase in task completion rates. |
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs (2024.acl-long)
Copied to clipboard
Kinjal Basu, Ibrahim Abdelaziz, Subhajit Chaudhury, Soham Dan, Maxwell Crouse, Asim Munawar, Vernon Austel, Sadhana Kumaravel, Vinod Muthusamy, Pavan Kapanipathi, Luis Lastras
| Challenge: | Existing methods to train and test large language models that involve calls to tools and APIs are lacking. |
| Approach: | They propose a large corpora for training and systematic testing of tool-augmented LLMs. |
| Outcome: | The proposed datasets mimic real-world scenarios involving API-tasks and slot filling. |