Papers by Shamik Bose

2 papers
MAPS: A Multilingual Benchmark for Agent Performance and Security (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks do not provide a comprehensive, multi-domain, security-aware evaluation of multilingual agentic AI systems.
Approach: They propose a multilingual benchmark suite to evaluate agentic AI systems across languages and tasks.
Outcome: The proposed framework evaluates agentic AI systems across languages and tasks.
CAIR: Counterfactual-based Agent Influence Ranker for Agentic AI Workflows (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to assess the influence of each agent on the AAW’s output perform only static structural analysis, which is unsuitable for inference time execution.
Approach: They propose to use an LLM-based agent influence Ranker to assess the influence level of each agent on the AAW's output and determine which agents are the most influential.
Outcome: The proposed method outperforms baseline methods and produces consistent rankings and relevancy of downstream tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations