Challenge: Existing LLM benchmarks focus on semantic complexity or quantitative competition, but rarely both simultaneously under economic scarcity.
Approach: They propose a benchmark that evaluates the capabilities of large language models (LLMs) in economically-relevant tasks through economic and trade competition.
Outcome: The proposed model evaluates the capabilities of large language models in economically-relevant tasks through economic and trade competition.

Similar Papers

Evaluating Large Language Models with Enterprise Benchmarks (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmarks lack domain-specific datasets for evaluating large language models . existing benchmarks often lack domain specific datasets, which can be difficult to convert to standardized metrics or regulatory issues.
Approach: They propose to use 25 publicly available domain-specific English benchmarks from diverse domains . they propose to combine a wide range of natural language processing tasks for holistic evaluation .
Outcome: The proposed framework includes 25 publicly available domain-specific English benchmarks from diverse enterprise domains like financial services, legal, climate, cyber security, and 2 public Japanese finance benchmarks.
LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly important for their intelligence evaluation.
Approach: They propose a game theory-based evaluation platform that measures LLMs’ decision-making strategies and social behaviors in classic game-theoretic settings.
Outcome: The proposed system cross-evaluates 15 leading LLMs using leaderboard rankings and scoring mechanisms.
MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have advanced natural language processing, demonstrating exceptional reasoning, tool usage, and memory capabilities.
Approach: They propose a competition-based benchmark framework specifically designed to assess LLMs within multi-agent environments.
Outcome: The proposed framework enhances the LLMs’ abilities in navigating complex social and cognitive dimensions by over threefold between the strongest and weakest LLM models.
MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition.
Approach: They propose a benchmark to evaluate LLM-based multi-agent systems across diverse, interactive scenarios.
Outcome: The proposed framework measures task completion and quality of collaboration and competition using novel, milestone-based key performance indicators.
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management (2026.findings-acl)

Copied to clipboard

Challenge: Large language models have shown promise in complex reasoning and tool-based decision making, but supply chain workflows require reliable long-horizon, multi-step orchestration grounded in domain-specific procedures.
Approach: They propose a unified real-world benchmark that assesses supply chain domain knowledge and long-horizon tool-based orchestration grounded in standard operating procedures.
Outcome: The proposed framework achieves the strongest and consistent tool-calling performance in real-world operational settings.
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient (2026.acl-long)

Copied to clipboard

Challenge: Using generic and efficient benchmark generators, human annotators are limited by inefficiency . current benchmark generator methods rely on seed signals, leading to long cycles and high costs .
Approach: They propose a framework to evaluate LLMs as generic benchmark generators and integrate them as BenchMaker.
Outcome: The proposed framework achieves comparable performance to human-annotated benchmarks on most metrics.
Agent Trading Arena: A Study on Numerical Understanding in LLM-Based Agents (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to large language models are limited to historical backtesting and static data.
Approach: a new large-language model is developed to simulate real-time trading in a virtual stock market . the agent trading arena simulates real-world bid-ask interactions and provides real-life trading scenarios .
Outcome: The Agent Trading Arena simulates real-world market conditions and directly impacts price dynamics.
TelAgentBench: A Multi-faceted Benchmark for Evaluating LLM-based Agents in Telecommunications (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are becoming powerful agentic systems . generic benchmarks fail to assess realistic, non-English performance .
Approach: They propose to evaluate five core agentic capabilities: Reasoning, Planning, Action (tool-use), Retrieval-Augmented Generation, and Instruction Following.
Outcome: The evaluations reveal significant performance disparities between models that employ explicit reasoning and those that do not.
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for LLM-based mobile agents are insufficient to evaluate their capabilities.
Approach: They propose a benchmark to evaluate LLM-based mobile agents' planning capabilities . they expand UI operations by incorporating 103 APIs to accelerate task completion .
Outcome: The proposed benchmarks are based on 103 collected APIs and real user queries . the data is categorized into three distinct groups: SAST, SAMT, and MAMT .
LLM Evaluate: An Industry-Focused Evaluation Tool for Large Language Models (2025.coling-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capability to solve a wide range of tasks in recent years.
Approach: They propose to build an on-premise system for LLM evaluation to address the challenges in the evaluation of LLMs in real-world industrial settings.
Outcome: The proposed evaluation system protects customer privacy and protects data integrity in real-world industrial environments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations