Challenge: Large language models (LLMs) have shown remarkable abilities in text generation, question answering, language translation, reasoning and many other tasks.
Approach: They propose a Large language model that can play chess games by transforming a game into a textual format with the best move represented in the Forsyth-Edwards Notation.
Outcome: The proposed model achieves professional-level Elo rating of 1788 in matches against the standard Elo-rated Stockfish when permitted to sample 10 times.

Similar Papers

Can Large Language Models Win the International Mathematical Games? (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated strong mathematical reasoning abilities, even in visual contexts.
Approach: They propose a benchmark of 2,183 high-quality mathematical problems in an open-ended format that enables a structured evaluation of LLMs’ mathematical and logical reasoning abilities.
Outcome: The new benchmark spans seven age groups and a skill-based taxonomy and enables a structured evaluation of LLMs’ mathematical and logical reasoning abilities.
AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are promising foundations to build generally-capable agents . however, the community lacks a unified interactive framework that covers diverse environments for comprehensive evaluation of agents.
Approach: They propose a framework that features 7 real-world scenarios, 14 environments, and 89 tasks for unified, real-time, and concurrent agent interaction.
Outcome: The proposed framework features 7 real-world scenarios, 14 environments, and 89 tasks for unified, real-time, and concurrent agent interaction.
Data and Model Centric Approaches for Expansion of Large Language Models to New languages (2025.emnlp-tutorials)

Copied to clipboard

Challenge: Existing LLMs mainly support English alongside a handful of high resource languages . this leaves a major gap for most low-resource languages despite increasing pace of research .
Approach: This tutorial examines approaches to expand the language coverage of LLMs . they look at tokenizer training, pre-training, instruction tuning, alignment, evaluation, etc.
Outcome: This tutorial examines approaches to expand the language coverage of LLMs . it provides an efficient and viable path to bring LLM technologies to low-resource languages .
Exploring Graph Learning Tasks with Pure LLMs: A Comprehensive Benchmark and Investigation (2026.findings-acl)

Copied to clipboard

Challenge: Recent studies focus on performance benchmarks without fully comparing LLMs to graph learning models.
Approach: They evaluate off-the-shelf and instruction-tuned graph learning models across a variety of scenarios.
Outcome: The proposed models outperform traditional graph learning models in few-shot settings, the authors show . their models out perform models with instruction tuning, and they show excellent generalization and robustness.
LLMs for Low Resource Languages in Multilingual, Multimodal and Dialectal Settings (2024.eacl-tutorials)

Copied to clipboard

Challenge: Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting .
Approach: They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models .
Outcome: The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects.
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have paved the way for complex tasks such as role-playing.
Approach: They propose a framework to benchmark, elicit, and enhance role-playing abilities in Large Language Models.
Outcome: The proposed framework improves role-playing abilities with 168,093 samples.
Adaptation of Large Language Models (2025.naacl-tutorial)

Copied to clipboard

Challenge: a tutorial on adaptation of large language models addresses the growing demand for models that go beyond static capabilities.
Approach: This tutorial will provide an overview of dynamic, domain-specific, and task-adaptive LLM adaptation techniques.
Outcome: This tutorial will outline dynamic, domain-specific, and task-adaptive LLM adaptation techniques.
LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating large language models use static datasets, leading to data leakage or overlooking the complexities of multi-agent interactions.
Approach: They propose a framework that evaluates the diverse capabilities of LLM agents in multi-agent dynamic environments.
Outcome: The proposed framework assesses the diverse capabilities of LLM agents in multi-agent dynamic environments.
Explore the Reasoning Capability of LLMs in the Chess Testbed (2025.naacl-short)

Copied to clipboard

Challenge: a recent study shows that large language models struggle with long-term, complex reasoning tasks.
Approach: They propose to integrate annotated strategy and tactic into large language models to improve reasoning capability.
Outcome: The proposed model performs better than GPT, Claude, and Gemini models . it integrates annotated strategy and tactic into the model .
Evaluating Large Language Models via Linguistic Profiling (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) undergo extensive evaluation against various benchmarks collected in established leaderboards to assess their performance across multiple tasks.
Approach: They propose a new evaluation methodology to test LLMs' sentence generation abilities under specific linguistic constraints.
Outcome: The proposed evaluation methodology is based on the 'linguistic profiling' approach and is not intended to be a task-oriented evaluation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations