Challenge: Existing approaches to solving text-based games require background knowledge as the context is important.
Approach: They propose a novel agent that organizes environment states and common sense by interactive objects with a dedicated graph encoder.
Outcome: The proposed agent outperforms baselines in text-based games by 17% of scores.

Similar Papers

Fire Burns, Sword Cuts: Commonsense Inductive Bias for Exploration in Text-based Games (2022.acl-short)

Copied to clipboard

Challenge: Existing RL agents are far away from solving text-based games due to their combinatorially large action spaces that hinders efficient exploration.
Approach: They propose an exploration technique that injects external commonsense knowledge, via a pretrained language model, into the agent during training when the agent is the most uncertain about its next action.
Outcome: The proposed method exhibits improvement on the collected game scores during the training in four out of nine games from Jericho.
PLM-based World Models for Text-based Games (2022.emnlp-main)

Copied to clipboard

Challenge: a new study shows that pre-trained world models provide a strong base for world models . worldformer is a text-based game environment that can be used to learn world models in text-driven games.
Approach: They propose to use pre-trained language models to build world models in text-based game environments.
Outcome: The proposed model outperforms state-of-the-art model-free algorithms in Atari games while retaining sample efficiency.
TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments (2025.findings-acl)

Copied to clipboard

Challenge: Existing GUI agents struggle to adapt to dynamic and interconnected nature of real-world digital environments, authors show .
Approach: They propose a benchmark to evaluate the transferability of GUI agents across three key dimensions . transBench includes 15 app categories with diverse functionalities .
Outcome: The proposed benchmark shows that existing GUI agents struggle to adapt to dynamic, interconnected environments.
Efficient Text-based Reinforcement Learning by Jointly Leveraging State and Commonsense Graph Representations (2021.acl-short)

Copied to clipboard

Challenge: Text-based games (TBGs) are useful benchmarks for evaluating progress in grounded language understanding and reinforcement learning (RL).
Approach: They propose an agent that induces a graph representation of the game state and jointly grounds it with a commonsense knowledge from ConceptNet.
Outcome: The proposed agent outperforms baseline agents in the proposed game .
Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction (2021.naacl-main)

Copied to clipboard

Challenge: Scholarly work in this area uses toy worlds and synthetic linguistic data, but grounded language learning offers several practical and scientific advantages.
Approach: They propose to model teacher-learner dynamics through natural interactions occurring between users and search engines.
Outcome: The proposed model is better than non-grounded models on compositionality and zero-shot inference tasks.
Find Someone Who: Visual Commonsense Understanding in Human-Centric Grounding (2022.findings-emnlp)

Copied to clipboard

Challenge: Visual scenes often involve multiple people and humans can distinguish between them based on context descriptions about what happened before, their mental/physical states, and intentions.
Approach: They propose a task that tests human-centric commonsense grounding models' ability to distinguish individuals given context descriptions about what happened before and their mental/physical states or intentions.
Outcome: The proposed model outperforms pre-trained and non-pretrained models on 130k commonsense descriptions annotated on 67k images.
Lexi: A tool for adaptive, personalized text simplification (C18-1)

Copied to clipboard

Challenge: Existing research on text simplification has aimed to develop generic solutions . instead, we need to develop customized simplification systems for individual users .
Approach: They propose a framework for adaptive lexical simplification and introduce Lexi, a free open-source tool for personalized text simplification.
Outcome: The proposed framework is based on a free open-source tool for adaptive, personalized text simplification.
Know your audience: specializing grounded language models with listener subtraction (2023.eacl-main)

Copied to clipboard

Challenge: Effective communication requires adapting to the idiosyncrasies of each communicative context.
Approach: They propose a method for specializing grounded language models without supervision . they fine-tune an attention-based adapter between a CLIP vision encoder and a large language model .
Outcome: The proposed method allows a speaker to adapt to the idiosyncracies of the listeners without supervision.
Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities (2026.acl-long)

Copied to clipboard

Challenge: Existing agentic RAG systems rely on large language models with billions of parameters.
Approach: They propose a method to elicit agentic RAG behaviors from compact models . they propose ARC, which uses cold-start initialization and teacher guidance .
Outcome: The proposed method outperforms the larger teacher model in some cases.
From Word to World: Can Large Language Models be Implicit Text-based World Models? (2026.acl-long)

Copied to clipboard

Challenge: Agentic learning increasingly hinges on interaction, yet real-world experience is expensive, limited, and often irreversible at inference time.
Approach: They propose a framework that reframes language modeling as next-state prediction under interaction.
Outcome: The proposed framework evaluates world models in text-based environments . it shows that sufficiently trained models capture coherent environment dynamics .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations