Seeded self-play for language learning (D19-64)

Copied to clipboard

Challenge: Current methods for learning human language are too data inefficient to learn it in this way.
Approach: They propose to train a meta-learning agent in simulation to interact with populations of pre-trained agents, each with their own distinct communication protocol.
Outcome: The proposed algorithm minimizes the number of on-policy interactions while learning human language while minimizing the number on-political interactions.

Similar Papers

Supervised Seeded Iterated Learning for Interactive Language Learning (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work has focused on word-based conversational agents that tend to invent their language rather than leveraging natural language.
Approach: They propose two methods to counter language drift by combining S2P and Seeded Iterated Learning to minimize their weaknesses.
Outcome: The proposed methods reduce late-stage training collapses and higher negative likelihood when evaluated on human corpus.
Self-imitation Learning for Action Generation in Text-based Games (2023.eacl-main)

Copied to clipboard

Challenge: Text-based games are situated systems where the game agents observe textual descriptions, and generate textual commands to interact with the environment.
Approach: They propose a confidence-based self-imitation model to generate action candidates for the RL agent by exploiting past valuable trajectories to adapt a pre-trained language model towards a target game.
Outcome: The proposed model performs well in multiple challenging games.
Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for embodied agents to learn and perform tasks use low-level instructions, which may not reflect natural human communication.
Approach: They propose to use different types of language inputs to facilitate reinforcement learning (RL) embodied agents.
Outcome: The proposed methods show that agents trained with diverse and informative language can achieve enhanced generalization and fast adaptation to new tasks in an open world.
Multi-agent Communication meets Natural Language: Synergies between Functional and Structural Language Learning (2020.acl-main)

Copied to clipboard

Challenge: a new method for combining multi-agent communication with traditional data-driven approaches to natural language learning is proposed . we combine the two types of learning with a goal of teaching agents to communicate with humans in natural language.
Approach: They propose a method that combines traditional data-driven approaches to natural language learning with multi-agent self-play environments.
Outcome: The proposed method outperforms other methods in communicating with humans in natural language.
Language Agents: Foundations, Prospects, and Risks (2024.emnlp-tutorials)

Copied to clipboard

Challenge: Language agents are autonomous agents that can follow language instructions to perform diverse tasks in real-world or simulated environments.
Approach: They propose to provide a conceptual framework for language agents and a comprehensive discussion on key topics.
Outcome: The proposed tutorial provides a conceptual framework of language agents and comprehensive discussion on important topic areas.
Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks (2026.acl-short)

Copied to clipboard

Challenge: Language models (LMs) are pre-trained on raw text datasets to generate text sequences token-by-token.
Approach: They propose a framework that integrates Language Learning Tasks alongside standard next-token prediction to stimulate the acquisition of morphological, syntactic, and semantic knowledge.
Outcome: The proposed framework improves performance on linguistic competence benchmarks while maintaining competitive performance on reasoning tasks.
Language Models are Few-Shot Butlers (2021.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models demonstrate strong performance in most NLP tasks when fine-tuned on small task-specific datasets.
Approach: They propose a two-stage procedure to learn from a small set of demonstrations and a simple reinforcement learning algorithm to improve by interacting with an environment.
Outcome: The proposed method improves with only 1.2% of the demonstrations and a simple reinforcement learning algorithm over existing methods in the ALFWorld environment.
Reader: Model-based language-instructed reinforcement learning (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models of RL are limited and need to be re-trained for every new problem.
Approach: They propose a model-based reinforcement learning approach to tackle the environment Read To Fight Monsters, a grounded policy learning problem.
Outcome: The proposed approach performs better than existing model-free SOTA agents in the read to fight monsters environment and is more sample efficient than existing models.
Countering Language Drift via Visual Grounding (D19-1)

Copied to clipboard

Challenge: Emergent multi-agent communication protocols are different from natural language . a long-standing goal of artificial intelligence research is to develop agents that can cooperate with other agents .
Approach: They propose to use syntactic and semantic constraints to improve communication . they propose to combine these constraints with auxiliary training constraints to reduce language drift .
Outcome: a new study shows that pre-trained agents retain English syntax while learning to convey intended meaning . the proposed training constraints can be used to mitigate language drift .
Co-evolution of language and agents in referential games (2021.eacl-main)

Copied to clipboard

Challenge: Referential games allow neural agents to learn language, but they do not take into account the learning biases of the learners.
Approach: They propose to model cultural and architectural evolution in a population of agents to take into account learning biases of the language learners and let them co-evolve.
Outcome: The proposed model outperforms cultural transmission in a population of agents and takes into account learning biases of the learners.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations