Challenge: Using logical forms, neural networks can sometimes require orders of magnitude more data to map from natural language instructions to state transitions (actions)
Approach: They propose to map from natural language instructions to state transitions (actions) they augment a baseline learner with an initial environment-learning phase that uses observations of language-free state transition to induce a suitable latent representation of actions before processing the instruction-following training data.
Outcome: The proposed model improves performance over systems whose representations are learned from limited instructional data alone.

Similar Papers

Robust Navigation with Language Pretraining and Stochastic Sampling (D19-1)

Copied to clipboard

Challenge: Existing methods to learn visual representations and action decoding schemes are limited to previously unseen instructions and environments.
Approach: They propose a stochastic sampling scheme to reduce the gap between the expert actions in training and sampled actions in test to correct its own mistakes.
Outcome: The proposed methods achieve 6% absolute gain over the previous best results on the Room-to-Room benchmark.
Pre-trained language model representations for language generation (N19-1)

Copied to clipboard

Challenge: Pre-trained language model representations have been successful in a wide range of language understanding tasks.
Approach: They propose to use pre-trained language model representations to integrate them into sequence to sequence models and apply it to machine translation and abstractive summarization.
Outcome: The proposed model is able to perform 5.3 BLEU in machine translation and 5.3 on the full text version of CNN/DailyMail.
Learning with Latent Language (N18-1)

Copied to clipboard

Challenge: Using the space of natural language strings as a parameter space is an effective way to capture natural task structure.
Approach: They propose to use natural language as a parameter space for few-shot learning problems including classification, transduction and policy search.
Outcome: The proposed model outperforms models with a linguistic parameterization on image classification, text editing, and reinforcement learning.
Boosting Inference Efficiency: Unleashing the Power of Parameter-Shared Pre-trained Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Parameter-shared pre-trained language models (PLMs) have emerged as a successful approach in resource-constrained environments.
Approach: They propose a method to enhance the inference efficiency of parameter-shared PLMs by pre-training models that can achieve even greater acceleration.
Outcome: The proposed method improves inference efficiency on autoregressive and autoencoding models.
Efficient Vision-Language pre-training via domain-specific learning for human activities (2024.emnlp-main)

Copied to clipboard

Challenge: Current vision-language models owe their success to large-scale pretraining on web-collected data.
Approach: They propose a domain-aligned pretraining strategy that aligns the downstream tasks to the downstream domain without additional data collection.
Outcome: The proposed method outperforms existing models on large-scale vision-language training datasets while preserving generalist knowledge.
Pretrain-KGE: Learning Knowledge Representation from Pretrained Language Models (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing knowledge graph embedding models suffer from limited knowledge representation due to sparse and noisy dataset annotations.
Approach: They propose to use pretrained language models to enhance knowledge representation by leveraging world knowledge from pretrained models.
Outcome: Extensive experiments show that the proposed framework can improve results over existing models.
Eliciting Instruction-tuned Code Language Models’ Capabilities to Utilize Auxiliary Function for Code Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Using auxiliary functions to implement functions is important for instruction-tuned models because it reduces the implementation difficulty of a target function compared to implementing them from scratch.
Approach: They propose several ways to provide auxiliary functions to the models by adding them to the query or providing a response prefix to incorporate the ability to utilize auxiliary function with the instruction following capability.
Outcome: The proposed models outperform the recent powerful language models, gpt-4o, in the code generation task.
Pretraining with Artificial Language: Studying Transferable Knowledge in Language Models (2022.acl-long)

Copied to clipboard

Challenge: Existing studies show that pretraining with an artificial language with nesting dependency structure provides some knowledge transferable to natural language.
Approach: They propose to pretrain artificial languages with structural properties that mimic natural language and then test their performance on downstream tasks.
Outcome: The proposed language models show strong performance across languages and languages.
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning (2026.eacl-long)

Copied to clipboard

Challenge: Curriculum learning has improved efficiency across machine learning domains, but remains underexplored for language model pretraining.
Approach: They present a systematic investigation of curriculum learning in LLM pretraining . they use vanilla curriculum learning, pacing-based sampling, and interleaved curricula .
Outcome: The proposed framework accelerates convergence in early and mid-training phases, reducing training steps by 18-45% to reach baseline performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations