Papers by Rogerio Feris

5 papers
Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Vision and language models (VLMs) have demonstrated remarkable zero-shot (ZS) performance in a variety of tasks.
Approach: They propose to integrate structured annotations into visual and textual representations to improve VLMs' understanding of compositional scenes.
Outcome: The proposed method improves VLMs on multiple VL datasets with only a mild degradation in ZS capabilities.
Synthetic Pre-Training Tasks for Neural Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: toxicity and bias can be addressed by pre-training with synthetic resources . BLEU scores are used to compare methods with real-world data .
Approach: They propose several ways to generate obfuscated data from large parallel corpus and concatenating phrase pairs from small word-aligned corpus with synthetic parallel data without real human language corpora.
Outcome: The proposed methods can be used to generate obfuscated data or synthetic parallel data without real human language corpora even with high levels of oblication.
Self-Specialization: Uncovering Latent Expertise within Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have demonstrated the effectiveness of self-alignment in which a large language model is aligned to follow general instructions using instructional data generated from the model itself.
Approach: They propose to use human-written seeds to align large language models to follow general instructions to achieve cross-task generalization.
Outcome: The proposed model outperforms base models and models that are generally instruction-tuned or have been adapted to the target domain by a large margin.
Activation Reward Models for Few-Shot Model Alignment (2026.findings-acl)

Copied to clipboard

Challenge: A common approach is to use reward models that enable reinforcement-learning post-training.
Approach: They propose a method that steers LLM activations to align with few-shot preference data without finetuning.
Outcome: The proposed method surpasses zero-shot, few-shot and voting-based benchmarks on reward hacking and noise signals.
LangNav: Language as a Perceptual Representation for Navigation (2024.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to vision-and-language navigation use visual features as the perceptual representation of a visual representation of an agent's egocentric panoramic view.
Approach: They propose to use off-the-shelf vision systems to convert an agent’s egocentric panoramic view into natural language descriptions.
Outcome: The proposed approach improves on the R2R VLN benchmark by using synthetic trajectories from a prompted language model and domain transfer where a policy learned on one simulated environment (ALFRED) is transferred to another (more realistic) environment and combining both vision- and language-based representations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations