Papers by Eugene Ie
Learning to Represent Image and Text with Denotation Graph (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in learning representations of visual and language information have been a problem with many applications. |
| Approach: | They propose to extract visual expressions from images aligned with linguistic expressions that describe the images to learn representations from implicit expressions. |
| Outcome: | The proposed representations lead to stronger empirical results on downstream tasks of cross-modal image retrieval, referring expression, and compositional attribute-object recognition. |
BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps (2020.acl-main)
Copied to clipboard
| Challenge: | Existing state-of-the-art VLN agents do not generalize well for long navigation tasks. |
| Approach: | They propose a VLN agent that is learned to navigate by decomposing long instructions into shorter ones and completing them sequentially. |
| Outcome: | The proposed agent can follow long instructions better than existing ones, but it does not generalize well. |
Chatbot Arena Estimate: towards a generalized performance benchmark for LLM capabilities (2025.naacl-industry)
Copied to clipboard
Lucas Spangher, Tianle Li, William F. Arnold, Nick Masiewicki, Xerxes Dotiwalla, Rama Kumar Pasumarthi, Peter Grabowski, Eugene Ie, Daniel Gruhl
| Challenge: | Existing benchmark aggregation methods, such as Elo-based systems, can be resource-intensive, public facing, and time-consuming. |
| Approach: | They propose a framework for aggregating performance across diverse benchmarks that generates a “Goodness” and a ‘Fastness” score. |
| Outcome: | The proposed framework achieves higher Pearson correlation with Chatbot Arena Elo scores than MMLU’s correlation with chatbot Arena scores, validating its reliability for real-world LLM evaluation. |
RLHF Algorithms Ranked: An Extensive Evaluation Across Diverse Tasks, Rewards, and Hyperparameters (2025.emnlp-industry)
Copied to clipboard
Lucas Spangher, Rama Kumar Pasumarthi, Nick Masiewicki, William F. Arnold, Aditi Kaushal, Dale Johnson, Peter Grabowski, Eugene Ie
| Challenge: | Proximal Policy Optimization (PPO) has fallen out of favor for Large Language Models (LLMs), but its complexity and inefficiency have spurred the investigation of simpler alternatives. |
| Approach: | They evaluate 17 RLHF algorithms on two benchmarks, OpenAI’s TL;DR Summarization and Anthropic’s Helpfulness / Harmlessness. |
| Outcome: | The proposed methods are based on OpenAI’s TL;DR Summarization and Anthropic’s Helpfulness / Harmlessness benchmarks with two different reward models and a Rules based reward model. |
Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation (P19-1)
Copied to clipboard
| Challenge: | Existing metrics for vision-and-language navigation focus on goal completion rather than the sequence of actions corresponding to the instructions. |
| Approach: | They propose to use a room-to-room dataset to measure the length of instruction followed by agents. |
| Outcome: | The proposed metric outperforms existing metrics for Room-to-Room tasks because it is direct-to goal shortest. |
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding (2020.emnlp-main)
Copied to clipboard
| Challenge: | Room-Across-Room (RxR) is a vision-and-language navigation dataset that addresses gaps in existing ones by addressing known biases in paths and eliciting more references to visible entities. |
| Approach: | They introduce a new Vision-and-Language Navigation (VLN) dataset that addresses biases in paths and elicits more references to visible entities. |
| Outcome: | The proposed model learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations. |
On the Evaluation of Vision-and-Language Navigation Instructions (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing instruction generators have not been evaluated using human wayfinders . BLEU, ROUGE, METEOR and CIDEr are ineffective for evaluating grounded navigation instructions. |
| Approach: | They propose an instruction-trajectory compatibility model that operates without reference instructions to improve wayfinding performance. |
| Outcome: | The proposed model shows the highest correlation with human wayfinding outcomes when scoring individual instructions. |
Improving Multi-Agent Debate with Sparse Communication Topology (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to multi-agent debates use a brute force algorithm, resulting in a computationally intensive process. |
| Approach: | They propose to extend the multi-agent debate framework to multi-modal reasoning and alignment labeling tasks, showcasing its broad applicability and effectiveness. |
| Outcome: | The proposed framework can achieve comparable or superior performance while significantly reducing computational costs. |