Papers by Ethan Mendes
Granular Privacy Control for Geolocation with Vision Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Vision Language Models (VLMs) are rapidly advancing in their capability to answer information-seeking questions. |
| Approach: | They develop a benchmark to evaluate the ability of VLMs to moderate geolocation dialogues with users. |
| Outcome: | a new benchmark evaluates the ability of VLMs to moderate geolocation conversations with users. |
Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments (2023.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations of human-in-the-loop systems to combat misinformation are often set up automatically using datasets that were retrospectively constructed. |
| Approach: | They propose a human-in-the-loop evaluation framework for fact-checking novel misinformation claims and identifying social media messages that support them. |
| Outcome: | The proposed framework is based on modern NLP methods for human-in-the-loop fact-checking in the domain of COVID-19 treatments. |
GeoRC: A Benchmark for Geolocation Reasoning Chains (2026.acl-long)
Copied to clipboard
Mohit Talreja, Joshua Diao, Jim James, Radu Casapu, Tejas Santanam, Ethan Mendes, Alan Ritter, Wei Xu, James Hays
| Challenge: | Vision Language Models (VLMs) are good at recognizing the global location of a photograph but are startlingly bad at explaining which image evidence led to their location prediction. |
| Approach: | They propose a benchmark for geolocation reasoning chains based on the global location prediction task in the popular GeoGuessr game. |
| Outcome: | The proposed benchmark compares LLM-as-a-judge and VLM-As-jumble strategies against human scoring. |
ChatHF: Collecting Rich Human Feedback from Real-time Conversations (2024.emnlp-demo)
Copied to clipboard
| Challenge: | We present an interactive framework for chatbot evaluation that integrates configurable annotation within a chat interface. |
| Approach: | They propose an interactive framework for chatbot evaluation that integrates configurable annotation within a chat interface. |
| Outcome: | The proposed framework supports fine-grained error detection and human evaluation at the same time. |