Papers by Ayush Goyal
ShopperBench: A Benchmark for Personalized Shopping with Persona-Guided Simulation (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing evaluation frameworks lack mechanisms to assess Personalized shopping agents' ability to adapt their strategies to heterogeneous user preferences and decisionmaking patterns. |
| Approach: | They propose a persona-guided benchmark that augments shopping trajectories with personas . they propose persona Fidelity, Persona-Query Alignment, and Path Consistency . |
| Outcome: | The proposed benchmark captures how shopper types navigate product search and selection . it measures persona Fidelity, Persona-Query Alignment, and Path Consistency . |
Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Framework (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) serve diverse global audiences, making it critical for responsible AI deployment across cultures. |
| Approach: | They propose a framework that conceptualizes alignment along three dimensions: Cultural Group, Cultural Elements and Awareness Scope. |
| Outcome: | The proposed framework reveals critical gaps between benchmarks and real-world cultural biases . region dominates cultural group representation, social and political relations dominates coverage . majority of datasets adopt majority-focused Awareness Scope approaches . |
CaM-Gen: Causally Aware Metric-Guided Text Generation (2022.findings-acl)
Copied to clipboard
| Challenge: | Content is created for a well-defined purpose, often described by a metric or signal . external metrics and content tend to have inherent relationships and not all of them may be of consequence. |
| Approach: | They propose a mechanism to guide generative models by user-defined target metrics . authors propose generative networks guided by causally significant aspects of text . |
| Outcome: | The proposed models beat baselines in terms of the target metric control while maintaining fluency and language quality of the generated text. |