Papers by Huilin Lu
STORM-BORN: A Challenging Mathematical Derivations Dataset Curated via a Human-in-the-Loop Multi-Agent Framework (2025.findings-acl)
Copied to clipboard
Wenhao Liu, Zhenyi Lu, Xinyu Hu, Jerry Zhang, Dailin Li, Jiacheng Cen, Huilin Cao, Haiteng Wang, Yuhan Li, Xie Kun, Dandan Li, Pei Zhang, Chengbo Zhang, Yuxiang Ren, Xiaohong Huang, Yan Ma
| Challenge: | Existing datasets suffer from outdated and insufficient challenging content, neglecting human-like reasoning, and limited reliability due to single-LLM generation. |
| Approach: | They propose a human-in-the-loop, multi-agent data generation framework that integrates reasoning-dense filters, multiagent collaboration, and human mathematicians’ evaluations to ensure the reliability and quality of the dataset. |
| Outcome: | The proposed framework improves accuracy and quality of the 2,000-synthesized datasets by integrating reasoning-dense filters, multi-agent collaboration, and human mathematicians’ evaluations. |
Auto-Evolve: Enhancing Large Language Model’s Performance via Self-Reasoning Framework (2024.findings-emnlp)
Copied to clipboard
Krishna Aswani, Huilin Lu, Pranav Patankar, Priya Dhalwani, Xue Tan, Jayant Ganeshmohan, Simon Lacasse
| Challenge: | Recent advances in prompt engineering strategies rely on static seed reasoning modules to simulate human approach to problem-solving. |
| Approach: | They propose a framework that enables LLMs to self-create dynamic reasoning modules and downstream action plan. |
| Outcome: | The proposed framework outperforms existing prompting strategies on a BigBench-Hard dataset and improves performance by 2.8% over existing methods. |
DeepResearch Retail: Benchmarking Tool-Augmented Deep Research in the E-Commerce Domain (2026.acl-industry)
Copied to clipboard
| Challenge: | Existing DR systems are largely web-centric and do not incorporate structured, domain-specific, and personalized information accessible through internal API tools. |
| Approach: | They propose a framework grounded in real-world e-commerce data for assessing Deep Research with tools in realistic commercial settings. |
| Outcome: | The proposed framework evaluates factual faithfulness and multidimensional response quality when reasoning over heterogeneous web and internal data sources. |