Papers by Yujia Bao
WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks (2026.findings-acl)
Copied to clipboard
| Challenge: | Large-language-model (LLM) agents are competent at straightforward web tasks, but struggle with complex tasks. |
| Approach: | They propose a general framework that decomposes web tasks into three subtasks . they show that WebDART lifts end-to-end success rates by 13.7 percentage points . |
| Outcome: | Evaluated on WebChoreArena, WebDART lifts success rates by 13.7 percentage points over previous state-of-the-art agents. |
Enhancing Retrieval Systems with Inference-Time Logical Reasoning (2025.acl-short)
Copied to clipboard
| Challenge: | Existing retrieval methods rely on transforming user queries into vector representations and retrieving documents based on cosine similarity and static embeddings. |
| Approach: | They propose an inference-time logical reasoning framework that incorporates logical thinking into retrieval process. |
| Outcome: | The proposed method outperforms traditional retrieval methods on synthetic and real-world benchmarks on synthetic queries and datasets. |
Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop (2026.acl-long)
Copied to clipboard
| Challenge: | Existing synthetic training loops for large language models cause performance drops and induce emerging biases . a large amount of generated content is posted to coding platforms, social media platforms and other platforms on the internet . |
| Approach: | They propose a self-consuming retraining loop where models are trained on their own outputs . they use a control loop to isolate and analyze feedback-driven bias evolution . |
| Outcome: | The proposed model increases preference bias and decreases disparate bias. |
Deriving Machine Attention from Human Rationales (D18-1)
Copied to clipboard
| Challenge: | Attention-based models are successful when trained on large amounts of data. |
| Approach: | They propose an approach to map human-annotated rationales to high-performing attention and use this to guide models trained in low-resource scenarios. |
| Outcome: | The proposed model yields over 15% error reduction on benchmark datasets. |
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs (2025.naacl-short)
Copied to clipboard
Forrest Sheng Bao, Miaoran Li, Renyi Qu, Ge Luo, Erana Wan, Yujia Tang, Weisi Fan, Manveer Singh Tamber, Suleman Kazi, Vivek Sourabh, Mike Qi, Ruixuan Tu, Chenyu Xu, Matthew Gonzales, Ofer Mendelevitch, Amin Ahmad
| Challenge: | Existing evaluations of hallucinations in large language models suffer from a lack of diversity and recency in the LLM and LLM families considered. |
| Approach: | They propose a summarization hallucination benchmark that challenges models to disagree on hallucines . they use models to generate answers or summaries from textual input . |
| Outcome: | The proposed model combines the best of 10 modern LLMs with ground truth annotations. |