Papers by Yujia Bao

5 papers
WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks (2026.findings-acl)

Copied to clipboard

Challenge: Large-language-model (LLM) agents are competent at straightforward web tasks, but struggle with complex tasks.
Approach: They propose a general framework that decomposes web tasks into three subtasks . they show that WebDART lifts end-to-end success rates by 13.7 percentage points .
Outcome: Evaluated on WebChoreArena, WebDART lifts success rates by 13.7 percentage points over previous state-of-the-art agents.
Enhancing Retrieval Systems with Inference-Time Logical Reasoning (2025.acl-short)

Copied to clipboard

Challenge: Existing retrieval methods rely on transforming user queries into vector representations and retrieving documents based on cosine similarity and static embeddings.
Approach: They propose an inference-time logical reasoning framework that incorporates logical thinking into retrieval process.
Outcome: The proposed method outperforms traditional retrieval methods on synthetic and real-world benchmarks on synthetic queries and datasets.
Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop (2026.acl-long)

Copied to clipboard

Challenge: Existing synthetic training loops for large language models cause performance drops and induce emerging biases . a large amount of generated content is posted to coding platforms, social media platforms and other platforms on the internet .
Approach: They propose a self-consuming retraining loop where models are trained on their own outputs . they use a control loop to isolate and analyze feedback-driven bias evolution .
Outcome: The proposed model increases preference bias and decreases disparate bias.
Deriving Machine Attention from Human Rationales (D18-1)

Copied to clipboard

Challenge: Attention-based models are successful when trained on large amounts of data.
Approach: They propose an approach to map human-annotated rationales to high-performing attention and use this to guide models trained in low-resource scenarios.
Outcome: The proposed model yields over 15% error reduction on benchmark datasets.
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs (2025.naacl-short)

Copied to clipboard

Challenge: Existing evaluations of hallucinations in large language models suffer from a lack of diversity and recency in the LLM and LLM families considered.
Approach: They propose a summarization hallucination benchmark that challenges models to disagree on hallucines . they use models to generate answers or summaries from textual input .
Outcome: The proposed model combines the best of 10 modern LLMs with ground truth annotations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations