Papers by Spencer Hong

3 papers
Watermark under Fire: A Robustness Evaluation of LLM Watermarking (2025.findings-emnlp)

Copied to clipboard

Challenge: Various watermarking methods have been proposed to identify LLM-generated texts . lack of unified evaluation platforms has left many critical questions unanswered .
Approach: They systematize existing LLM watermarkers and watermark removal attacks and develop a unified platform that integrates them.
Outcome: The proposed systematizes existing LLM watermarkers and watermark removal attacks, mapping out their design spaces.
RAFFLES: Reasoning-based Attribution of Faults for LLM Systems (2026.eacl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on simple metrics and end-to-end outcomes, but they struggle with longer contexts.
Approach: They propose an offline evaluation architecture that incorporates iterative reasoning to evaluate the quality of the candidate faults and rationales of the Judge.
Outcome: The proposed architecture outperforms baseline evaluation frameworks with two datasets to identify step-level faults in multi-agent systems and ReasonEval datasets.
DF-RAG: Query-Aware Diversity for Retrieval-Augmented Generation (2026.findings-eacl)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) is a common technique for grounding language models in domain-specific information.
Approach: They propose a new retrieval technique that incorporates diversity into the retrieval step to improve performance on reasoning-intensive QA benchmarks.
Outcome: The proposed method outperforms baselines on reasoning-intensive QA benchmarks by 4–10%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations