Papers by Hamed Hassani

5 papers
Evaluating the Performance of Large Language Models via Debates (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are evolving and impacting various fields . current methods for evaluation are based on fixed, domain-specific questions or rely on human input, making them unscalable.
Approach: They propose a benchmarking framework based on debates between LLMs, judged by another LLM.
Outcome: The proposed framework achieves rankings that align closely with popular rankings based on human input eliminating the need for costly crowdsourcing.
Uncertainty in Language Models: Assessment through Rank-Calibration (2024.emnlp-main)

Copied to clipboard

Challenge: Language Models (LMs) have shown promising performance in natural language generation . however, it is crucial to correctly quantify their level of uncertainty in responding to inputs.
Approach: They propose a framework to quantify uncertainty and confidence for Large Language Models . they use a Rank-calibration framework to measure uncertainty and confident responses .
Outcome: The proposed framework assesses uncertainty and confidence measures for LMs.
Watermark Smoothing Attacks against Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Watermarking is a key technique for detecting AI-generated text.
Approach: They propose a method to selectively smooth watermarks by leveraging the relationship between the model’s confidence and detectability.
Outcome: The proposed method selectively smoothes watermark traces while preserving text quality.
Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities (2026.acl-long)

Copied to clipboard

Challenge: Uncertainty quantification (UQ) for large language models is a key building block for daily applications.
Approach: They propose a general formulation of agent UQ that subsumes broad classes of existing UQ setups.
Outcome: The proposed framework is based on the first general formulation of agent UQ that subsumes broad classes of existing setups.
Adaptively profiling models with task elicitation (2025.emnlp-main)

Copied to clipboard

Challenge: Language model evaluations fail to characterize consequential failure modes, forcing experts to inspect outputs and build new benchmarks.
Approach: They propose a method that automatically builds new evaluations to profile model behavior.
Outcome: The proposed method finds that language models fail in hundreds of tasks . it also finds that o3-mini is prone to hallucination when fabrications are repeated .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations