Papers by Dongha Lim

2 papers
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations assume tool use in short contexts, offering limited insight into model behavior during realistic long-term interactions.
Approach: a benchmark is a tool to test long-term tool use in large language models . the tool includes multiple tasks execution contexts and realistic noise .
Outcome: a new benchmark tests the tool use capabilities in long-term interactions.
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have led to their adaptation as conversational agents.
Approach: They propose a new benchmark that uses 8K multi-choice questions to assess the personality of Large Language Models.
Outcome: The proposed personality test outperforms existing personality tests for LLMs in reliability and validity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations