Papers by Séb Arnold

3 papers
Using Linguistic Entrainment to Evaluate Large Language Models for Use in Cognitive Behavioral Therapy (2025.findings-naacl)

Copied to clipboard

Challenge: Entrainment is a communication process that builds a strong relationship between a mental health therapist and their client.
Approach: They evaluate the linguistic entrainment of an LLM in a mental health dialog setting and compare it to trained therapists and non-expert online peer supporters.
Outcome: The proposed model outperforms humans in a cognitive behavioral therapy setting.
LOFT: Scalable and More Realistic Long-Context Evaluation (2025.findings-naacl)

Copied to clipboard

Challenge: Long-context language models (LCLMs) can be used to perform tasks traditionally reliant on external tools like retrieval systems or databases.
Approach: They propose a benchmark to evaluate LCLMs' performance on in-context retrieval and reasoning tasks using a set of tokens.
Outcome: The proposed model outperforms state-of-the-art retrieval and RAG systems on in-context retrieval tasks while still requiring prompting strategies.
Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations (2025.emnlp-main)

Copied to clipboard

Challenge: a lack of trust in graders on graduate-level physics and Olympiad-level math makes them unreliable grader.
Approach: They propose to use a grader LM to evaluate the candidate LMs.
Outcome: The proposed approach outperforms human graders on *RewardBench* and human expert grader on Olympiad-level math problems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations