Papers by Alfy Samuel

4 papers
Lessons from the Field: An Adaptable Lifecycle Approach to Applied Dialogue Summarization (2026.eacl-industry)

Copied to clipboard

Challenge: Summarization of multi-party dialogues is a critical capability in industry . but generating high-quality summaries in practice is challenging . prior work has focused on static datasets and benchmarks, a condition rare in practical scenarios .
Approach: They present an agentic system to summarize multi-party interactions using static datasets.
Outcome: The proposed system can summarize multi-party interactions using a set of complex requirements.
TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs (2025.emnlp-demos)

Copied to clipboard

Challenge: Generative Large Language Models (LLMs) produce untruthful outputs, referred to as hallucinations, which are often referred as false positives.
Approach: They propose an open-source Python library with over 30 truthfulness prediction methods.
Outcome: The proposed methods span diverse trade-offs in computational cost, access level, grounding document requirements, and supervision type (self-supervised or supervised).
A Comparison of Independent and Joint Fine-tuning Strategies for Retrieval-Augmented Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Multiple fine-tuning strategies exist with different costs and benefits for RAG pipelines.
Approach: They evaluate several RAG fine-tuning strategies with different costs and benefits . embedding and generator models can be fine- tuned to increase performance .
Outcome: The proposed techniques improve quality metrics, but have different computational costs.
An Automatic Method to Estimate Correctness of RAG (2025.coling-industry)

Copied to clipboard

Challenge: Existing methods to assess the correctness of RAG models fail to capture the model’s internal state during answer generation.
Approach: They propose a method to predict the correctness of RAG models by modeling the model’s uncertainty on quantified perturbations of input.
Outcome: Extensive experiments across multiple large language models show that the proposed approach quantifies RAG robustness by aligning predictions with ground truth with a MSE 0.002 while offering flexibility for diverse qualitative metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations