Papers by Jan Trienes

5 papers
InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification (2024.acl-long)

Copied to clipboard

Challenge: Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness.
Approach: They propose a framework to characterize and recover simplification-induced information loss in form of question-and-answer (QA) pairs.
Outcome: The proposed framework characterizes and recovers simplification-induced information loss in form of question-and-answer (QA) pairs.
Behavioral Analysis of Information Salience in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models excel at text summarization, but the exact notion of salience remains unclear.
Approach: They propose a framework to derive and investigate information salience in Large Language Models (LLMs) using length-controlled summarization as a behavioral probe into the content selection process.
Outcome: The proposed framework derives a proxy for how models prioritize information in large language models.
Marcel: A Lightweight and Open-Source Conversational Agent for University Student Support (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing systems that provide contextually relevant information are difficult to deploy in a university setting . a number of universities are developing or using chatbots to support prospective students .
Approach: They propose a conversational agent called Marcel that uses retrieval-augmented generation to provide contextually relevant information.
Outcome: The proposed system is designed to provide fast and personalized responses while reducing workload.
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence (2024.acl-long)

Copied to clipboard

Challenge: FactPICO is a factuality benchmark for plain language summarization of medical texts describing randomized controlled trials . existing metrics for factual summarizing medical evidence are poorly correlated with expert judgments on the instance level.
Approach: They propose a factuality benchmark for plain language summarization of medical texts . they assess factuality of critical elements of RCTs in those summaries .
Outcome: The proposed benchmark assesses the factuality of medical summaries using LLMs . the summary summators are based on 345 plain language summaires with fine-grained evaluation .
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained language models can struggle in specialized domains such as medicine . existing generalpurpose pre-tried models can be used and refined through further pre-training on domainspecific unlabeled data.
Approach: They pre-trained German medical language models on 2.4B tokens from translated public data and 3B token of German clinical data.
Outcome: The proposed models outperform clinical models on various downstream tasks in germany . the authors show that continuous pre-training can match or exceed clinical models trained from scratch .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations