Papers by Luca Ragazzi

9 papers
Can Large Language Models Win the International Mathematical Games? (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated strong mathematical reasoning abilities, even in visual contexts.
Approach: They propose a benchmark of 2,183 high-quality mathematical problems in an open-ended format that enables a structured evaluation of LLMs’ mathematical and logical reasoning abilities.
Outcome: The new benchmark spans seven age groups and a skill-based taxonomy and enables a structured evaluation of LLMs’ mathematical and logical reasoning abilities.
Discriminative Marginalized Probabilistic Neural Method for Multi-Document Summarization of Medical Literature (2022.acl-long)

Copied to clipboard

Challenge: Existing solutions for multi-document summarization ignore potential summary-relevant contents, causing problems in the medical domain.
Approach: They propose a discriminative marginalized probabilistic method to generate a multi-document summary from a cluster of topic-related medical documents using token probability marginalization.
Outcome: The proposed method outperforms the current state-of-the-art on a biomedical dataset for multi-document summarization.
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? (2026.acl-long)

Copied to clipboard

Challenge: Recent advances have seen large language models (LLMs) achieve remarkable performance across high-stakes specialized domains.
Approach: They propose a diagnostic framework that evaluates legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness) they uncover a sharp domain asymmetry when applied to a benchmark that encodes temporal validity and normative relationships.
Outcome: The proposed framework evaluates legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness) it shows that legal LLMs struggle to assess when retrieved citations are useful or misleading, exhibiting overconfidence in perturbed contexts and sensitivity to superficial formatting cues.
What Are You Token About? Differentiable Perturbed Top-k Token Selection for Scientific Document Summarization (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets suffer from a deficiency in source heterogeneity, hindering effective model training and generalizability.
Approach: They propose a dataset that includes technical and lay summaries from multiple journals . they propose 'prinepert' transformer-based model that prunes irrelevant tokens in end-to-end learning.
Outcome: The proposed model achieves a 2x speed-up compared to a state-of-the-art linear transformer, remaining comparable in effectiveness.
LLMs (Almost) Never Abstain Under Medical Uncertainty (2026.acl-long)

Copied to clipboard

Challenge: Medical multiple-choice question answering (MCQA) benchmarks assume large language models should always commit to an answer.
Approach: They propose a benchmark to evaluate medical abstention under uncertainty . they remove the gold answer and introduce an explicit "I abstain" option . results highlight abstinence as a critical but overlooked dimension of medical decision-making evaluation .
Outcome: The new benchmark evaluates medical abstention under uncertainty.
ReMedQA: Are We Done With Medical Multiple-Choice Benchmarks? (2026.eacl-long)

Copied to clipboard

Challenge: Multiple-choice question answering (MCQA) benchmarks show near-human accuracy . but a single accuracy score is a poor proxy for competence .
Approach: They propose a medical multiple-choice question answering (MCQA) benchmark that augments three standard medical MCQA datasets with open-ended answers and systematically perturbed options.
Outcome: The proposed benchmarks show that high MCQA accuracy masks low reliability . MCQ is the dominant paradigm for assessing medical knowledge in large language models .
“What do you call a dog that is incontrovertibly true? Dogma”: Testing LLM Generalization through Humor (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown strong performance in NLP tasks like text summarization and question answering.
Approach: They propose a new humor-based question-answering benchmark to assess LLMs’ reasoning through carefully crafted puns.
Outcome: Experiments on pun comprehension, resolution, and generation reveal that most LLMs struggle with generalization, even on simple tasks, consistently underperforming the human baseline.
MemeWeaver: Inter-Meme Graph Reasoning for Sexism and Misogyny Detection (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods to detect hate speech on social media are limited by heuristic graph construction, shallow modality fusion, and instance-level reasoning.
Approach: They propose a multimodal framework for detecting sexism and misogyny using a graph reasoning mechanism that can be used to train multiple visual-textual fusion strategies.
Outcome: The proposed framework outperforms state-of-the-art methods on MAMI and EXIST benchmarks while achieving faster training convergence.
Unknown Claims: Generation of Fact-Checking Training Examples from Unstructured and Structured Data (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for fact-checking are labor-intensive and time-consuming.
Approach: They propose a framework that generates training instances for FC systems automatically using textual and tabular content.
Outcome: The proposed framework generates training instances for FC systems using textual and tabular content.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations