Papers by Alessio Cocchieri

8 papers
Can Large Language Models Win the International Mathematical Games? (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated strong mathematical reasoning abilities, even in visual contexts.
Approach: They propose a benchmark of 2,183 high-quality mathematical problems in an open-ended format that enables a structured evaluation of LLMs’ mathematical and logical reasoning abilities.
Outcome: The new benchmark spans seven age groups and a skill-based taxonomy and enables a structured evaluation of LLMs’ mathematical and logical reasoning abilities.
OpenBioNER: Lightweight Open-Domain Biomedical Named Entity Recognition Through Entity Type Description (2025.findings-naacl)

Copied to clipboard

Challenge: Biomedical Named Entity Recognition (BioNER) is a computationally expensive and limited tool . specialized 7B NER LLMs and GPT-4o can't match textual spans with entity types .
Approach: They propose a lightweight BERT-based cross-encoder architecture that can identify any biomedical entity using only its description.
Outcome: The proposed system outperforms existing models that match textual spans with entity types rather than descriptions on biomedical benchmarks.
To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering (2024.acl-long)

Copied to clipboard

Challenge: Medical open-domain question answering requires substantial access to specialized knowledge.
Approach: They propose a framework that generates multiple-choice questions from a set of open-book parameters and a small-scale reader that can outcompete closed-book questions by 706x using fewer parameters.
Outcome: The proposed framework outperforms closed-book models on MedQA-USMLE, MedMCQA, and MMLU while using up to 706x fewer parameters.
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? (2026.acl-long)

Copied to clipboard

Challenge: Recent advances have seen large language models (LLMs) achieve remarkable performance across high-stakes specialized domains.
Approach: They propose a diagnostic framework that evaluates legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness) they uncover a sharp domain asymmetry when applied to a benchmark that encodes temporal validity and normative relationships.
Outcome: The proposed framework evaluates legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness) it shows that legal LLMs struggle to assess when retrieved citations are useful or misleading, exhibiting overconfidence in perturbed contexts and sensitivity to superficial formatting cues.
ZeroNER: Fueling Zero-Shot Named Entity Recognition via Entity Type Descriptions (2025.findings-acl)

Copied to clipboard

Challenge: Existing zero-shot learning methods rely on entity type names for generalization . current solutions require large datasets and prioritize a handful of commonly occurring types .
Approach: They propose a description-driven framework that enhances hard zero-shot NER in low-resource settings.
Outcome: The proposed framework outperforms existing models by up to 16% in the F1 score . it also surpasses baseline models that use type names alone .
LLMs (Almost) Never Abstain Under Medical Uncertainty (2026.acl-long)

Copied to clipboard

Challenge: Medical multiple-choice question answering (MCQA) benchmarks assume large language models should always commit to an answer.
Approach: They propose a benchmark to evaluate medical abstention under uncertainty . they remove the gold answer and introduce an explicit "I abstain" option . results highlight abstinence as a critical but overlooked dimension of medical decision-making evaluation .
Outcome: The new benchmark evaluates medical abstention under uncertainty.
ReMedQA: Are We Done With Medical Multiple-Choice Benchmarks? (2026.eacl-long)

Copied to clipboard

Challenge: Multiple-choice question answering (MCQA) benchmarks show near-human accuracy . but a single accuracy score is a poor proxy for competence .
Approach: They propose a medical multiple-choice question answering (MCQA) benchmark that augments three standard medical MCQA datasets with open-ended answers and systematically perturbed options.
Outcome: The proposed benchmarks show that high MCQA accuracy masks low reliability . MCQ is the dominant paradigm for assessing medical knowledge in large language models .
“What do you call a dog that is incontrovertibly true? Dogma”: Testing LLM Generalization through Humor (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown strong performance in NLP tasks like text summarization and question answering.
Approach: They propose a new humor-based question-answering benchmark to assess LLMs’ reasoning through carefully crafted puns.
Outcome: Experiments on pun comprehension, resolution, and generation reveal that most LLMs struggle with generalization, even on simple tasks, consistently underperforming the human baseline.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations