Papers by Arianna Graciotti

3 papers
GOLEMcoref: A Multilingual Coreference Dataset of Fiction (2026.acl-short)

Copied to clipboard

Challenge: Despite considerable progress, most research still focuses predominantly on English . fictional texts bring additional challenges not covered by standard benchmark datasets .
Approach: They present a multilingual coreference dataset of 827k fanfiction tokens in 7 languages . they discuss their annotation scheme and language-specific challenges .
Outcome: The proposed dataset includes full stories of diverse lengths, ranging from 500 to 17k words.
Latent vs Explicit Knowledge Representation: How ChatGPT Answers Questions about Low-Frequency Entities (2024.lrec-main)

Copied to clipboard

Challenge: In this paper, we compare two different approaches to the free-form Question Answering task.
Approach: They propose to use a new benchmark to test knowledge representations on a dynamic benchmark.
Outcome: The proposed benchmark is particularly challenging and the best model answers only on 50% of the questions.
KE-MHISTO: Towards a Multilingual Historical Knowledge Extraction Benchmark for Addressing the Long-Tail Problem (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models struggle when probed for long-tail knowledge due to the inherent sparsity of such data.
Approach: They propose a multilingual benchmark for Entity Linking and Question Answering in the domain of historical music knowledge that provides broader coverage of long-tail knowledge.
Outcome: The proposed model provides broader coverage of long-tail knowledge compared to existing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations