Papers by Simone Balloccu
ARQA: A Benchmark for Grounded Table–Text QA in Enterprise Annual Reports (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing QA benchmarks focus on retrieval or single-modality reasoning . annual reports are a company's definitive record of performance . |
| Approach: | They propose an annual report QA benchmark that compares QAs with lookups, arithmetics, and insights. |
| Outcome: | The proposed benchmarks show strong factual retrieval but persistent weaknesses in grounded arithmetic and causal reasoning. |
Are Experts Needed? On Human Evaluation of Counselling Reflection Generation (2023.acl-long)
Copied to clipboard
| Challenge: | Language models have been used to generate reflections automatically, but human evaluation is challenging due to the cost of hiring experts. |
| Approach: | They ask laypeople and experts to annotate synthetic reflections and human reflections from actual therapists. |
| Outcome: | The proposed method shows that laypeople and experts are reliable annotators and have moderate-to-strong inter-group correlation. |
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs (2024.eacl-long)
Copied to clipboard
| Challenge: | Lack of access to model details has raised concerns about data contamination among researchers. |
| Approach: | They conduct the first systematic analysis of work using OpenAI’s GPT-3.5 and GPT-4, the most prominently used LLMs today, in the context of data contamination. |
| Outcome: | The proposed models have been exposed to 4.7M samples from 263 benchmarks during the first year after their release. |
Ask the experts: sourcing a high-quality nutrition counseling dataset through Human-AI collaboration (2024.findings-emnlp)
Copied to clipboard
Simone Balloccu, Ehud Reiter, Karen Li, Rafael Sargsyan, Vivek Kumar, Diego Reforgiato, Daniele Riboni, Ondrej Dusek
| Challenge: | Large Language Models (LLMs) are being used by end-users for various tasks, including sensitive ones such as health counseling, disregarding potential safety concerns. |
| Approach: | They use ChatGPT to crowd-source dietary struggles and work with nutrition experts to generate supportive text using ChatGPS. |
| Outcome: | The proposed model outperforms other models on dietary struggles and mental health tasks. |