Papers by Simon Mille
Standard Quality Criteria Derived from Current NLP Evaluations for Guiding Evaluation Design and Grounding Comparability and AI Compliance Assessments (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations do not evaluate the same aspect of quality, resulting in unclear comparability and low repeatability. |
| Approach: | They propose to use a standard set of qualitycriterion names and definitions to establish comparability of existing evaluations. |
| Outcome: | The proposed taxonomy combines 114 quality criteria from 3 surveys of 933 evaluations in NLP and is used to establish comparability of existing evaluations and guide the design of new evaluations. |
Quantified Reproducibility Assessment of NLP Results (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods for reproducibility assessment are based on concepts and definitions from metrology. |
| Approach: | They propose a method for quantified reproducibility assessment that is based on metrology. |
| Outcome: | The proposed method produces comparable scores across multiple studies . authors find that it facilitates insights into causes of variation between studies - and conclusions can be drawn about improvements. |
A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization (2023.acl-long)
Copied to clipboard
Lining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc
| Challenge: | Using crowdsourcing, it is difficult to obtain high-quality annotations for difficult tasks. |
| Approach: | They propose a recruitment pipeline to recruit high-quality Amazon Mechanical Turk workers . they filter out subpar workers before they carry out the evaluations . |
| Outcome: | The proposed method can filter out subpar workers before they carry out evaluations and obtain high-agreement annotations with similar constraints on resources. |
The Second Multilingual Surface Realisation Shared Task (SR’19): Overview and Evaluation Results (D19-63)
Copied to clipboard
| Challenge: | EMNLP’19 Workshop on Multilingual Surface Realisation aims to stimulate the exploration of advanced neural networks for multilingual sentence generation from Universal Dependency (UD) structures. |
| Approach: | They present results from the SR'19 Shared Task, a multilingual surface realisation task organised as part of the EMNLP'19 Workshop on Multilingual Surface Realisation. |
| Outcome: | The SR'19 shared task was organised as part of the EMNLP'19 Workshop on Multilingual Surface Realisation . it consisted of two tracks with different levels of complexity . the shallow track was offered in eleven, and the deep track in three languages . |
Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP (2023.findings-acl)
Copied to clipboard
| Challenge: | reproducibility of human evaluations is rarely queried in NLP . authors estimate that just 5% of humanevaluations are repeatable . |
| Approach: | They propose to make human evaluations more repeatable and more reproducible . they estimate that just 5% of human evaluation experiments are repeatable . |
| Outcome: | The results show that human evaluations are rarely queried or formally tested in NLP . the authors estimate that just 5% of human evaluation experiments are repeatable . |
On the Role of Summary Content Units in Text Summarization Evaluation (2024.naacl-short)
Copied to clipboard
Marcel Nawrath, Agnieszka Nowak, Tristan Ratz, Danilo Walenta, Juri Opitz, Leonardo Ribeiro, João Sedoc, Daniel Deutsch, Simon Mille, Yixin Liu, Sebastian Gehrmann, Lining Zhang, Saad Mahamood, Miruna Clinciu, Khyathi Chandu, Yufang Hou
| Challenge: | a human written summary content unit (SCU) is used to judge the quality of a summary . a pyramid evaluation method is based on SCUs that decompose a reference summary into concise sentences . |
| Approach: | They propose to use automated SCUs to evaluate the quality of a candidate summary . they propose to generate SCU approximations from AMR meaning representations and large language models . |
| Outcome: | The proposed method can be fully automated, but lacks the human effort to validate it. |
Assessing the Syntactic Capabilities of Transformer-based Multilingual Language Models (2021.findings-acl)
Copied to clipboard
| Challenge: | Multilingual Transformer-based language models have been shown to be excellent learners in crosslingual transfer tasks. |
| Approach: | They evaluate the syntactic generalization capabilities of BERT and RoBERTa models on English and Spanish tests. |
| Outcome: | The proposed models perform well on English and Spanish tests, and the proposed tests are compared against models on the same language and models on two different languages. |
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)
Copied to clipboard
Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina Mcmillan-major, Anna Shvets, Ashish Upadhyay, Bernd Bohnet, Bingsheng Yao, Bryan Wilie, Chandra Bhagavatula, Chaobin You, Craig Thomson, Cristina Garbacea, Dakuo Wang, Daniel Deutsch, Deyi Xiong, Di Jin, Dimitra Gkatzia, Dragomir Radev, Elizabeth Clark, Esin Durmus, Faisal Ladhak, Filip Ginter, Genta Indra Winata, Hendrik Strobelt, Hiroaki Hayashi, Jekaterina Novikova, Jenna Kanerva, Jenny Chim, Jiawei Zhou, Jordan Clive, Joshua Maynez, João Sedoc, Juraj Juraska, Kaustubh Dhole, Khyathi Raghavi Chandu, Laura Perez Beltrachini, Leonardo F . R. Ribeiro, Lewis Tunstall, Li Zhang, Mahim Pushkarna, Mathias Creutz, Michael White, Mihir Sanjay Kale, Moussa Kamal Eddine, Nico Daheim, Nishant Subramani, Ondrej Dusek, Paul Pu Liang, Pawan Sasanka Ammanamanchi, Qi Zhu, Ratish Puduppully, Reno Kriz, Rifat Shahriyar, Ronald Cardenas, Saad Mahamood, Salomey Osei, Samuel Cahyawijaya, Sanja Štajner, Sebastien Montella, Shailza Jolly, Simon Mille, Tahmid Hasan, Tianhao Shen, Tosin Adewumi, Vikas Raunak, Vipul Raheja, Vitaly Nikolaev, Vivian Tsai, Yacine Jernite, Ying Xu, Yisi Sang, Yixin Liu, Yufang Hou
| Challenge: | Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work. |
| Approach: | They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations. |
| Outcome: | The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work. |
LLM Multi-Agent Systems for Long Triple Set Data-to-Text Generation (2026.findings-acl)
Copied to clipboard
Chinonso Cynthia Osuji, Simon Mille, Mark Andrade, Jane Adkins, Ornait O’Connell, Elaine Uí Dhonnchadha, Bláithín Heffernan, Fírinne Nic an tSaoir, Anya Belz, Thiago Castro Ferreira, Brian Davis
| Challenge: | Existing data-to-text benchmarks that do not involve content selection feature short input-output pairs designed for sentence or paragraph-level generation with reference texts spanning only a few dozen tokens. |
| Approach: | They propose a system that generates multi-paragraph outputs in English and Irish . they compare a multi-agent configuration against a single-task variant . |
| Outcome: | The proposed framework generates multi-paragraph outputs in English and Irish . human evaluation and LLM-as-a-judge score better in both languages . |
Automatic Paper Analysis and Categorisation for Systematic Reviews with Combined Reasoning-Augmented SFT and DAPO RL (2026.findings-acl)
Copied to clipboard
| Challenge: | Automating systematic reviews is expensive and time consuming, a study finds . automatic approaches are being explored but their performance has been poor . |
| Approach: | They propose to use reasoning-enhanced fine-tuning and DAPO reinforcement learning to automate systematic reviews. |
| Outcome: | The proposed methods significantly improve the performance of LLMs, the authors find . they find that reasoning-enhanced fine-tuning reduces time required for annotation by 80% . |
Back-Translation as Strategy to Tackle the Lack of Corpus in Natural Language Generation from Semantic Representations (D19-63)
Copied to clipboard
| Challenge: | Abstract Meaning Representation and Brazilian Portuguese (BP) are selected as semantic representation and language, respectively. |
| Approach: | They propose to use Brazilian Portuguese and Abstract Meaning Representation as semantic representations for NLG. |
| Outcome: | The proposed methods were evaluated on two datasets (one automatically generated and another human-generated) to compare the performance in a real context. |