Papers by Simon Mille

11 papers
Standard Quality Criteria Derived from Current NLP Evaluations for Guiding Evaluation Design and Grounding Comparability and AI Compliance Assessments (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluations do not evaluate the same aspect of quality, resulting in unclear comparability and low repeatability.
Approach: They propose to use a standard set of qualitycriterion names and definitions to establish comparability of existing evaluations.
Outcome: The proposed taxonomy combines 114 quality criteria from 3 surveys of 933 evaluations in NLP and is used to establish comparability of existing evaluations and guide the design of new evaluations.
Quantified Reproducibility Assessment of NLP Results (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for reproducibility assessment are based on concepts and definitions from metrology.
Approach: They propose a method for quantified reproducibility assessment that is based on metrology.
Outcome: The proposed method produces comparable scores across multiple studies . authors find that it facilitates insights into causes of variation between studies - and conclusions can be drawn about improvements.
A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization (2023.acl-long)

Copied to clipboard

Challenge: Using crowdsourcing, it is difficult to obtain high-quality annotations for difficult tasks.
Approach: They propose a recruitment pipeline to recruit high-quality Amazon Mechanical Turk workers . they filter out subpar workers before they carry out the evaluations .
Outcome: The proposed method can filter out subpar workers before they carry out evaluations and obtain high-agreement annotations with similar constraints on resources.
The Second Multilingual Surface Realisation Shared Task (SR’19): Overview and Evaluation Results (D19-63)

Copied to clipboard

Challenge: EMNLP’19 Workshop on Multilingual Surface Realisation aims to stimulate the exploration of advanced neural networks for multilingual sentence generation from Universal Dependency (UD) structures.
Approach: They present results from the SR'19 Shared Task, a multilingual surface realisation task organised as part of the EMNLP'19 Workshop on Multilingual Surface Realisation.
Outcome: The SR'19 shared task was organised as part of the EMNLP'19 Workshop on Multilingual Surface Realisation . it consisted of two tracks with different levels of complexity . the shallow track was offered in eleven, and the deep track in three languages .
Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP (2023.findings-acl)

Copied to clipboard

Challenge: reproducibility of human evaluations is rarely queried in NLP . authors estimate that just 5% of humanevaluations are repeatable .
Approach: They propose to make human evaluations more repeatable and more reproducible . they estimate that just 5% of human evaluation experiments are repeatable .
Outcome: The results show that human evaluations are rarely queried or formally tested in NLP . the authors estimate that just 5% of human evaluation experiments are repeatable .
On the Role of Summary Content Units in Text Summarization Evaluation (2024.naacl-short)

Copied to clipboard

Challenge: a human written summary content unit (SCU) is used to judge the quality of a summary . a pyramid evaluation method is based on SCUs that decompose a reference summary into concise sentences .
Approach: They propose to use automated SCUs to evaluate the quality of a candidate summary . they propose to generate SCU approximations from AMR meaning representations and large language models .
Outcome: The proposed method can be fully automated, but lacks the human effort to validate it.
Assessing the Syntactic Capabilities of Transformer-based Multilingual Language Models (2021.findings-acl)

Copied to clipboard

Challenge: Multilingual Transformer-based language models have been shown to be excellent learners in crosslingual transfer tasks.
Approach: They evaluate the syntactic generalization capabilities of BERT and RoBERTa models on English and Spanish tests.
Outcome: The proposed models perform well on English and Spanish tests, and the proposed tests are compared against models on the same language and models on two different languages.
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)

Copied to clipboard

Challenge: Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work.
Approach: They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations.
Outcome: The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work.
LLM Multi-Agent Systems for Long Triple Set Data-to-Text Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing data-to-text benchmarks that do not involve content selection feature short input-output pairs designed for sentence or paragraph-level generation with reference texts spanning only a few dozen tokens.
Approach: They propose a system that generates multi-paragraph outputs in English and Irish . they compare a multi-agent configuration against a single-task variant .
Outcome: The proposed framework generates multi-paragraph outputs in English and Irish . human evaluation and LLM-as-a-judge score better in both languages .
Automatic Paper Analysis and Categorisation for Systematic Reviews with Combined Reasoning-Augmented SFT and DAPO RL (2026.findings-acl)

Copied to clipboard

Challenge: Automating systematic reviews is expensive and time consuming, a study finds . automatic approaches are being explored but their performance has been poor .
Approach: They propose to use reasoning-enhanced fine-tuning and DAPO reinforcement learning to automate systematic reviews.
Outcome: The proposed methods significantly improve the performance of LLMs, the authors find . they find that reasoning-enhanced fine-tuning reduces time required for annotation by 80% .
Back-Translation as Strategy to Tackle the Lack of Corpus in Natural Language Generation from Semantic Representations (D19-63)

Copied to clipboard

Challenge: Abstract Meaning Representation and Brazilian Portuguese (BP) are selected as semantic representation and language, respectively.
Approach: They propose to use Brazilian Portuguese and Abstract Meaning Representation as semantic representations for NLG.
Outcome: The proposed methods were evaluated on two datasets (one automatically generated and another human-generated) to compare the performance in a real context.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations