Papers by Jacopo Amidei
Rethinking the Agreement in Human Evaluation Tasks (C18-1)
Copied to clipboard
| Challenge: | In natural language processing, IAA is often viewed as a means of assessing the quality of data on a task, in particular, the reliability. |
| Approach: | They propose a new approach to use agreement metrics in natural language generation evaluation tasks to reduce subjective bias. |
| Outcome: | The proposed approach is based on the inter-annotator agreement (IAA) of natural language generation tasks. |
Identifying Annotator Bias: A new IRT-based method for bias identification (2020.coling-main)
Copied to clipboard
| Challenge: | An important factor that can affect the Inter Annotator Agreement (IAA) is the presence of annotator bias. |
| Approach: | They propose a new interpretation and application of the Item Response Theory to detect annotators' bias and characterise annotation disagreement. |
| Outcome: | The proposed method can be used to spot outliers, improve annotation guidelines and provide a better picture of the annotation reliability. |
Coherence of Argumentative Dialogue Snippets: A New Method for Large Scale Evaluation with an Application to Inference Anchoring Theory (2025.findings-emnlp)
Copied to clipboard
| Challenge: | illocutionary acts and propositional relations impact dialogue coherence, whereas propositional acts do not. |
| Approach: | They propose a method for testing the components of theories of dialogue coherence through utterance substitution and apply it to Inference Anchoring Theory (IAT) |
| Outcome: | The proposed method is applied to 933 dialogue snippets and 87 annotators. |
Exploring the Impact of Language Switching on Personality Traits in LLMs (2025.coling-main)
Copied to clipboard
| Challenge: | Using three personality tests, we examine the extent to which LLMs align with humans when personality shifts are associated with language changes. |
| Approach: | They propose to use the Eysenck Personality Questionnaire-Revised to examine whether LLMs align with humans when personality shifts are associated with language changes. |
| Outcome: | The results show that language-switching affects personality traits in multilingual individuals, and that it is not translation-related. |
Opening up Minds with Argumentative Dialogues (2022.findings-emnlp)
Copied to clipboard
Youmna Farag, Charlotte Brand, Jacopo Amidei, Paul Piwek, Tom Stafford, Svetlana Stoyanchev, Andreas Vlachos
| Challenge: | Recent research on argumentative dialogues has focused on persuading people to take some action, changing their stance on the topic of discussion, or winning debates. |
| Approach: | They present a dataset of 183 argumentative dialogues about veganism, Brexit and COVID-19 vaccination. |
| Outcome: | The proposed model is significantly better on other dialogue properties such as engagement and clarity. |
Similarity or deeper understanding? Analyzing the TED-Q dataset of evoked questions (2020.coling-main)
Copied to clipboard
| Challenge: | TED-Q datasets are annotated with the questions they implicitly evoke, based on a dataset of TED talks . we test whether relation between a discourse and questions it evokes is one of similarity or association . |
| Approach: | They construct a binary classification task from TED-Q and fit a BERT-based classifier alongside models based on different notions of similarity. |
| Outcome: | The proposed classifier outperforms similarity-based models in the TED-Q dataset. |