Papers by Daniela Occhipinti
Fine-tuning with HED-IT: The impact of human post-editing for dialogical language models (2024.findings-acl)
Copied to clipboard
Daniela Occhipinti, Michele Marchi, Irene Mondella, Huiyuan Lai, Felice Dell’Orletta, Malvina Nissim, Marco Guerini
| Challenge: | a recent study has focused on the quality of data generated by automatic methods for fine-tuning Language Models in languages less resourced than English. |
| Approach: | They investigate whether human intervention improves the quality of machine-generated dialogues . they use a large-scale dataset to fine-tune three different sizes of an LM . |
| Outcome: | The results show that human intervention can improve the quality of training data . larger models are less sensitive to data quality, while smaller models are more sensitive . |
When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation (2025.acl-long)
Copied to clipboard
| Challenge: | In recent years, large language models (LLMs) have proven effective in generating coherent and contextually appropriate responses. |
| Approach: | They examine the ability of a model to adapt to the interlocutor's profile by masking or disclosing information about interlucutor . |
| Outcome: | The proposed model generalises well across topics, but struggles with unfamiliar interlocutors. |
PRODIGy: a PROfile-based DIalogue Generation dataset (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing profiles-based dialogue datasets lack explicit profile representations or are difficult to collect. |
| Approach: | They propose a dataset that brings together multiple profiles for each speaker, and then integrates them together to provide a more comprehensive profile dimension set for generative language models. |
| Outcome: | The PRODIGy dataset provides a more comprehensive profile dimension set for each speaker. |