Papers by Leonardo Zilio
An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)
Copied to clipboard
| Challenge: | a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures . |
| Approach: | They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners . |
| Outcome: | The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied . |
PLOD: An Abbreviation Detection Dataset for Scientific Documents (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing datasets for abbreviation detection and extraction are limited. |
| Approach: | They propose to use a large-scale dataset for abbreviation detection and extraction that contains 160k+ segments automatically annotated with abbrevian and long forms. |
| Outcome: | The proposed dataset has an F1 score of 0.92 for abbreviations and 0.89 for detecting their corresponding long forms. |
Character-level Language Models for Abbreviation and Long-form Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | Abbreviations and long forms are textual elements that are present in scientific communication . non-recognition of abbreviation and long form can lead to a negative impact on information retrieval . |
| Approach: | They propose to train and test language models for automatically identifying abbreviations and long forms . they use existing datasets annotated with abbrevations and their associated long forms to test them . |
| Outcome: | The proposed model can detect abbreviations and long forms on biomedical data . the proposed model improves on a previously untested dataset with biomedically-annotated datasets . |
SW4ALL: a CEFR Classified and Aligned Corpus for Language Learning (L18-1)
Copied to clipboard
| Challenge: | Learning a second language requires exposition to texts, especially for the acquisition of vocabulary. |
| Approach: | They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia . |
| Outcome: | The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency. |
Automatic Identification of COVID-19-Related Conspiracy Narratives in German Telegram Channels and Chats (2024.lrec-main)
Copied to clipboard
Philipp Heinrich, Andreas Blombach, Bao Minh Doan Dang, Leonardo Zilio, Linda Havenstein, Nathan Dykes, Stephanie Evert, Fabian Schäfer
| Challenge: | Existing methods to identify and track conspiracy narratives are difficult to track and use because of their short-lived nature. |
| Approach: | They analysed 1,000 German Telegram posts tagged with 14 fine-grained conspiracy narrative labels by three independent annotators. |
| Outcome: | The proposed methods compare well with off-the-shelf methods and human performance. |
Investigating Productive and Receptive Knowledge: A Profile for Second Language Learning (C18-1)
Copied to clipboard
| Challenge: | Literature on receptive and productive vocabulary often ignores grammar in second language acquisition studies. |
| Approach: | They use two corpora to investigate divergences in grammatical structures in texts . they set a polarity to the divergence scores to indicate whether there is overuse or underuse . |
| Outcome: | The proposed system will help language learners to activate more of their passive knowledge in writing texts. |