Papers by Leonardo Zilio

6 papers
An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)

Copied to clipboard

Challenge: a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures .
Approach: They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners .
Outcome: The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied .
PLOD: An Abbreviation Detection Dataset for Scientific Documents (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for abbreviation detection and extraction are limited.
Approach: They propose to use a large-scale dataset for abbreviation detection and extraction that contains 160k+ segments automatically annotated with abbrevian and long forms.
Outcome: The proposed dataset has an F1 score of 0.92 for abbreviations and 0.89 for detecting their corresponding long forms.
Character-level Language Models for Abbreviation and Long-form Detection (2024.lrec-main)

Copied to clipboard

Challenge: Abbreviations and long forms are textual elements that are present in scientific communication . non-recognition of abbreviation and long form can lead to a negative impact on information retrieval .
Approach: They propose to train and test language models for automatically identifying abbreviations and long forms . they use existing datasets annotated with abbrevations and their associated long forms to test them .
Outcome: The proposed model can detect abbreviations and long forms on biomedical data . the proposed model improves on a previously untested dataset with biomedically-annotated datasets .
SW4ALL: a CEFR Classified and Aligned Corpus for Language Learning (L18-1)

Copied to clipboard

Challenge: Learning a second language requires exposition to texts, especially for the acquisition of vocabulary.
Approach: They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia .
Outcome: The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency.
Automatic Identification of COVID-19-Related Conspiracy Narratives in German Telegram Channels and Chats (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to identify and track conspiracy narratives are difficult to track and use because of their short-lived nature.
Approach: They analysed 1,000 German Telegram posts tagged with 14 fine-grained conspiracy narrative labels by three independent annotators.
Outcome: The proposed methods compare well with off-the-shelf methods and human performance.
Investigating Productive and Receptive Knowledge: A Profile for Second Language Learning (C18-1)

Copied to clipboard

Challenge: Literature on receptive and productive vocabulary often ignores grammar in second language acquisition studies.
Approach: They use two corpora to investigate divergences in grammatical structures in texts . they set a polarity to the divergence scores to indicate whether there is overuse or underuse .
Outcome: The proposed system will help language learners to activate more of their passive knowledge in writing texts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations