Papers by Marco Passarotti

6 papers
Representing Compounding with OntoLex. An Evaluation of Vocabularies for Word Formation Resources (2024.lrec-main)

Copied to clipboard

Challenge: OntoLex is a de facto standard for the modelling of lexical resources in the framework of Linguistic Linked Open Data.
Approach: They propose to use OntoLex to convert Linked Open Data into compounds by using the RDF model.
Outcome: The proposed model can be applied to all resources harmonized in that format, potentially allowing for the conversion into Linked Open Data of a large amount of structured data.
Odi et Amo. Creating, Evaluating and Extending Sentiment Lexicons for Latin. (2020.lrec-1)

Copied to clipboard

Challenge: a new paper aims to provide sentiment analysis tools for ancient languages . the current sentiment analysis resources only cover modern languages based on textual typologies .
Approach: They propose to use manually-curated Latin lexicons to evaluate sentiment analysis tools . they propose a gold standard and a silver standard for evaluating lexical items .
Outcome: The proposed lexicons are evaluated using a gold standard and a silver standard for sentiment analysis.
Exploring Neural Topic Modeling on a Classical Latin Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Using topic modeling, it is possible to study Latin literature through methods and tools that support distant reading.
Approach: They propose to use topic modeling to investigate thematic distribution of Latin corpus . they train, optimize and compare two neural models to evaluate which performs better .
Outcome: The proposed model is compared with two neural models with a Classical Latin corpus and shows that it is coherent and interpretable.
Modelling and Linking an Old Latin-Portuguese Dictionary to the LiLa Knowledge Base (2024.lrec-main)

Copied to clipboard

Challenge: lexical and lexicographic information of Antonio Velez's bilingual Latin-Portuguese dictionary was modelled using the Lexicon Model for Ontologies and its lexicog module.
Approach: This paper describes steps undertaken to include data from Antonio Velez’s bilingual Latin-Portuguese dictionary into the LiLa Knowledge Base of interoperable linguistic resources for Latin.
Outcome: The proposed model includes lexical and lexicographic information from the source dictionary with those of the LiLa collection of Latin lemmas.
The Index Thomisticus Treebank as Linked Data in the LiLa Knowledge Base (2022.lrec-1)

Copied to clipboard

Challenge: a series of Latin treebanks with word-by-word account of syntax and morphology of Latin texts have been published only in recent years.
Approach: They propose to publish Latin treebanks that contain morphology and syntax annotations . they propose to use principles of the Linguistic Linked Open Data community .
Outcome: The proposed approach enables interoperability between corpora and lexical resources for Latin . language learning and corpus-based research are the most obvious applications .
A New Latin Treebank for Universal Dependencies: Charters between Ancient Latin and Romance Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, Latin features the most data and the most treebanks of all the ancient languages of UD .
Approach: They introduce a Latin treebank that follows the Universal Dependencies (UD) annotation standard . they use a translation of the late Latin Charter Treebank 2 (LLCT2) into the UD style .
Outcome: The proposed treebank is based on the Universal Dependencies (UD) annotation standard.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations