Papers by Dan Tufiș

6 papers
Collection and Annotation of the Romanian Legal Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Currently, the corpus contains more than 140k documents representing the legislative body of Romania.
Approach: They present a Romanian legislative corpus which is a valuable linguistic asset for machine translation systems.
Outcome: The Romanian legislative corpus contains more than 140k documents representing the legislative body of Romania.
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)

Copied to clipboard

Challenge: The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains.
Approach: They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains .
Outcome: The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed .
The Reference Corpus of the Contemporary Romanian Language (CoRoLa) (L18-1)

Copied to clipboard

Challenge: a four-year project focused on the creation of a big corpus for contemporary Romanian language is underway . the corpus is the largest publicly available corpus of contemporary Romania .
Approach: a four-year project is focusing on the creation of a big corpus for Romanian language . the corpus is the largest publicly available corpus of the language based in the country . authors propose to use the corpora as a tool to query and listen to the results .
Outcome: a four-year project has created the largest publicly available corpus of Romanian language . the corpus is the result of a project focused on the creation of 'corola.racai.ro' the written component contains 1,257,752,812 tokens, distributed in several languages .
A Bird’s-eye View of Language Processing Projects at the Romanian Academy (L18-1)

Copied to clipboard

Challenge: a recent article outlines five projects that address contemporary Romanian language . the authors argue that a constant accumulation of human expertise is needed to develop complex projects.
Approach: a new article gives a general overview of five AI language-related projects at the Romanian Academy . they focus on the creation of a contemporary Romanian language text and speech corpus and language related applications .
Outcome: a new article gives an overview of five AI language-related projects at the Romanian Academy . the projects address contemporary Romanian language, as well as language related applications .
The MARCELL Legislative Corpus (2020.lrec-1)

Copied to clipboard

Challenge: MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification.
Approach: They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents .
Outcome: The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations