Papers by Prokopis Prokopidis

4 papers
SciPar: A Collection of Parallel Corpora from Scientific Abstracts (2022.lrec-1)

Copied to clipboard

Challenge: SciPar is a collection of parallel corpora created from openly available metadata of bachelor theses, master theses and doctoral dissertations hosted in institutional repositories, digital libraries and national archives.
Approach: They propose to harvest and process openly available metadata from repositories to extract bilingual titles and abstracts from scientific publications.
Outcome: The proposed corpora could be useful for cross-lingual plagiarism detection or adapting Machine Translation systems for translation of scientific texts and academic writing in general.
Discovering Parallel Language Resources for Training MT Engines (L18-1)

Copied to clipboard

Challenge: Web crawling is an efficient way for compiling the monolingual, parallel and/or domain-specific corpora needed for machine translation and other HLT applications.
Approach: They propose a system for compiling monolingual, parallel and/or domain-specific corpora . ILSP-FC is a web crawling system that generates bilingual lexica and terminology lists .
Outcome: The ILSP Focused Crawler is a system developed by researchers at the IL SP/Athena RIC for the acquisition of such resources.
Constructing Parallel Corpora from COVID-19 News using MediSys Metadata (2022.lrec-1)

Copied to clipboard

Challenge: Using the COVID-19 related dataset, we generated parallel corpora in 26 languages and 26 languages.
Approach: They propose to exploit COVID-19 related metadata to generate parallel corpora using the EMM/MediSys processing chain of news articles.
Outcome: The proposed corpora were based on the COVID-19 related dataset created with the Europe Media Monitor (EMM) / Medical Information System (MediSys) .
Krikri: Advancing Open Large Language Models for Greek (2025.findings-emnlp)

Copied to clipboard

Challenge: Llama-Krikri-8B is a cutting-edge Large Language Model for the Greek language based on Meta's Llma 3.1-8B.
Approach: They propose to use Llama-Krikri-8B to train Greek language models . it has 8 billion parameters and is capable of handling polytonic text and Ancient Greek .
Outcome: The proposed model is based on Meta's Llama 3.1-8B and has 8 billion parameters and is capable of handling polytonic text and Ancient Greek.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations