Papers by Prokopis Prokopidis
SciPar: A Collection of Parallel Corpora from Scientific Abstracts (2022.lrec-1)
Copied to clipboard
Dimitrios Roussis, Vassilis Papavassiliou, Prokopis Prokopidis, Stelios Piperidis, Vassilis Katsouros
| Challenge: | SciPar is a collection of parallel corpora created from openly available metadata of bachelor theses, master theses and doctoral dissertations hosted in institutional repositories, digital libraries and national archives. |
| Approach: | They propose to harvest and process openly available metadata from repositories to extract bilingual titles and abstracts from scientific publications. |
| Outcome: | The proposed corpora could be useful for cross-lingual plagiarism detection or adapting Machine Translation systems for translation of scientific texts and academic writing in general. |
Discovering Parallel Language Resources for Training MT Engines (L18-1)
Copied to clipboard
| Challenge: | Web crawling is an efficient way for compiling the monolingual, parallel and/or domain-specific corpora needed for machine translation and other HLT applications. |
| Approach: | They propose a system for compiling monolingual, parallel and/or domain-specific corpora . ILSP-FC is a web crawling system that generates bilingual lexica and terminology lists . |
| Outcome: | The ILSP Focused Crawler is a system developed by researchers at the IL SP/Athena RIC for the acquisition of such resources. |
Constructing Parallel Corpora from COVID-19 News using MediSys Metadata (2022.lrec-1)
Copied to clipboard
Dimitrios Roussis, Vassilis Papavassiliou, Sokratis Sofianopoulos, Prokopis Prokopidis, Stelios Piperidis
| Challenge: | Using the COVID-19 related dataset, we generated parallel corpora in 26 languages and 26 languages. |
| Approach: | They propose to exploit COVID-19 related metadata to generate parallel corpora using the EMM/MediSys processing chain of news articles. |
| Outcome: | The proposed corpora were based on the COVID-19 related dataset created with the Europe Media Monitor (EMM) / Medical Information System (MediSys) . |
Krikri: Advancing Open Large Language Models for Greek (2025.findings-emnlp)
Copied to clipboard
Dimitris Roussis, Leon Voukoutis, Georgios Paraskevopoulos, Sokratis Sofianopoulos, Prokopis Prokopidis, Vassilis Papavassileiou, Athanasios Katsamanis, Stelios Piperidis, Vassilis Katsouros
| Challenge: | Llama-Krikri-8B is a cutting-edge Large Language Model for the Greek language based on Meta's Llma 3.1-8B. |
| Approach: | They propose to use Llama-Krikri-8B to train Greek language models . it has 8 billion parameters and is capable of handling polytonic text and Ancient Greek . |
| Outcome: | The proposed model is based on Meta's Llama 3.1-8B and has 8 billion parameters and is capable of handling polytonic text and Ancient Greek. |