Papers with Greek
GR-NLP-TOOLKIT: An Open-Source NLP Toolkit for Modern Greek (2025.coling-demos)
Copied to clipboard
Lefteris Loukas, Nikolaos Smyrnioudis, Chrysa Dikonomaki, Spiros Barbakos, Anastasios Toumazatos, John Koutsikakis, Manolis Kyriakakis, Mary Georgiou, Stavros Vassos, John Pavlopoulos, Ion Androutsopoulos
| Challenge: | GR-NLP-TOOLKIT is an open-source natural language processing toolkit for modern Greek. |
| Approach: | They present GR-NLP-TOOLKIT, an open-source natural language processing toolkit for Greek. |
| Outcome: | The toolkit provides state-of-the-art performance in five core NLP tasks . it can be easily installed in Python and is accessible through a demonstration platform on HuggingFace . |
NeuronMoE: Efficient Cross-Lingual Extension via Neuron-Guided Mixture-of-Experts (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing approaches allocate experts based on layer-level similarity, yet language processing exhibits fine-grained specialization at individual neurons. |
| Approach: | They propose a method that analyzes language-specific neurons to guide expert allocation per layer based on cross-lingual neuron diversity. |
| Outcome: | The proposed method reduces the complexity of the model by 40% while matching the performance of the LayerMoE baseline. |
OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across Languages (2022.acl-long)
Copied to clipboard
| Challenge: | a new study examines the performance of pretraining for sign language recognition in low-resource settings. |
| Approach: | They propose using pose extracted through pretrained models as the standard modality of data to reduce training time and enable efficient inference. |
| Outcome: | The proposed model reduces training time and allows efficient inference in sign languages. |
Detecting de minimis Code-Switching in Historical German Books (2020.coling-main)
Copied to clipboard
| Challenge: | Code-switching has drawn scholarly attention in computational linguistics and natural language processing from many different perspectives. |
| Approach: | They propose to compare informal code-switching to its appearance in more formal registers by annotating and inspecting the German textarchives. |
| Outcome: | The proposed classifiers can help reduce errors when speech recognition is applied to a large corpus with rare embedded languages. |
Krikri: Advancing Open Large Language Models for Greek (2025.findings-emnlp)
Copied to clipboard
Dimitris Roussis, Leon Voukoutis, Georgios Paraskevopoulos, Sokratis Sofianopoulos, Prokopis Prokopidis, Vassilis Papavassileiou, Athanasios Katsamanis, Stelios Piperidis, Vassilis Katsouros
| Challenge: | Llama-Krikri-8B is a cutting-edge Large Language Model for the Greek language based on Meta's Llma 3.1-8B. |
| Approach: | They propose to use Llama-Krikri-8B to train Greek language models . it has 8 billion parameters and is capable of handling polytonic text and Ancient Greek . |
| Outcome: | The proposed model is based on Meta's Llama 3.1-8B and has 8 billion parameters and is capable of handling polytonic text and Ancient Greek. |
Nunc profana tractemus. Detecting Code-Switching in a Large Corpus of 16th Century Letters (2022.lrec-1)
Copied to clipboard
Martin Volk, Lukas Fischer, Patricia Scheurer, Bernard Silvan Schroffenegger, Raphael Schwitter, Phillip Ströbel, Benjamin Suter
| Challenge: | a corpus of 16th century letters from and to the Zurich reformer Heinrich Bullinger has been preserved . a recent study investigated code-switching in these 8600 letters . |
| Approach: | They investigate the automatic detection of code-switching in a 16th century letter exchange . they use a popular language identifier to bootstrap a word-based language classifier . |
| Outcome: | The proposed language classifier bootstraps with a popular identifier on a small training corpus of 150 sentences per language. |
Using Semantic Role Labeling to Improve Neural Machine Translation (2022.lrec-1)
Copied to clipboard
| Challenge: | despite progress in machine translation, some form of language understanding may be desirable . current systems rely on pattern recognition, but some form may be useful . |
| Approach: | They use semantic role labeling to annotate a standard parallel corpus with semantic roles . they then train a neural machine translation system using the annotated corpus and original unannotated text . |
| Outcome: | The proposed system improves BLEU scores for English, French, German, Greek and Spanish. |
Lemmatisation & Morphological Analysis of Unedited Greek: Do Simple Tasks Need Complex Solutions? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing systems for part-of-speech tagging of unedited Greek text have outperformed traditional methods for morphological analysis and lemmatisation. |
| Approach: | They propose to combine nominal features into a single label and combine the three most distinctive features of verbs into another unified label. |
| Outcome: | The proposed models outperform traditional models in lemmatisation and morphological analysis and show that multi-task learning improves performance by transferring parameters. |
Explain the Flag: Contextualizing Hate Speech Beyond Censorship (2026.findings-acl)
Copied to clipboard
Jason Liartis, Eirini Kaldeli, Lamprini Gyftokosta, Eleftherios Chelioudakis, Orfeas Menis Mastromichalakis
| Challenge: | a hybrid approach to detect and explain hate speech combines large language models with vocabularies to detect hate speech in three languages . authors: the spread of hate speech online has serious personal, social, and legal consequences . eu has launched initiatives to analyze, regulate, and counteract online hate speech, authors say . |
| Approach: | They propose a hybrid approach that combines Large Language Models with vocabularies to detect hate speech in English, French, and Greek. |
| Outcome: | The proposed approach outperforms baselines in English, French, and Greek . it uses large language models and vocabularies to detect and explain hate speech . human evaluation shows that the proposed approach is accurate and clear . |
GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek (2026.findings-acl)
Copied to clipboard
Yang Zhang, Mersin Konomi, Christos Xypolopoulos, Konstantinos Divriotis, Konstantinos Skianis, Giannis Nikolentzos, Giorgos Stamou, Guokan Shang, Michalis Vazirgiannis
| Challenge: | Existing evaluation benchmarks for large language models are limited for Greek . Existing datasets are often machine-translated from English, failing to capture Greek linguistic and cultural characteristics. |
| Approach: | They propose a native-sourced benchmark for massive multitask language understanding in Greek . they publicize 16,857 samples and reserve 4,948 samples for a private leaderboard . |
| Outcome: | The proposed model is based on 21,805 multiple-choice questions across 45 subject areas . the model is publicly released and reserved for a private leaderboard . |
A Supervised Part-Of-Speech Tagger for the Greek Language of the Social Web (2020.lrec-1)
Copied to clipboard
| Challenge: | Part-of-speech tagging is a fundamental part of NLP, but it is not widely used in unstructured text processing. |
| Approach: | They propose to use part-of-speech tags to extract information from unstructured social text in Greek and a supervised part-off-seech tagger to do so. |
| Outcome: | The proposed method performs better on unstructured microblogging text than existing methods on structured text processing. |
Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection (2026.findings-acl)
Copied to clipboard
Zhiwei Liu, Yupeng Cao, Yuechen Jiang, Mohsinul Kabir, Polydoros Giannouris, Chen Xu, Ziyang Xu, Tianlei Zhu, Md. Tariquzzaman, Triantafillos Papadopoulos, Yan Wang, Lingfei Qian, Xueqing Peng, Zhuohan Xie, Ye Yuan, Saeed Almheiri, Abdulrazzaq Alnajjar, Ming-Bin Chen, Harry Stuart, Paul Thompson, Prayag Tiwari, Alejandro Lopez-Lira, Xue Liu, Jimin Huang, Sophia Ananiadou
| Challenge: | Existing research on LLM biases has focused on direct questioning or general-purpose settings . pronounced behavioral biase despite their growing deployment in financial analysis, forecasting, and decision support. |
| Approach: | They propose a benchmark to evaluate behavioral biases of large language models in MFMD . they use a multilingual financial misinformation dataset to integrate these with misinformation claims . |
| Outcome: | The proposed benchmark evaluates behavioral biases of large language models across economic scenarios. |
Enriching Grammatical Error Correction Resources for Modern Greek (2022.lrec-1)
Copied to clipboard
| Challenge: | Davidson and Kilgarriff, 2011) have focused on the English language, but there are limited efforts to expand GEC in other languages. |
| Approach: | They develop and test a multilingual text-to-text transformer for Greek . they provide a model that can be fully-fledged for Greek with annotation corrections . |
| Outcome: | The proposed model achieves 52.63% F0.5 on part of the Greek Native Corpus, 16% below the winning system on English GEC. |
When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical Selection (2021.emnlp-main)
Copied to clipboard
| Challenge: | Using manual content to learn languages is expensive and time consuming. |
| Approach: | They propose a method for automatically identifying fine-grained lexical distinctions and extracting rules explaining them in a human- and machine-readable format. |
| Outcome: | The proposed method is able to identify fine-grained distinctions and explain them in a human- and machine-readable format. |
Probing for Reading Times (2026.acl-long)
Copied to clipboard
Eleftheria Tsipidi, Samuel Kiegeland, Francesco Ignazio Re, Tianyang Xu, Mario Giulianelli, Karolina Stanczak, Ryan Cotterell
| Challenge: | a large body of work on probing has demonstrated that language model representations encode a wealth of linguistic information, but it remains unclear whether they also capture cognitive signals about human processing. |
| Approach: | They use regularized linear regression to compare language model representations against scalar predictors. |
| Outcome: | The representations from early layers outperform surprisal in predicting early-pass measures such as first fixation and gaze duration. |
Evaluating Inflectional Complexity Crosslinguistically: a Processing Perspective (L18-1)
Copied to clipboard
| Challenge: | a cognitively motivated method for evaluating the inflectional complexity of a language is proposed . authors argue that some languages are inflectionally more complex than others . |
| Approach: | They propose a cognitively motivated method for evaluating inflectional complexity of a language . they use a recurrent self-organising neural network to learn "raw" inflected word forms . |
| Outcome: | The proposed method is independent of meta-linguistic issues and language-specific typological aspects. |
Prompting Scientific Names for Zero-Shot Species Recognition (2023.emnlp-main)
Copied to clipboard
| Challenge: | We use visionlanguage models (VLMs) to recognize images of common objects in a zero-shot fashion, but it is underexplored how to use CLIP for zero- shot species recognition of highly specialized concepts. |
| Approach: | They propose a method to translate scientific names to common English names and use them in prompts to improve their performance. |
| Outcome: | The proposed method performs poorly for species recognition with prompts that use scientific names, e.g., “a photo of Lepus Timidus” (which is a scientific name in Latin) and additionally use them in the prompts. |
Handwritten Paleographic Greek Text Recognition: A Century-Based Approach (2022.lrec-1)
Copied to clipboard
| Challenge: | achieving high accuracy HTR results for Greek manuscripts is still a major challenge . Optical character recognition software is notoriously difficult to use for handwritten text . |
| Approach: | They propose to use Greek manuscripts as a source for a new model to assess HTR accuracy. |
| Outcome: | The proposed model can be used to improve the recognition rate of Greek manuscripts. |
Sentiment Analysis of Homeric Text: The 1st Book of Iliad (2022.lrec-1)
Copied to clipboard
| Challenge: | Sentiment analysis studies focus more on online customer reviews and social media texts, but are less on literary studies. |
| Approach: | They propose to model the perceived sentiment of Iliad verses using a deep learning masked language model and a pre-trained model to estimate the sentiment of the poem. |
| Outcome: | The proposed model shows that sentiment estimators can be used as mechanical annotators, thus facilitating the distant reading of Homeric text. |
Lemmatisation of Medieval Greek: Against the Limits of Transformer’s Capabilities? (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing lemmatisation algorithms display an accuracy drop of around 30pp when tested on unedited, Byzantine Greek epigrams. |
| Approach: | They propose to use transformer-based embeddings and a dictionary look-up to lemmatise unedited, Byzantine Greek epigrams. |
| Outcome: | The proposed method outperforms existing methods and provides detailed error analysis revealing why unedited, Byzantine Greek is so challenging for lemmatisation. |
Still All Greeklish to Me: Greeklish to Greek Transliteration (2024.lrec-main)
Copied to clipboard
| Challenge: | Greeklish is a writing form that is used to avoid switching languages on multilingual keyboards . even native Greek speakers may struggle to understand Greeklished . |
| Approach: | They propose to use Greeklish to avoid switching languages on multilingual keyboards . they propose to train models on Greek datasets using the Greek alphabet . |
| Outcome: | The proposed model outperforms existing models on Greeklish data. |
Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance (2025.emnlp-main)
Copied to clipboard
Xueqing Peng, Triantafillos Papadopoulos, Efstathia Soufleri, Polydoros Giannouris, Ruoyu Xiang, Yan Wang, Lingfei Qian, Jimin Huang, Qianqian Xie, Sophia Ananiadou
| Challenge: | Greek is the dominant language of the world's merchant navy and is a key language for international trade. |
| Approach: | They propose to develop a Greek financial evaluation benchmark and a financial LLM fine-tuned on Greek-specific financial data to bridge this gap. |
| Outcome: | The proposed benchmarks surpass GPT-4 by 8.33%, GPT- 4o by 26.83%, and Deepseek-V3 by 67.74%. |
Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Terms (2025.emnlp-main)
Copied to clipboard
| Challenge: | ailsntua researchers examine whether machine translation systems exhibit gender biases that reinforce societal stereotypes. |
| Approach: | They propose a probability-based metric to evaluate gender bias by analyzing aggregated model responses. |
| Outcome: | The proposed metric evaluates whether translations in Greek and French align with or diverge from societal stereotypes. |