Papers with Basque

20 papers
XLTime: A Cross-Lingual Knowledge Transfer Framework for Temporal Expression Extraction (2022.findings-naacl)

Copied to clipboard

Challenge: Temporal Expression Extraction (TEE) is essential for understanding time in natural language.
Approach: They propose a framework for multilingual Temporal Expression Extraction that leverages pre-trained language models to prompt cross-language knowledge transfer from English to non-English languages.
Outcome: The proposed framework outperforms the existing SOTA methods on French, Spanish, Portuguese, and Basque by large margins.
BasqueGLUE: A Natural Language Understanding Benchmark for Basque (2022.lrec-1)

Copied to clipboard

Challenge: Natural Language Understanding (NLU) benchmarks are costly to develop and language-dependent . basqueGLUE is the first benchmark for Basque, a less-resourced language .
Approach: They propose a benchmark for Basque, a less-resourced language, using existing datasets.
Outcome: The proposed benchmarks take into account a wide and diverse set of NLU tasks that require some form of language understanding beyond the detection of superficial clues.
Basque and Spanish Counter Narrative Generation: Data Creation and Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Davidson et al.: hate speech is a growing media presence, but research on generating CNs has been limited . he says a new dataset for CN generation is available for basque and spanish . this dataset is based on a multilingual encoder-decoder model .
Approach: They propose a new Basque and Spanish dataset for automatic CN generation . they use machine translation and professional post-edition to generate CNs in both languages .
Outcome: The proposed datasets show that training on post-edited data improves generation over monolingual settings . similar results in zero-shot crosslingual evaluations show multilingual data augmentation outperforms training in English and Spanish .
Linguistic Appropriateness and Pedagogic Usefulness of Reading Comprehension Questions (2020.lrec-1)

Copied to clipboard

Challenge: Existing evaluation measures for automatic generation of reading comprehension questions focus on linguistic quality only, ignoring educational value and appropriateness of questions.
Approach: They propose a new evaluation scheme where questions are structured in a hierarchical way . they also create and evaluate two new evaluation data sets for Basque and German .
Outcome: The proposed evaluation scheme can be applied, but expert annotators are needed.
Bilingual Sentiment Embeddings: Joint Projection of Sentiment Across Languages (P18-1)

Copied to clipboard

Challenge: Existing approaches to sentiment analysis in low-resource languages lack annotated corpora or do not capture sentiment information.
Approach: They propose a model that represents sentiment in a source and target language without annotated corpus.
Outcome: The proposed model outperforms state-of-the-art methods on four out of six setups and captures complementary information to machine translation.
XNLIeu: a dataset for cross-lingual NLI in Basque (2024.naacl-long)

Copied to clipboard

Challenge: XNLI is a popular benchmark used to evaluate cross-lingual Natural Language Understanding (NLU) in languages such as English, Basque and other low-resource languages.
Approach: They expand XNLI to include Basque, a low-resource language that can benefit from transfer-learning approaches.
Outcome: The proposed dataset includes Basque, a low-resource language that can benefit from transfer-learning approaches.
Not Enough Data to Pre-train Your Language Model? MT to the Rescue! (2023.findings-acl)

Copied to clipboard

Challenge: In recent years, transformer-based language models (LMs) have become the default approach for many NLP tasks.
Approach: They compare the performance of transformer-based language models with machine-translated corpora.
Outcome: The proposed model can be improved with real data, but further research is needed.
BasqBBQ: A QA Benchmark for Assessing Social Biases in LLMs for Basque, a Low-Resource Language (2025.coling-main)

Copied to clipboard

Challenge: Existing pre-trained language models can propagate social biases in under-resourced languages like Basque.
Approach: They propose a benchmark to assess biases in Basque using a multiple-choice question-answering task.
Outcome: The proposed dataset is the first to assess biases in Basque across eight domains . larger models achieve better accuracy, but ambiguous cases remain challenging .
Better, Faster, Stronger Sequence Tagging Constituent Parsers (N19-1)

Copied to clipboard

Challenge: Existing efforts to speed up constituent parsing have focused on chart-based or shift-reduce parsers.
Approach: They propose to use auxiliary losses and sentence-level fine-tuning to mitigate greedy decoding issues.
Outcome: The proposed model surpasses the performance of sequence tagging constituent parsers on the English and Chinese Penn Treebank datasets and reduces their parsing time even further.
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of Basque parliamentary transcripts is released to study political discourse in contrasting languages . a corpus containing political discourses from public institutions can be used for computational social science research .
Approach: They present a corpus from Basque parliamentary transcripts and enrich it with metadata related to relevant attributes of speakers and speeches.
Outcome: The proposed corpus is characterized by heavy Basque-Spanish code-switching . it provides interesting insights about language use of political representatives across time, parties and gender .
Konbitzul: an MWE-specific database for Spanish-Basque (L18-1)

Copied to clipboard

Challenge: Multiword Expressions (MWEs) are combinations of words which express a single meaning.
Approach: They present an online database of verb+noun MWEs in Spanish and Basque.
Outcome: The proposed database helps to identify occurrences of MWEs in multiple morphosyntactic variants and improve translation quality in rule-based MT.
Injecting structural hints: Using language models to study inductive biases in language learning (2023.findings-emnlp)

Copied to clipboard

Challenge: a recent study examines the cognitive inductive biases that make language learning possible.
Approach: They structurally bias transformer language models by pretraining on synthetic data . they then evaluate their inductive biases by fine-tuning on three different languages .
Outcome: The proposed method predisposes transformer models to three types of inductive biases . it also fine-tunes the models on three typologically-distant human languages .
Event Extraction in Basque: Typologically Motivated Cross-Lingual Transfer-Learning Analysis (2024.lrec-main)

Copied to clipboard

Challenge: Using a multilingual language model, Event Extraction tasks require humans to follow complicated guidelines and follow complicated rules.
Approach: They propose a multilingual multilingual language model that is trained in a source language and applied to a target language.
Outcome: The proposed model is based on a multilingual event extraction dataset for Basque . it shows that the shared linguistic characteristic between source and target languages does have an impact on transfer quality.
Give your Text Representation Models some Love: the Case for Basque (2020.lrec-1)

Copied to clipboard

Challenge: Word embeddings and pre-trained language models are expensive to train and are often used by small companies and research groups to build their own.
Approach: They propose to use word embeddings and pre-trained language models to build rich representations of text and improve NLP tasks.
Outcome: The proposed models perform better than publicly available versions in downstream NLP tasks for Basque.
Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on Basque (2025.naacl-long)

Copied to clipboard

Challenge: Large language models are typically optimized for resource-rich languages like English . however, the proprietary nature of these models makes them impractical for many researchers and developers.
Approach: They propose to develop large language models that can follow instructions in Basque . they focus on three key stages: pre-training, instruction tuning, and alignment with human preferences .
Outcome: The proposed models improve natural language understanding (NLU) of the foundational model by 12 points . the results show that the models can follow instructions in Basque with human preferences .
How Well Can BERT Learn the Grammar of an Agglutinative and Flexible-Order Language? The Case of Basque. (2024.lrec-main)

Copied to clipboard

Challenge: Neural Language Models (NLMs) have demonstrated effectiveness in acquiring skills related to human language use.
Approach: They hypothesize that languages with complex grammar present substantial challenges during the pre-training phase . they constructed a test set that measures grammatical knowledge of BERT models trained under various pre-training configurations using corpus size, model size, number of epochs, and lemmatization.
Outcome: The proposed model is based on a student-based minimal pairs test set with a grammatically correct and an incorrect sentence.
Do LLMs learn a true syntactic universal? (2024.emnlp-main)

Copied to clipboard

Challenge: linguistics literature has debated whether large multilingual language models learn language universals . Typological generalizations are a key battleground in such debates - e.g. van der Hulst, 2023, chapter 7).
Approach: They consider a candidate universal for language universals, the Final-over-Final Condition . they suggest that modern language models may need additional sources of bias to become truly human-like .
Outcome: The proposed model only seems to recognize the Final-over-Final Condition in German, Russian, Hungarian and Serbian .
Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque (2025.emnlp-main)

Copied to clipboard

Challenge: Instructing language models with user intent requires large instruction datasets limited to a limited set of languages.
Approach: They propose to use existing LLMs and synthetically generated instructions to train models with user intent.
Outcome: The proposed model outperforms base non-instructed models on Basque without Basque instructions.
Truth Knows No Language: Evaluating Truthfulness Beyond English (2025.acl-long)

Copied to clipboard

Challenge: a new benchmark evaluates the truthfulness of large language models (LLMs) based on imitative falsehoods.
Approach: They propose a professionally translated extension of the TruthfulQA benchmark . it evaluates truthfulness in Basque, Catalan, Galician, and Spanish .
Outcome: The proposed extension of the TruthfulQA benchmark evaluates truthfulness in Basque, Catalan, Galician, and Spanish.
Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages? (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study focused on complex, high-level tasks, but LMentry is limited to English . a multilingual evaluation of large language models is needed to address this gap, authors say .
Approach: They propose a compact benchmark that enables systematic evaluation of large language models . they propose to use tasks that are trivial for humans but remain surprisingly difficult for LLMs .
Outcome: The proposed benchmark is limited to English, leaving its insights linguistically narrow.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations