Papers with Swedish

29 papers
Reducing Confusion in Active Learning for Part-Of-Speech Tagging (2021.tacl-1)

Copied to clipboard

Challenge: Existing algorithms for annotating parts of speech are not optimal for all languages.
Approach: They propose to use a data selection algorithm to select useful training samples to minimize annotation cost.
Outcome: The proposed strategy outperforms existing strategies on six typologically diverse languages.
Acceptability Judgements via Examining the Topology of Attention Maps (2022.findings-emnlp)

Copied to clipboard

Challenge: Acceptability judgments are a key component of generative linguistics, but their ability to judge grammatical acceptability has not been explored.
Approach: They propose to exploit the geometric properties of the attention graph to evaluate the grammatical acceptability of sentences using topological data analysis.
Outcome: The proposed approach outperforms nine statistical and Transformer LM baselines on the BLiMP benchmark and the human-level performance on the same benchmark.
Beyond the English Web: Zero-Shot Cross-Lingual and Lightweight Monolingual Classification of Registers (2021.eacl-srw)

Copied to clipboard

Challenge: Existing studies on register classification for web documents have limited results due to skewed datasets and low performance.
Approach: They propose two new register-annotated corpora for French and Swedish . they show that deep pre-trained language models perform strongly in these languages .
Outcome: The proposed models outperform existing models in English and Finnish and can match or surpass existing models.
Guiding Zero-Shot Paraphrase Generation with Fine-Grained Control Tokens (2023.starsem-1)

Copied to clipboard

Challenge: Sequence-to-sequence paraphrase generation models struggle with the generation of diverse paraphrases.
Approach: They propose a translation-based guided paraphrase generation model that learns useful features for promoting surface form variation in generated paraphrases from cross-lingual parallel data.
Outcome: The proposed model learns useful features for promoting surface form variation in generated paraphrases from cross-lingual parallel data.
Constructing Multimodal Language Learner Texts Using LARA: Experiences with Nine Languages (2020.lrec-1)

Copied to clipboard

Challenge: LARA is an open source project that aims to support easy conversion of plain texts into online versions suitable for use by language learners.
Approach: They propose to support easy conversion of plain texts into online versions suitable for use by language learners.
Outcome: The proposed platform is suitable for creating texts in multiple languages via crowdsourcing techniques that can be used for teaching a language via reading and listening.
Sentences with Gapping: Parsing and Reconstructing Elided Predicates (N18-1)

Copied to clipboard

Challenge: Sentences with gapping lack an overt predicate to indicate the relation between two or more arguments.
Approach: They propose two methods for parsing to a Universal Dependencies graph representation that explicitly encodes the elided material with additional nodes and edges.
Outcome: The proposed methods reconstruct elided material from dependency trees with high accuracy when the parser correctly predicts the existence of a gap.
An Evaluation of Neural Machine Translation Models on Historical Spelling Normalization (C18-1)

Copied to clipboard

Challenge: In this paper, we apply different NMT models to the problem of historical spelling normalization for five languages . we find that NMT model is much better than SMT in terms of character error rate .
Approach: They propose to use NMT models to solve the problem of historical spelling normalization in five languages.
Outcome: The proposed method improves historical spelling normalization for five languages.
XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE (2023.acl-short)

Copied to clipboard

Challenge: Existing approaches to the Word in Context task use cross-encoders, which prevent the possibility of deriving comparable word embeddings.
Approach: They propose a Lexical Semantic Change Detection model that extends SBERT, highlighting the target word in the sentence.
Outcome: The proposed model outperforms the state-of-the-art on the multilingual benchmarks for SemEval-2020 Task 1 - Lexical Semantic Change (LSC) Detection and the RuShiftEval shared task.
Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender Bias (2020.emnlp-main)

Copied to clipboard

Challenge: English challenge datasets highlight gender-ambiguous occurrences of ‘doctor’ as male doctors, but they are not useful for other languages.
Approach: They propose to build multi-task challenge datasets for detecting gender bias that lead to unambiguously wrong model predictions for languages with type B reflexivization.
Outcome: The proposed dataset can detect gender bias in languages with type B reflexivization and spans four languages and four NLP tasks.
Open Subtitles Paraphrase Corpus for Six Languages (L18-1)

Copied to clipboard

Challenge: Opusparcus is a new corpus of paraphrases for six European languages . it is based on movie and TV subtitles, which are colloquial and informal .
Approach: They propose to use opensubtitles2016 paraphrase corpus for six European languages . they extract paraphrases from movie and TV subtitles from the corpus .
Outcome: The new corpus is available in German, English, Finnish, French, Russian, and Swedish . it is extracted from the OpenSubtitles2016 corpus, which contains subtitles from movies and TV shows .
Paraphrase Generation and Evaluation on Colloquial-Style Sentences (2020.lrec-1)

Copied to clipboard

Challenge: a new study investigates the quality and novelty of generated paraphrases . paraphrase models can be used for information retrieval and data mining .
Approach: They use state-of-the-art neural machine translation models trained on the Opusparcus corpus to generate paraphrases in six languages.
Outcome: The proposed model outperforms the existing model on human evaluation in five of the six languages.
Can Word Sense Distribution Detect Semantic Changes of Words? (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to detect semantic variations of words are not accurate for time-sensitive predictions.
Approach: They propose to use pretrained static sense embeddings to annotate a word's occurrence with a sense id to compare its distributions.
Outcome: The proposed method compares word sense distributions across two corpora to predict meaning change . the results show that pretrained LLMs can detect changes in words over time .
How Conservative are Language Models? Adapting to the Introduction of Gender-Neutral Pronouns (2022.naacl-main)

Copied to clipboard

Challenge: a recent study shows that gender-neutral pronouns are not associated with processing difficulties . linguistic scholars have observed how technology has altered the course of language evolution .
Approach: They show that gender-neutral pronouns in Danish, English and Swedish are not associated with processing difficulties.
Outcome: a new study shows that gender-neutral pronouns are not associated with human processing difficulties . the findings suggest that such conservativity in language models may limit widespread adoption .
“We will Reduce Taxes” - Identifying Election Pledges with Language Models (2021.findings-acl)

Copied to clipboard

Challenge: a political party's manifestos are published before any election, but do they follow through? a new study uses neural models to distinguish between actual pledges and general statements .
Approach: They use election manifestos of Swedish and Indian political parties to learn neural models that distinguish actual pledges from generic positions.
Outcome: The proposed model can predict election year and manifesto's party, while context introduces noise.
Building Sentiment Lexicons for Mainland Scandinavian Languages Using Machine Translation and Sentence Embeddings (2022.lrec-1)

Copied to clipboard

Challenge: a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages is proposed . a number of experiments with Scandinavian language datasets yield state-of-the-art results using a rule-based sentiment analysis algorithm.
Approach: They propose a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages.
Outcome: The proposed method is based on the English Sentiwordnet and a thesaurus in one of the target languages.
Collecting Linguistic Resources for Assessing Children’s Pronunciation of Nordic Languages (2024.lrec-main)

Copied to clipboard

Challenge: Using annotated corpora of languages is difficult for children learning a foreign language . most effort is directed to the most popular languages and adult learners .
Approach: They collect annotated corpora of languages spoken by children in three Nordic countries . they hope to make the data available for future research .
Outcome: The collected data will be used to develop and evaluate computer assisted pronunciation assessment systems for non-native children learning a Nordic language (L2) and for L1 children with speech sound disorder (SSD).
Better, Faster, Stronger Sequence Tagging Constituent Parsers (N19-1)

Copied to clipboard

Challenge: Existing efforts to speed up constituent parsing have focused on chart-based or shift-reduce parsers.
Approach: They propose to use auxiliary losses and sentence-level fine-tuning to mitigate greedy decoding issues.
Outcome: The proposed model surpasses the performance of sequence tagging constituent parsers on the English and Chinese Penn Treebank datasets and reduces their parsing time even further.
A Parallel WordNet for English, Swedish and Bulgarian (2020.lrec-1)

Copied to clipboard

Challenge: a new WordNet resource for Swedish and Bulgarian is created that is tightly aligned with the Princeton WordNet.
Approach: They propose a WordNet resource for Swedish and Bulgarian that is tightly aligned with Princeton WordNet.
Outcome: The proposed resource is tightly aligned with the Princeton WordNet for Swedish and Bulgarian . the new resource is open-source and in its development used only existing resources.
The FISKMÖ Project: Resources and Tools for Finnish-Swedish Machine Translation and Cross-Linguistic Research (2020.lrec-1)

Copied to clipboard

Challenge: Finnish and Swedish are the two official languages of Finland.
Approach: They propose to compile a massive corpus of translated material between Finnish and Swedish . they also aim to develop open and freely accessible translation services for those two languages .
Outcome: The project aims to develop open and freely accessible translation services for Finnish and Swedish.
Multilingual Culture-Independent Word Analogy Datasets (2020.lrec-1)

Copied to clipboard

Challenge: In text processing, deep neural networks use word embeddings as an input.
Approach: They propose to use benchmark datasets to compare the quality of word embeddings in text processing . they use a word analogy task in Croatian, English, Estonian, Finnish, Latvian, Lithuanian, Russian, Slovenian, and Swedish .
Outcome: The proposed datasets are culturally independent and cross-lingual for the languages used.
Swap and Predict – Predicting the Semantic Changes in Words across Corpora by Context Swapping (2023.findings-emnlp)

Copied to clipboard

Challenge: Detecting semantic changes of words is an important task for various NLP applications that must make time-sensitive predictions.
Approach: They propose a method that randomly swaps contexts between two different corpora to detect whether a given word changes its meaning . they then use a pretrained masked language model to generate contextualised word embeddings of w, which are then used to predict the semantic changes of words in four languages .
Outcome: The proposed method achieves significant performance improvements compared to baselines for the English semantic change prediction task.
High Quality ELMo Embeddings for Seven Less-Resourced Languages (2020.lrec-1)

Copied to clipboard

Challenge: Recent results show that deep neural networks using contextual embeddings outperform non-contextual embedders on a majority of text classification tasks.
Approach: They propose to use contextual embeddings for seven languages to train new embeddables . they also show that existing embeddibles for listed languages shall be improved .
Outcome: The proposed embeddings outperform non-contextual embeddables on a majority of text classification tasks.
SemR-11: A Multi-Lingual Gold-Standard for Semantic Similarity and Relatedness for Eleven Languages (L18-1)

Copied to clipboard

Challenge: SemR-11 is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Approach: This paper describes a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Outcome: The dataset is a multi-lingual dataset for evaluating semantic similarity and relatedness for 11 languages.
Training a Swedish Constituency Parser on Six Incompatible Treebanks (2020.lrec-1)

Copied to clipboard

Challenge: Syntactic parsing is a widely used intermediate step in several natural language processing tasks.
Approach: They propose to use a function-tagged constituent treebank for Swedish which includes discontinuous constituents to improve the accuracy.
Outcome: The proposed parser can be trained on additional treebanks that use other annotation models.
The EDGeS Diachronic Bible Corpus (2020.lrec-1)

Copied to clipboard

Challenge: EDGeS is a diachronic and parallel corpus of Bible translations in Dutch, English, German and Swedish . it is intended to be used for longitudinal studies of complex verb constructions in Germanic .
Approach: They present the EDGeS Diachronic Bible Corpus, a diachronic corpus of Bible translations in Dutch, English, German and Swedish . they use a synchronically and synchronly parallel corpus to study complex verb constructions in Germanic .
Outcome: The EDGeS is a diachronic and parallel corpus of Bible translations in Dutch, English, German and Swedish spanning six and a half centuries.
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model (2025.emnlp-main)

Copied to clipboard

Challenge: Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content.
Approach: They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set .
Outcome: The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu.
Nested Noun Phrase Identification Using BERT (2024.lrec-main)

Copied to clipboard

Challenge: a number of methods for identifying noun phrases have been developed . chunking-like methods do not represent the fact that noun phrase can be nested.
Approach: They propose a method of finding all noun phrases in a sentence nested to an arbitrary depth using the BERT model for token classification.
Outcome: The proposed method achieves very good results for both Swedish and English . it is based on the BERT model for token classification .
Data-Constrained Synthesis of Training Data for De-Identification (2025.acl-long)

Copied to clipboard

Challenge: sensitive domains lack widely available datasets due to privacy risks . recent studies have focused on evaluating the privacy of the synthetic text .
Approach: They domain-adapt LLMs to clinical domain and generate synthetic clinical texts . they then generate NER models that can be annotated with tags for PII .
Outcome: The proposed model performs better than the original model using smaller datasets.
Transformer-based Swedish Semantic Role Labeling through Transfer Learning (2024.lrec-main)

Copied to clipboard

Challenge: Semantic Role Labeling (SRL) is a task in natural language understanding where the goal is to extract semantic roles for a given sentence.
Approach: They propose to build a Transformer-based SRL system for Swedish by exploring multilingual and cross-lingual transfer learning methods and leveraging the Swedish FrameNet resource.
Outcome: The proposed model outperforms two different cross-lingual transfer models and shows that the multilingual learning outperformed the other models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations