Papers by Alina Wróblewska
Polish Corpus of Annotated Descriptions of Images (L18-1)
Copied to clipboard
| Challenge: | a new dataset of image descriptions is presented in Polish . the dataset is too small for training a sophisticated language-vision system. |
| Approach: | They propose to use a Polish dataset to analyze image descriptions . the descriptions are morphosyntactically analysed and annotated by human annotators . |
| Outcome: | The proposed model learns about the inter-modal correspondences between language and vision. |
Towards the Conversion of National Corpus of Polish to Universal Dependencies (2020.lrec-1)
Copied to clipboard
| Challenge: | a paper aims at enriching the manually annotated part of National Corpus of Polish with a syntactic layer. |
| Approach: | They enrich manually annotated part of Polish National Corpus with a syntactic layer and a UD dependency graph. |
| Outcome: | The proposed model outperforms a model trained on a smaller set of gold-standard trees in predicting part-of-speech tags, morphological features, lemmata and labelled dependency trees. |
Empirical Linguistic Study of Sentence Embeddings (P19-1)
Copied to clipboard
| Challenge: | a new method of analysing sentence embeddings shows that linguistic information is retained in the vector representations of sentences. |
| Approach: | They propose a method of analysing the content of sentence embeddings based on probing tasks and contrasting languages. |
| Outcome: | The proposed method is based on probing tasks and classification datasets for two contrasting languages. |
NLPre: A Revised Approach towards Language-centric Benchmarking of Natural Language Preprocessing Systems (2024.lrec-main)
Copied to clipboard
| Challenge: | GLUE benchmarking system enables ongoing evaluation of multiple NLPre tools while credibly tracking their performance. |
| Approach: | They propose a language-centric benchmarking system that enables ongoing evaluation of multiple NLPre tools while credibly tracking their performance. |
| Outcome: | The proposed system is configured for Polish and integrated with the thoroughly assembled NLPre-PL benchmark. |
COMBO: State-of-the-Art Morphosyntactic Analysis (2021.emnlp-demo)
Copied to clipboard
| Challenge: | COMBO is an end-to-end NLP system for accurate part-of-speech tagging, morphological analysis, and (enhanced) dependency parsing. |
| Approach: | They propose a fully neural NLP system for accurate part-of-speech tagging, morphological analysis, lemmatisation, and (enhanced) dependency parsing. |
| Outcome: | The proposed system predicts categorical morphosyntactic features whilst also exposes their vector representations, extracted from hidden layers. |