Papers with Polish
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing (2025.emnlp-main)
Copied to clipboard
Zhisheng Zheng, Puyuan Peng, Anuj Diwan, Cong Phuoc Huynh, Xiaohang Sun, Zhu Liu, Vimal Bhat, David Harwath
| Challenge: | Autoregressive language model for multilingual speech editing and zero-shot text-to-speech synthesis is available in 11 languages. |
| Approach: | They introduce an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech synthesis across 11 languages. |
| Outcome: | The model generates high-quality, natural-sounding speech, even with limited per-language data . it shows robust performance in diverse linguistic settings, even in limited per language data compared to other models . |
Acquiring Verb Classes Through Bottom-Up Semantic Verb Clustering (L18-1)
Copied to clipboard
| Challenge: | Existing methods for creating verbal classifications are limited or non-existent in most languages . a range of automatic verb classification approaches have been proposed, but high-quality resources are needed . |
| Approach: | They propose to use top-up semantic clustering to extract syntactic and semantic information from verbs in English, Polish and Croatian. |
| Outcome: | The proposed classifications in English, Polish and Croatian are compared with other languages. |
Better, Faster, Stronger Sequence Tagging Constituent Parsers (N19-1)
Copied to clipboard
| Challenge: | Existing efforts to speed up constituent parsing have focused on chart-based or shift-reduce parsers. |
| Approach: | They propose to use auxiliary losses and sentence-level fine-tuning to mitigate greedy decoding issues. |
| Outcome: | The proposed model surpasses the performance of sequence tagging constituent parsers on the English and Chinese Penn Treebank datasets and reduces their parsing time even further. |
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | This work presents a corpus manually annotated with named entities for six Slavic languages . |
| Approach: | They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models . |
| Outcome: | The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier. |
Manual Clustering and Spatial Arrangement of Verbs for Multilingual Evaluation and Typology Analysis (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods to learn general language representations from large volumes of unlabeled text have been used to improve multilingual NLP. |
| Approach: | They propose to use a spatial arrangement method to generate large-scale evaluation datasets that balance cross-lingual alignment with language specificity. |
| Outcome: | The proposed method produces semantic verb classes and fine-grained similarity scores for nearly 130 thousand verb pairs. |
Learning Word Vectors for 157 Languages (L18-1)
Copied to clipboard
| Challenge: | Distributed word representations, or word vectors, have been used in natural language processing for many tasks. |
| Approach: | They propose to use the encyclopedia Wikipedia and the common crawl corpus to train distributed word representations on large corpora and use them in downstream tasks. |
| Outcome: | The proposed model performs very well on 10 languages for which evaluation dataset exists. |
Modeling the Impact of Syntactic Distance and Surprisal on Cross-Slavic Text Comprehension (2022.lrec-1)
Copied to clipboard
| Challenge: | Using symmetric measures of insertion, deletion and movement of syntactic units, we investigate phonetic and orthographic asymmetries between selected languages. |
| Approach: | They focus on the syntactic variation and measure syntaktic distances between nine Slavic languages using symmetric measures of insertion, deletion and movement of syntak units in parallel sentences of the fable “The North Wind and the Sun”. |
| Outcome: | The proposed measures are validated on spoken and written cloze tests for Slavic native speakers to determine whether variations in syntax lead to slower or impeded intercomprehension of Slav texts. |