Papers by Paul Cook
Joint Training for Learning Cross-lingual Embeddings with Sub-word Information without Parallel Corpora (2020.starsem-1)
Copied to clipboard
| Challenge: | Existing methods for learning cross-lingual word embeddings incorporate sub-word information during training. |
| Approach: | They propose a method that incorporates sub-word information during training to learn cross-lingual word embeddings from monolingual data and a bilingual lexicon. |
| Outcome: | The proposed method improves on bilingual lexicon induction, monolingual word similarity, and document classification using low-resource languages. |
Leveraging distributed representations and lexico-syntactic fixedness for token-level prediction of the idiomaticity of English verb-noun combinations (P18-2)
Copied to clipboard
| Challenge: | Verb-noun combinations (VNCs) are ambiguous between literal and idiomatic usages in English. |
| Approach: | They propose and evaluate models for classifying verb-noun combinations as idiomatic or literal, based on averaging word embeddings and a variety of approaches to forming distributed representations. |
| Outcome: | The proposed model outperforms a previous model based on skip-thoughts and averaging word embeddings. |
Evaluating Sub-word Embeddings in Cross-lingual Models (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing approaches to learning sub-word embeddings for out-of-vocabulary words have not considered sub- word embedds in cross-lingual models. |
| Approach: | They propose to use sub-word embeddings to form cross-lingual embeddables for out-of-vocabulary (OOV) words for which no embeddibles are available. |
| Outcome: | The proposed bilingual lexicon induction task shows that sub-word embeddings can be leveraged to form cross-lingual embeddables for OOV words. |
Evaluating Approaches to Personalizing Language Models (2020.lrec-1)
Copied to clipboard
| Challenge: | a large amount of text is not available for training a user-specific language model, which suggests a need to personalize language models with only a small amount of data. |
| Approach: | They propose three approaches to personalize a language model that was trained on a large background corpus using a relatively small amount of text from an individual user. |
| Outcome: | The proposed techniques outperform language model adaptation based on demographic factors. |
Evaluating a Joint Training Approach for Learning Cross-lingual Embeddings with Sub-word Information without Parallel Corpora on Lower-resource Languages (2021.starsem-1)
Copied to clipboard
| Challenge: | Cross-lingual word embeddings provide a way for information to be transferred between languages. |
| Approach: | They propose a joint training approach that incorporates sub-word information during training to learn cross-lingual embeddings. |
| Outcome: | The proposed method improves bilingual lexicon induction, especially for out-of-vocabulary words (OOVs) it is able to represent out- of-vocal words (OVs) and is more isomorphic than previous methods. |
Evaluating the Impact of Sub-word Information and Cross-lingual Word Embeddings on Mi’kmaq Language Modelling (2020.lrec-1)
Copied to clipboard
| Challenge: | Mi'kmaq is an Indigenous language spoken primarily in Eastern Canada. |
| Approach: | They consider n-gram and RNN language models for Mi'kmaq and use them to investigate their performance. |
| Outcome: | The proposed model performs better than word-level models, but does not improve over word-based models. |
Towards Language Technology for Mi’kmaq (L18-1)
Copied to clipboard
| Challenge: | Mi'kmaq is a polysynthetic Indigenous language spoken primarily in Eastern Canada . |
| Approach: | They construct and analyze a web corpus of Mi'kmaq and evaluate several approaches to language modelling . they argue that natural language processing could aid efforts to preserve Indigenous languages . |
| Outcome: | The proposed language model is based on a web corpus of Mi'kmaq . the model is well-suited to morphologically-rich languages, the authors argue . |
Leveraging a Bilingual Dictionary to Learn Wolastoqey Word Representations (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing word embeddings for lowresource languages require large corpora of running text to learn high quality representations. |
| Approach: | They leverage a bilingual dictionary to learn Wolastoqey word embeddings by encoding their corresponding English definitions into vector representations using pretrained English word and sequence representation models. |
| Outcome: | The proposed model outperforms baseline models without language-specific training or fine-tuning. |
WaCadie: Towards an Acadian French Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing corpora do not exist for many languages and language varieties, such as Acadian French. |
| Approach: | They propose to build a corpus of Acadian French using web-as-corpus methodologies . they use domain crawling, social media scraping, and search engines to create corpus . |
| Outcome: | The proposed corpus includes some traces of Acadian French, but it is not available for many languages and language varieties, such as Acadinian French. |