Papers by Paul Cook

9 papers
Joint Training for Learning Cross-lingual Embeddings with Sub-word Information without Parallel Corpora (2020.starsem-1)

Copied to clipboard

Challenge: Existing methods for learning cross-lingual word embeddings incorporate sub-word information during training.
Approach: They propose a method that incorporates sub-word information during training to learn cross-lingual word embeddings from monolingual data and a bilingual lexicon.
Outcome: The proposed method improves on bilingual lexicon induction, monolingual word similarity, and document classification using low-resource languages.
Leveraging distributed representations and lexico-syntactic fixedness for token-level prediction of the idiomaticity of English verb-noun combinations (P18-2)

Copied to clipboard

Challenge: Verb-noun combinations (VNCs) are ambiguous between literal and idiomatic usages in English.
Approach: They propose and evaluate models for classifying verb-noun combinations as idiomatic or literal, based on averaging word embeddings and a variety of approaches to forming distributed representations.
Outcome: The proposed model outperforms a previous model based on skip-thoughts and averaging word embeddings.
Evaluating Sub-word Embeddings in Cross-lingual Models (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches to learning sub-word embeddings for out-of-vocabulary words have not considered sub- word embedds in cross-lingual models.
Approach: They propose to use sub-word embeddings to form cross-lingual embeddables for out-of-vocabulary (OOV) words for which no embeddibles are available.
Outcome: The proposed bilingual lexicon induction task shows that sub-word embeddings can be leveraged to form cross-lingual embeddables for OOV words.
Evaluating Approaches to Personalizing Language Models (2020.lrec-1)

Copied to clipboard

Challenge: a large amount of text is not available for training a user-specific language model, which suggests a need to personalize language models with only a small amount of data.
Approach: They propose three approaches to personalize a language model that was trained on a large background corpus using a relatively small amount of text from an individual user.
Outcome: The proposed techniques outperform language model adaptation based on demographic factors.
Evaluating a Joint Training Approach for Learning Cross-lingual Embeddings with Sub-word Information without Parallel Corpora on Lower-resource Languages (2021.starsem-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings provide a way for information to be transferred between languages.
Approach: They propose a joint training approach that incorporates sub-word information during training to learn cross-lingual embeddings.
Outcome: The proposed method improves bilingual lexicon induction, especially for out-of-vocabulary words (OOVs) it is able to represent out- of-vocal words (OVs) and is more isomorphic than previous methods.
Evaluating the Impact of Sub-word Information and Cross-lingual Word Embeddings on Mi’kmaq Language Modelling (2020.lrec-1)

Copied to clipboard

Challenge: Mi'kmaq is an Indigenous language spoken primarily in Eastern Canada.
Approach: They consider n-gram and RNN language models for Mi'kmaq and use them to investigate their performance.
Outcome: The proposed model performs better than word-level models, but does not improve over word-based models.
Towards Language Technology for Mi’kmaq (L18-1)

Copied to clipboard

Challenge: Mi'kmaq is a polysynthetic Indigenous language spoken primarily in Eastern Canada .
Approach: They construct and analyze a web corpus of Mi'kmaq and evaluate several approaches to language modelling . they argue that natural language processing could aid efforts to preserve Indigenous languages .
Outcome: The proposed language model is based on a web corpus of Mi'kmaq . the model is well-suited to morphologically-rich languages, the authors argue .
Leveraging a Bilingual Dictionary to Learn Wolastoqey Word Representations (2022.lrec-1)

Copied to clipboard

Challenge: Existing word embeddings for lowresource languages require large corpora of running text to learn high quality representations.
Approach: They leverage a bilingual dictionary to learn Wolastoqey word embeddings by encoding their corresponding English definitions into vector representations using pretrained English word and sequence representation models.
Outcome: The proposed model outperforms baseline models without language-specific training or fine-tuning.
WaCadie: Towards an Acadian French Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora do not exist for many languages and language varieties, such as Acadian French.
Approach: They propose to build a corpus of Acadian French using web-as-corpus methodologies . they use domain crawling, social media scraping, and search engines to create corpus .
Outcome: The proposed corpus includes some traces of Acadian French, but it is not available for many languages and language varieties, such as Acadinian French.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations