Papers by Piotr Bojanowski

7 papers
Training Hybrid Language Models by Marginalizing over Segmentations (P19-1)

Copied to clipboard

Challenge: Statistical language modeling is the problem of estimating a probability distribution over text data.
Approach: They propose to marginalize over the segmentations efficiently to compute the true probability of a sequence.
Outcome: The proposed model marginalizes over the segmentations to compute the true probability of a sequence on three datasets comprising seven languages.
Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion (D18-1)

Copied to clipboard

Challenge: Existing approaches to learn orthogonal matrix aligning bilingual lexicons are suboptimal . resulting models suffer from "hubness problem" because word vectors tend to be nearest neighbors of abnormally high number of other words.
Approach: They propose a unified formulation that directly optimizes a retrieval criterion in an end-to-end fashion.
Outcome: The proposed approach outperforms the state-of-the-art on word translation on standard benchmarks.
Adaptive Attention Span in Transformers (P19-1)

Copied to clipboard

Challenge: We extend the maximum context size of a neural network called Transformer to 8k characters.
Approach: They propose a self-attention mechanism that can learn its optimal attention span . this allows for models with longer context and the capability to catch longer dependencies.
Outcome: The proposed model achieves state-of-the-art performance on text8 and enwiki8 using 8k characters with no loss of performance, and maintains control over memory footprint and computational time.
Colorless Green Recurrent Networks Dream Hierarchically (N18-1)

Copied to clipboard

Challenge: Recurrent neural networks (RNNs) can induce non-trivial properties of language.
Approach: They investigate whether RNNs can track hierarchical syntactic structure . they include nonsensical sentences where RNN cannot rely on semantic cues .
Outcome: The proposed models can predict long-distance agreement in nonsensical sentences in Italian and English.
Learning Word Vectors for 157 Languages (L18-1)

Copied to clipboard

Challenge: Distributed word representations, or word vectors, have been used in natural language processing for many tasks.
Approach: They propose to use the encyclopedia Wikipedia and the common crawl corpus to train distributed word representations on large corpora and use them in downstream tasks.
Outcome: The proposed model performs very well on 10 languages for which evaluation dataset exists.
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.
Misspelling Oblivious Word Embeddings (N19-1)

Copied to clipboard

Challenge: Existing word embeddings have limited applicability to malformed texts . misspellings are frequent and embeddable for words that have not been observed at training time .
Approach: They propose a method to learn word embeddings that are resilient to misspellings . they use FastText with subwords to train embeddables on a new dataset .
Outcome: The proposed method is tested on a publicly available dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations