Papers by Armand Joulin

10 papers
Training Hybrid Language Models by Marginalizing over Segmentations (P19-1)

Copied to clipboard

Challenge: Statistical language modeling is the problem of estimating a probability distribution over text data.
Approach: They propose to marginalize over the segmentations efficiently to compute the true probability of a sequence.
Outcome: The proposed model marginalizes over the segmentations to compute the true probability of a sequence on three datasets comprising seven languages.
Loss in Translation: Learning Bilingual Word Mapping with a Retrieval Criterion (D18-1)

Copied to clipboard

Challenge: Existing approaches to learn orthogonal matrix aligning bilingual lexicons are suboptimal . resulting models suffer from "hubness problem" because word vectors tend to be nearest neighbors of abnormally high number of other words.
Approach: They propose a unified formulation that directly optimizes a retrieval criterion in an end-to-end fashion.
Outcome: The proposed approach outperforms the state-of-the-art on word translation on standard benchmarks.
Adaptive Attention Span in Transformers (P19-1)

Copied to clipboard

Challenge: We extend the maximum context size of a neural network called Transformer to 8k characters.
Approach: They propose a self-attention mechanism that can learn its optimal attention span . this allows for models with longer context and the capability to catch longer dependencies.
Outcome: The proposed model achieves state-of-the-art performance on text8 and enwiki8 using 8k characters with no loss of performance, and maintains control over memory footprint and computational time.
CCMatrix: Mining Billions of High-Quality Parallel Sentences on the Web (2021.acl-long)

Copied to clipboard

Challenge: Using a curated common crawl corpus, we were able to mine 10.8 billion parallel sentences out of which only 2.9 billions are aligned with English.
Approach: They use 32 snapshots of a curated common crawl corpus totaling 71 billion unique sentences to mine 10.8 billion parallel sentences out of which only 2.9 billions are aligned with English.
Outcome: The proposed system outperforms the best single systems on the WMT’19 test set for English-German/Russian/Chinese and outperformed the best submission at the 2020 WAT workshop.
Time Sensitive Knowledge Editing through Efficient Finetuning (2024.acl-short)

Copied to clipboard

Challenge: Existing locate-and-edit knowledge editing methods suffer from two limitations: they are infeasible for large scale KE in practice and require long run-time.
Approach: They propose to use parametric fine-tuning techniques to update obsolete knowledge and induce new knowledge into LLMs.
Outcome: The proposed methods improve the performance of KE and knowledge update in a temporal dataset with knowledge update and knowledge injection examples.
Learning Word Vectors for 157 Languages (L18-1)

Copied to clipboard

Challenge: Distributed word representations, or word vectors, have been used in natural language processing for many tasks.
Approach: They propose to use the encyclopedia Wikipedia and the common crawl corpus to train distributed word representations on large corpora and use them in downstream tasks.
Outcome: The proposed model performs very well on 10 languages for which evaluation dataset exists.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data (2020.lrec-1)

Copied to clipboard

Challenge: Pre-training text representations have led to significant improvements in many areas of natural language processing.
Approach: They propose a pipeline to extract monolingual datasets from Common Crawl . pipeline follows data processing introduced in fastText that deduplicates documents .
Outcome: The proposed pipeline performs standard document deduplication and language identification similar to the pipeline introduced in fastText and a filtering step to select documents close to high quality corpora like Wikipedia.
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.
Target Conditioning for One-to-Many Generation (2020.findings-emnlp)

Copied to clipboard

Challenge: Neural Machine Translation models lack diversity in their generated translations, even when paired with search algorithm, like beam search.
Approach: They propose to model one-to-many mapping by conditioning a decoder on a latent variable that represents the domain of target sentences.
Outcome: The proposed method can scale to any number of domains without affecting performance or training time.
Cooperative Learning of Disjoint Syntax and Semantics (N19-1)

Copied to clipboard

Challenge: Existing models that learn to jointly infer an expression’s syntactic structure and its semantics fail to learn the correct parsing strategy on mathematical expressions generated from a simple context-free grammar.
Approach: They propose a recursive model that learns to jointly infer an expression’s syntactic structure and its semantics without requiring a formal supervision.
Outcome: The proposed model performs competitively on several natural language tasks, such as Natural Language Inference and Sentiment Analysis.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations