Papers by Julien Abadji

2 papers
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing web crawling pipelines are used to collect large corpora raw data, but the main way to collect such data is through manual data extraction.
Approach: They propose to use a web crawler to extract and classify data from a multilingual web corpus and an automated annotation pipeline to improve it.
Outcome: The proposed version of OSCAR could be used to pre-train large generative language models and other applications in Natural Language Processing and Digital Humanities.
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that multimodal large language models can learn from text-image data.
Approach: They propose to train multimodal large language models on large amounts of text-image data . they also show a boost in few-shot learning performance across various multilingual tasks .
Outcome: The proposed dataset is not public and is only in English . it is the first large-scale multilingual and multimodal document corpus crawled from the web.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations