Papers by Antoine Doucet

4 papers
A Dataset for Multi-lingual Epidemiological Event Extraction (2020.lrec-1)

Copied to clipboard

Challenge: Using the Web, we propose a corpus for information extraction and text classification.
Approach: They propose to use a corpus for information extraction and natural language processing (NLP) tasks such as text classification.
Outcome: The proposed corpus can be used for information extraction and natural language processing tasks such as text classification.
Multi-TimeLine Summarization (MTLS): Improving Timeline Summarization by Generating Multiple Summaries (2021.acl-long)

Copied to clipboard

Challenge: Existing work on Time-Line Summarization (TLS) has focused on improving the performance of summarization but its drawbacks are as follows: a homogeneous dataset makes it hard to generalize; output is usually a single timeline regardless of the size and complexity of the input dataset.
Approach: They propose a task that generates a time-line for each story given a news article . they propose MTLS task that can be generalized to other news articles .
Outcome: The proposed task can generate bet-ter results than Time-Line Summarization (TLS) the proposed task is based on previous evaluation methods.
Dataset for Temporal Analysis of English-French Cognates (2020.lrec-1)

Copied to clipboard

Challenge: Using computational techniques to study language evolution has gained much attention . comparing two or more languages can shed light on how they co-evolve .
Approach: They propose to use a dataset to investigate the similarity in evolution between languages by comparing cognates across time.
Outcome: The proposed dataset is the first to use computational approaches and large data to make a cross-language diachronic analysis.
Multilingual Epidemiological Text Classification: A Comparative Study (2020.coling-main)

Copied to clipboard

Challenge: a comparative study of multilingual text classification models analyzes the performance of different models based on different languages . low-resource languages are highly influenced by typology of the languages on which the models have been trained or fine-tuned but also by their size.
Approach: They compare machine and deep learning models with a dataset of epidemiological news articles . they find that the performance of the models is proportionate to the training data size .
Outcome: The proposed model outperforms baseline models on a multilingual text classification task . low-resource languages are highly influenced by typology of languages and their size .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations