Papers by Antoine Doucet
A Dataset for Multi-lingual Epidemiological Event Extraction (2020.lrec-1)
Copied to clipboard
| Challenge: | Using the Web, we propose a corpus for information extraction and text classification. |
| Approach: | They propose to use a corpus for information extraction and natural language processing (NLP) tasks such as text classification. |
| Outcome: | The proposed corpus can be used for information extraction and natural language processing tasks such as text classification. |
Multi-TimeLine Summarization (MTLS): Improving Timeline Summarization by Generating Multiple Summaries (2021.acl-long)
Copied to clipboard
| Challenge: | Existing work on Time-Line Summarization (TLS) has focused on improving the performance of summarization but its drawbacks are as follows: a homogeneous dataset makes it hard to generalize; output is usually a single timeline regardless of the size and complexity of the input dataset. |
| Approach: | They propose a task that generates a time-line for each story given a news article . they propose MTLS task that can be generalized to other news articles . |
| Outcome: | The proposed task can generate bet-ter results than Time-Line Summarization (TLS) the proposed task is based on previous evaluation methods. |
Dataset for Temporal Analysis of English-French Cognates (2020.lrec-1)
Copied to clipboard
| Challenge: | Using computational techniques to study language evolution has gained much attention . comparing two or more languages can shed light on how they co-evolve . |
| Approach: | They propose to use a dataset to investigate the similarity in evolution between languages by comparing cognates across time. |
| Outcome: | The proposed dataset is the first to use computational approaches and large data to make a cross-language diachronic analysis. |
Multilingual Epidemiological Text Classification: A Comparative Study (2020.coling-main)
Copied to clipboard
| Challenge: | a comparative study of multilingual text classification models analyzes the performance of different models based on different languages . low-resource languages are highly influenced by typology of the languages on which the models have been trained or fine-tuned but also by their size. |
| Approach: | They compare machine and deep learning models with a dataset of epidemiological news articles . they find that the performance of the models is proportionate to the training data size . |
| Outcome: | The proposed model outperforms baseline models on a multilingual text classification task . low-resource languages are highly influenced by typology of languages and their size . |