Papers by Martin d’Hoffschmidt

2 papers
On the importance of pre-training data volume for compact language models (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in language modeling have led to computationally intensive and resource-demanding state-of-the-art models.
Approach: They investigate the impact of pre-training data volume on compact language models . they use a French question answering task to train models with as little as 100 MB of text .
Outcome: The results show that pre-training data volume can improve models with as little as 100 MB of text . the results suggest that the model performance is poorer with less data than with larger datasets .
FQuAD: French Question Answering Dataset (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in the field of language modeling have improved state-of-the-art results on many natural language processing tasks.
Approach: They propose to use a French Question Answering Dataset to track progress of French Question answering models.
Outcome: The proposed model achieves an F1 score of 92.2 and an exact match ratio of 82.1 on the test set.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations