Papers by Martin d’Hoffschmidt
On the importance of pre-training data volume for compact language models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in language modeling have led to computationally intensive and resource-demanding state-of-the-art models. |
| Approach: | They investigate the impact of pre-training data volume on compact language models . they use a French question answering task to train models with as little as 100 MB of text . |
| Outcome: | The results show that pre-training data volume can improve models with as little as 100 MB of text . the results suggest that the model performance is poorer with less data than with larger datasets . |
FQuAD: French Question Answering Dataset (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in the field of language modeling have improved state-of-the-art results on many natural language processing tasks. |
| Approach: | They propose to use a French Question Answering Dataset to track progress of French Question answering models. |
| Outcome: | The proposed model achieves an F1 score of 92.2 and an exact match ratio of 82.1 on the test set. |