Papers by Nicolas Ballier

4 papers
A Manually Annotated Resource for the Investigation of Nasal Grunts (2020.lrec-1)

Copied to clipboard

Challenge: acoustic annotation of nasal grunts is described in the whole CID corpus of the french language . acculturation of non-lexical conversational sounds has been debated for a long time .
Approach: They propose an annotation framework for nasal grunts of the whole French CID corpus . they characterise acoustic cues and visual cue conventions followed for the annotation .
Outcome: The proposed framework is based on the entire French CID corpus.
Beyond WER: Probing Whisper’s Sub‐token Decoder Across Diverse Language Resource Levels (2025.emnlp-main)

Copied to clipboard

Challenge: Large multilingual automatic speech recognition models achieve remarkable performance, but the internal mechanisms of the end-to-end pipeline remain underexplored.
Approach: They propose to analyze Whisper's multilingual decoder to uncover systematic decoding disparities masked by aggregate error rates.
Outcome: The proposed model performs better on higher resource languages, but lower resource languages fare worse on these metrics.
The Learnability of the Annotated Input in NMT Replicating (Vanmassenhove and Way, 2018) with OpenNMT (2020.lrec-1)

Copied to clipboard

Challenge: reproducibility of experiments is a key issue in Neural Networks, which are fed with variable samples of training data.
Approach: They reproduce some of the experiments related to neural network training for Machine Translation as reported in . they annotated a sample from the EN-FR and EN-DE Europarl with syntactic and semantic annotations to train neural networks with the Nematus Neural Machine Translation toolkit.
Outcome: The results obtained were lower than the original paper, but on a more limited set of annotations.
Logging Keystrokes in Writing by English Learners (2024.lrec-main)

Copied to clipboard

Challenge: Essay writing is a skill commonly taught and practised in schools.
Approach: They collect and analyse data representing the essay writing process from start to finish by recording every keystroke from multiple writers participating in the study.
Outcome: The data collected from 1,006 writers is compared against a standard dataset of texts, keystroke logs and metadata for public release.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations