Papers by Valentin Barriere

9 papers
StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos (2025.findings-emnlp)

Copied to clipboard

Challenge: a new multimodal dataset of stand-up comedies is proposed to improve humor detection . the dataset is the biggest available for this type of task, and the most diverse .
Approach: They propose a method to enhance the automatic laughter detection based on Audio Speech Recognition errors.
Outcome: The proposed method improves existing models of humor detection by using audio speech recognition errors.
Opinions in Interactions : New Annotations of the SEMAINE Database (2022.lrec-1)

Copied to clipboard

Challenge: a new method for the detection of opinions in interactions is proposed . a dataset of dyadic interactions is annotated continuously in two affective dimensions related to the emotions .
Approach: They propose to annotate opinions over a multimodal corpus of dyadic interactions . they use a d-acting algorithm to annnotate the opinions of a speaker .
Outcome: The proposed method allows to obtain a precise annotation regarding the opinion of a speaker.
Deep Natural Language Feature Learning for Interpretable Prediction (2023.emnlp-main)

Copied to clipboard

Challenge: Using a small transformer language model, we can break down a complex task into a set of intermediary easier sub-tasks.
Approach: They propose a method to break down a main task into a set of intermediary easier sub-tasks, which are formulated in natural language as binary questions related to the final target task.
Outcome: The proposed method breaks down a complex task into a set of easier sub-tasks, which are formulated in natural language as binary questions related to the final target task.
A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research shows that named entities influence PLMs in many applications.
Approach: They propose a method to quantify biases associated with named entities from various countries using Twitter data instead of templates or specific datasets.
Outcome: The proposed method shows positive biases related to the language spoken in a country across all classifiers.
Are Text Classifiers Xenophobic? A Country-Oriented Bias Detection Method with Least Confounding Variables (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for detecting biases are biased because of confounding variables . authors propose a method to detect the biased classifier on any type of unlabeled data .
Approach: They propose a method to detect biases of a specific fine-tuned classifier on unlabeled data.
Outcome: The proposed method detects biases on unlabeled data on named entity perturbations . it uses name-entity recognition on target-domain data and morphosynctactically different languages spoken in relation to countries of the target groups .
Improving Sentiment Analysis over non-English Tweets using Multilingual Transformers and Automatic Translation for Data-Augmentation (2020.coling-main)

Copied to clipboard

Challenge: Existing models for sentiment analysis over tweets require a substantial amount of text to adapt to a domain where the syntax is different.
Approach: They propose to use a multilingual transformer model to train over tweets in five different languages to adapt the model to non-English languages.
Outcome: The proposed model improves over small corpora of tweets in non-English languages.
Adapting Bias Evaluation to Domain Contexts using Generative Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to assess social bias in NLP systems face limitations in scalability and fidelity across domains.
Approach: They propose a domain-adaptive framework that uses prompting with Large Language Models to automatically transform template-based bias datasets into domain-specific variants.
Outcome: The proposed framework improves the accuracy and contextual relevance of bias evaluations in socially relevant datasets.
CoFE: A New Dataset of Intra-Multilingual Multi-target Stance Classification from an Online European Participatory Democracy Platform (2022.aacl-short)

Copied to clipboard

Challenge: Stance Recognition is a useful tool for many real-life applications, from misinformation detection to poll verification.
Approach: They propose to use an online debating platform where users can submit proposals and comment over proposals or over other comments.
Outcome: The proposed dataset contains 4.2k proposals and 20k comments on various topics.
The Touché23-ValueEval Dataset for Identifying Human Values behind Arguments (2024.lrec-main)

Copied to clipboard

Challenge: Cultural norms can influence the prioritization of values, leading to distinct perspectives on debatable topics.
Approach: They present a Touché23-ValueEval dataset that annotates 4780 new arguments and annotated 54 human values.
Outcome: The Touché23-ValueEval dataset doubles the original Webis-ArgValués-22 dataset to 9324 arguments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations