Papers by Federico Bianchi

16 papers
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Prior studies assess instruction adherence in the model’s main responses, but it is also critical for large reasoning models to follow user instructions throughout their reasoning process.
Approach: They propose a systematic benchmark for assessing reasoning instruction following to assess the model's adherence to instructions.
Outcome: The proposed benchmark reduces the risk of undesirable shortcuts, hallucinations, or reward hacking within reasoning traces.
Twitter-Demographer: A Flow-based Tool to Enrich Twitter Data (2022.emnlp-demos)

Copied to clipboard

Challenge: 199 million people communicate on twitter daily, making it essential to study policy and decision-making.
Approach: They propose a flow-based tool to augment Twitter data with additional information about tweets and users.
Outcome: The proposed tool is designed to enhance Twitter data with additional information about tweets and users.
“You Sound Just Like Your Father” Commercial Machine Translation Systems Include Stylistic Biases (2020.acl-main)

Copied to clipboard

Challenge: a recent study shows that machine translations make older and more male characters sound older and older than the original.
Approach: They propose to use demographicallyrepresentative data to examine how text is translated . they show that authors sound older and more male than the original .
Outcome: The results suggest that translation models reflect demographic bias in the training data.
BERTective: Language Models and Contextual Information for Deception Detection (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods to classify texts as truthful or deceptive are limited by the context of the text being analyzed.
Approach: They propose to use a corpus of Italian dialogues to classify texts as truthful or deceptive.
Outcome: The proposed models show that not all contexts are equally useful to the task.
SocioProbe: What, When, and Where Language Models Learn about Sociodemographics (2022.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have outperformed other models on a wide range of tasks . however, there is still little understanding of their knowledge of higher-level aspects of language .
Approach: They investigate whether pre-trained language models have knowledge of sociodemographics . they use traditional probing techniques to probe the knowledge of single-GPU PLMs based on multiple English data sets .
Outcome: The results show that pre-trained language models outperform other models on a wide range of tasks.
HONEST: Measuring Hurtful Sentence Completion in Language Models (2021.naacl-main)

Copied to clipboard

Challenge: 4.3% of the time, language models complete a sentence with a hurtful word . authors propose a score to quantify the amount of hurtful sentence completions in a language model.
Approach: They propose a score to measure hurtful sentence completions in language models . they use a template- and lexicon-based bias evaluation methodology for six languages .
Outcome: The proposed score measures the amount of hurtful sentences in language models.
Cross-lingual Contextualized Topic Models with Zero-shot Learning (2021.eacl-main)

Copied to clipboard

Challenge: Existing topic models are language-specific and cannot be transferred in a transferable manner.
Approach: They propose a zero-shot cross-lingual topic model that learns topics on one language and predicts them for unseen documents in different languages.
Outcome: The proposed model learns topics on one language and predicts them for unseen documents in different languages.
SWEAT: Scoring Polarization of Topics across Different Corpora (2021.emnlp-main)

Copied to clipboard

Challenge: Using two additional wordsets, we compute the relative polarization of a topical wordsetting across two distributional representations.
Approach: They propose a new measure to compute the relative polarization of a topical wordset across two distributional representations using two additional wordsetes deemed to have opposite valence to represent two different poles.
Outcome: The proposed measure is validated by a case study and validated in a randomized controlled trial.
Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence (2021.acl-short)

Copied to clipboard

Challenge: Recent neural topic models extract words from documents, but they are not coherent . coherence is crucial for topic models, but many use bag-of-words document representations as input . pre-trained language models are becoming ubiquitous in natural language processing .
Approach: They combine contextualized representations with neural topic models to produce more coherent topics . they say that future improvements in language models will translate into better topic models .
Outcome: The proposed approach produces more meaningful and coherent topics than bag-of-words models and recent neural models.
Query2Prod2Vec: Grounded Word Embeddings for eCommerce (2021.naacl-industry)

Copied to clipboard

Challenge: Query2Prod2Vec is a model that grounds lexical representations for product search in product embeddings.
Approach: They propose a model that grounds lexical representations for product search in product embeddings.
Outcome: The proposed model is more accurate than existing methods from the literature . it is also more efficient than existing embedding methods in the context of high-traffic websites.
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory (2026.eacl-long)

Copied to clipboard

Challenge: Unlike fine-tuning or static retrieval methods, DC adapts LMs’ problem-solving skills on the fly, without modifying their underlying parameters.
Approach: They propose a lightweight framework that endows a black-box LM with a persistent, evolving memory.
Outcome: The proposed framework enables models to store and reuse accumulated strategies, code snippets, and general problem-solving insights at inference time.
Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction (2021.naacl-main)

Copied to clipboard

Challenge: Scholarly work in this area uses toy worlds and synthetic linguistic data, but grounded language learning offers several practical and scientific advantages.
Approach: They propose to model teacher-learner dynamics through natural interactions occurring between users and search engines.
Outcome: The proposed model is better than non-grounded models on compositionality and zero-shot inference tasks.
“It’s Not Just Hate”: A Multi-Dimensional Perspective on Detecting Harmful Speech Online (2022.emnlp-main)

Copied to clipboard

Challenge: Detecting offensive content is becoming a critical task in natural language processing . but most datasets use a single binary label for hate or incivility, even though each concept is multi-faceted . a more fine-grained multi-label approach addresses conceptual and performance issues .
Approach: They propose to use a dataset to annotate offensive online speech with six labels . they propose to apply a more fine-grained approach to predicting incivility and hateful content .
Outcome: The proposed approach outperforms or matches benchmark datasets on the annotated tweets.
On the Gap between Adoption and Understanding in NLP (2021.findings-acl)

Copied to clipboard

Challenge: a recent paper argues that current publications foster a gap between adoption and understanding of models . it also makes it easier to meet publication demands with method papers, argues the paper .
Approach: They argue that current NLP publication models foster a gap between adoption and understanding of models . they argue that everlarger models make it harder to explain how our methods work .
Outcome: The authors argue that current publications foster a gap between adoption and understanding of models . they argue that the rise of everlarger models makes it harder to explain how our methods work .
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Hate speech datasets focus on English-language content, hindering effective models . annotating hateful content is expensive, time-consuming and potentially harmful to annotators.
Approach: They propose to use ISO 639-1 codes to fine-tune models on one source language and apply them to another language.
Outcome: The proposed approach performs well on some tasks, but fails on many others.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are now being used by millions of people across the world.
Approach: They propose a test suite called XSTest to identify such eXaggerated Safety behaviours in a systematic way.
Outcome: The proposed test suite identifies eXaggerated Safety behaviours in a systematic way.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations