Papers by Valentin Hofmann

14 papers
Dynamic Contextualized Word Embeddings (2021.acl-long)

Copied to clipboard

Challenge: Static word embeddings that represent words by a single vector cannot capture word meaning in different linguistic and extralinguistic contexts.
Approach: They propose dynamic contextualized word embeddings that represent words as a function of linguistic and extralinguistic contexts.
Outcome: The proposed model models time and social space jointly, making them attractive for NLP tasks involving semantic variability.
Modeling Ideological Salience and Framing in Polarized Online Groups with Graph Neural Networks and Structured Sparsity (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods to detect ideological divides in social media rely on knowing in advance the political orientation of text . fascist and mainstream are among the most polarized concepts in reddit in 2019 .
Approach: They propose a minimally supervised method that leverages the network structure of online discussion forums to detect polarized concepts.
Outcome: The proposed framework captures temporal ideological dynamics such as right-wing and left-wing radicalization using graph neural networks and sparsity learning.
Predicting the Growth of Morphological Families from Social and Linguistic Factors (2020.acl-main)

Copied to clipboard

Challenge: a burst in token frequency of the word "trump" in social media before the 2016 presidential election is a prime indicator of topical dynamics.
Approach: They propose a task of Morphological Family Expansion Prediction to predict the size of a morphological family by analyzing a reddit corpus.
Outcome: The proposed task predicts the increase in the size of a morphological family on a reddit corpus.
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance (2026.tacl-1)

Copied to clipboard

Challenge: Large language models are helping millions of users write texts about diverse issues . issue bias is where an LLM tends to present just one perspective on a given issue .
Approach: They construct a set of 2.49m realistic English-language prompts to measure issue bias in LLM writing assistance using 3.9k templates and 212 political issues from real user interactions.
Outcome: The proposed model aligns more with US Democrat than Republican voter opinion on a subset of issues.
Superbizarre Is Not Superb: Derivational Morphology Improves BERT’s Interpretation of Complex Words (2021.acl-long)

Copied to clipboard

Challenge: Pretrained language models (PLMs) are based on fixed-size vocabularies of words and subwords that are generated by compression algorithms such as bytepair encoding.
Approach: They propose to use BERT as an example PLM to study its semantic representations of English derivatives to test their hypothesis.
Outcome: The proposed model outperforms BERT on a series of semantic probing tasks.
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks (2025.acl-long)

Copied to clipboard

Challenge: a study aims to assess the fairness and robustness of Large Language Models in dialectal queries . speakers of "non-standard" dialects are known to experience implicit and explicit discrimination .
Approach: They propose to use a benchmark to assess the fairness of large language models in dialects . they hire speakers with computer science backgrounds to rewrite seven popular benchmarks based on AAVE .
Outcome: The proposed benchmarks show that most models show significant brittleness and unfairness to queries in AAVE.
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race (2025.acl-long)

Copied to clipboard

Challenge: et al., 2012) show value-aligned language models exhibit stereotypes in word association tasks . ignoring racial nuances can perpetuate subtle biases in LMs .
Approach: They propose a bias mitigation strategy that incentivizes representation of racial concepts in early model layers.
Outcome: The proposed approach incentivizes representation of racial concepts in early model layers . it reduces implicit bias by reducing the number of ambiguous inputs, the authors show .
An Embarrassingly Simple Method to Mitigate Undesirable Properties of Pretrained Language Model Tokenizers (2022.acl-short)

Copied to clipboard

Challenge: a standard tokenizer does not cover all characters of a word but preserves key aspects of its morphological structure . a novel method to improve tokenization of pretrained language models is proposed .
Approach: They propose a method to improve the tokenization of pretrained language models . they use the vocabulary of a standard tokenizer but preserves morphological structure .
Outcome: The proposed method improves tokenization of pretrained language models on morphological gold segmentations and text classification tasks.
Large Language Models Discriminate Against Speakers of German Dialects (2025.emnlp-main)

Copied to clipboard

Challenge: In Germany, more than 40% of the population speaks a regional dialect . however, dialect speakers face negative societal stereotypes .
Approach: They construct a corpus that pairs sentences from seven regional German dialects with their standard German counterparts to assess their dialect usage bias.
Outcome: The proposed model reproduces dialect usage bias in association task and decision task.
A Graph Auto-encoder Model of Derivational Morphology (2020.acl-main)

Copied to clipboard

Challenge: Existing words that conform to morphological patterns of a language differ in how likely they are to be actually created by speakers.
Approach: They propose to model the morphological well-formedness of derivatives by combining syntactic and semantic information with associative information from the mental lexicon.
Outcome: The proposed model models the morphological well-formedness of derivatives in English .
The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative (2022.emnlp-main)

Copied to clipboard

Challenge: Construction Grammar posits constructions as the central building blocks of language . human-like performance of pretrained language models on many NLP tasks has been alleged .
Approach: They propose to use construction grammar to posit constructions as the central building blocks of language . they conduct experiments with three pretrained language models to examine their ability to classify and understand English comparative correlative .
Outcome: The proposed models are able to recognise the English comparative correlative (CC) but fail to use its meaning.
Counting the Bugs in ChatGPT’s Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models (LLMs) ignore the remarkable ability of humans to generalize and focus only on English.
Approach: They conduct the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages.
Outcome: The proposed model massively underperforms purpose-built systems, particularly in English.
CaMEL: Case Marker Extraction without Labels (2022.acl-long)

Copied to clipboard

Challenge: Existing models for morphological case marking and semantic content are not isomorphic.
Approach: They propose a model that extracts case markers from a multilingual corpus using a noun phrase chunker and an alignment system.
Outcome: The proposed model can extract case markers in 83 languages and visualise similarities and differences between case systems and annotate fine-grained deep cases in languages where they are not overtly marked.
DagoBERT: Generating Derivational Morphology with a Pretrained Language Model (2020.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models (PLMs) generate derivationally complex words, but it is unclear what they learn about other aspects of language.
Approach: They propose to use BERT to examine its derivational capabilities in different settings, from unmodified pretrained models to full finetuning.
Outcome: The proposed model outperforms the state-of-the-art in derivation generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations