Papers by Niklas Stoehr
World Models for Math Story Problems (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent efforts to solve math story problems have lacked accurate representations of mathematical concepts. |
| Approach: | They propose a graph-based semantic formalism for solving math story problems . they combine existing datasets and annotate a corpus of 1,019 problems with MathWorld . |
| Outcome: | The proposed model can be used to solve math story problems with pre-trained language models . the model can also be used for generating new problems by using the model as a design space . |
An Ordinal Latent Variable Model of Conflict Intensity (2023.acl-long)
Copied to clipboard
| Challenge: | Advances in automated event extraction yield massive data sets of “who did what to whom” micro-records that enable data-driven approaches to monitoring conflict. |
| Approach: | They propose a probabilistic generative model that assumes each observed event is associated with a latent intensity class. |
| Outcome: | The proposed model obtains comparatively good held-out predictive performance on a conflictual to cooperative scale. |
Activation Scaling for Steering and Interpreting Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a successful intervention should flip the correct with the wrong token, while remaining sparse. |
| Approach: | They propose to use activation scaling to flip the correct with the wrong token . they use gradient-based optimization to learn and evaluate a specific kind of efficient intervention . |
| Outcome: | The proposed method performs comparable with steering vectors but is much less minimal. |
Context versus Prior Knowledge in Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies have investigated how often a model will rely on prior knowledge over conflicting contextual information in answering questions. |
| Approach: | They propose two mutual information-based metrics to measure a model’s dependency on a context and on its prior about an entity. |
| Outcome: | The proposed metrics show that language models can integrate prior knowledge and new information in a predictable way across different questions and contexts. |
Classifying Dyads for Militarized Conflict Analysis (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing research examines the origins of militarized conflict by examining bi-lateral relationships between entity pairs and multi-lateral relations among multiple entities. |
| Approach: | They propose to use Wikipedia to model dyadic and systemic causes to compare their correlations with conflict between two entities. |
| Outcome: | The proposed graphs show that Wikipedia articles of allies are semantically more similar than enemies. |
Unsupervised Contrast-Consistent Ranking with Language Models (2024.eacl-long)
Copied to clipboard
| Challenge: | Language models contain ranking-based knowledge and are powerful solvers of in-context ranking tasks. |
| Approach: | They propose to use a model to elicit language models' ranking knowledge without supervision by using a pairwise, pointwise and listwise prompting method. |
| Outcome: | The proposed method is inspired by an unsupervised probing method called Contrast-Consistent Search (CCS). |
The Architectural Bottleneck Principle (2022.emnlp-main)
Copied to clipboard
| Challenge: | a recent study examined how much information a model's representations contain . a new approach to probing is to look exactly like the component . |
| Approach: | They propose a new probing principle that aims to estimate how much information a model could extract from its representations. |
| Outcome: | The proposed probes extract syntactic information from the representations of a neural network . the proposed probe is based on the architectural bottleneck principle . |
UniMorph 4.0: Universal Morphology (2022.lrec-1)
Copied to clipboard
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
| Challenge: | The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages. |
| Approach: | They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. |
| Outcome: | The proposed schema has added 66 new languages, including 24 endangered languages. |
Sentiment as an Ordinal Latent Variable (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing dictionaries are limited in coverage and sentiment scales vary widely; some are discrete others continuous. |
| Approach: | They propose a Bayesian generative model that learns a composite sentiment dictionary as an interpolation between six existing dictionaries with different scales. |
| Outcome: | The proposed model learns a composite sentiment dictionary as an interpolation between six existing dictionaries with different scales. |
What About the Precedent: An Information-Theoretic Analysis of Common Law (2021.naacl-main)
Copied to clipboard
| Challenge: | In common law, the outcome of a new case is determined mostly by precedent cases, rather than by existing statutes. |
| Approach: | They propose to model the argumentation of precedent cases and compare them to a case out-come classification task to determine how the precedent influences the outcome of a new case. |
| Outcome: | The proposed method compared arguments of two longstanding jurisprudential views on the European Court of Human Rights (ECtHR) and the precedent cases. |
Generalizing Backpropagation for Gradient-Based Interpretability (2023.acl-long)
Copied to clipboard
| Challenge: | Several feature-attribution methods for interpreting deep neural networks rely on computing the gradients of a model’s output with respect to its inputs, but they reveal little about the inner workings of the model itself. |
| Approach: | They propose a generalized backpropagation algorithm that generalizes the gradient computation of a model to efficiently compute other interpretable statistics about the gradient graph of neural networks. |
| Outcome: | The proposed generalized algorithm can be used to compute other interpretable statistics about the gradient graph of a neural network, such as the highest-weighted path and entropy. |
Extracting Victim Counts from Text (2023.eacl-main)
Copied to clipboard
| Challenge: | Using tagging and regex methods, data on injured, displaced, or abused victims is difficult . data on earthquake injuries and deaths is scarce, subjective, or biased . |
| Approach: | They compare tagging approaches to extract injured, displaced, or abused victims . they discuss calibration and investigate out-of-distribution and few-shot performance . |
| Outcome: | The proposed model is among the first to apply numeracy-focused large language models in a real-world use case with a positive impact. |
Measuring scalar constructs in social science with LLMs (2025.emnlp-main)
Copied to clipboard
Hauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel, Niklas Stoehr, Elliott Ash, Alexander Miserlis Hoyle
| Challenge: | Valid scalar measurement of skalar constructs is a fundamental task in text analysis. |
| Approach: | They evaluate four approaches to measuring scalar constructs using large language models . pairwise comparisons produced better measurements than prompting LLMs, they say . validation of skalar measurement enables wide range of substantive applications in social science research . |
| Outcome: | The proposed methods improve on pairwise comparisons and finetuning . the proposed methods can be used in social science research . |