Papers by Dorottya Demszky
From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment (2026.findings-acl)
Copied to clipboard
Ivo Bueno, Babette Bühler, Philipp Stark, Tim Fütterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
| Challenge: | a framework for sentence-level interpretability of rubric-based scoring is proposed . aaron e. smith: automated scoring models provide little insight into why scores are produced . |
| Approach: | They propose a framework for sentence-level interpretability of rubric-based scoring that combines Shapley-value attributions with rationales generated by large language models. |
| Outcome: | The proposed framework compares fine-tuned pretrained language models with large language models . it shows that fine- tuned models outperform LLMs in prediction accuracy but exhibit label compression toward mid-scale scores . |
Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations (2024.findings-emnlp)
Copied to clipboard
| Challenge: | POSR is a task of breaking down conversations into segments and linking each segment to the relevant reference item. |
| Approach: | They propose a task that breaks down conversations into segments and links each segment to the relevant reference item. |
| Outcome: | The proposed method outperforms independent segmentation pipelines and large language models on joint metrics. |
Learning to Recognize Dialect Features (2021.naacl-main)
Copied to clipboard
| Challenge: | linguistics do not characterize dialects as simple categories, but as collections of correlated features. |
| Approach: | They propose two multitask learning approaches based on pretrained transformers to detect dialect features in speech and text. |
| Outcome: | The proposed models learn to recognize many features with high accuracy on 22 dialect features of Indian English. |
Tell, Don’t Show: Leveraging Language Models’ Abstractive Retellings to Model Literary Themes (2025.findings-acl)
Copied to clipboard
| Challenge: | Literature challenges traditional bag-of-words approaches for topic modeling because narrative language focuses on immersive sensory details instead of abstractive description or exposition. |
| Approach: | They propose a topic modeling approach that prompts generative language models to *tell* what passages *show*, thereby translating narratives’ surface forms into higher-level concepts and themes. |
| Outcome: | The proposed model can translate narratives’ surface forms into higher-level concepts and themes than by running LDA alone or directly asking LMs to list topics. |
GoEmotions: A Dataset of Fine-Grained Emotions (2020.acl-main)
Copied to clipboard
| Challenge: | Existing datasets for language-based emotion classification are limited and small . existing datasets lack quality annotations for many different emotion categories . |
| Approach: | They propose to use a large manually annotated dataset to study emotion expressions . they conduct transfer learning experiments with existing emotion benchmarks to test their model . |
| Outcome: | The proposed model achieves an average F1-score of .46, leaving room for improvement. |
IDEAlign: Comparing Ideas of Large Language Models to Domain Experts (2026.eacl-long)
Copied to clipboard
| Challenge: | Large language models are increasingly used to produce open-ended, interpretive annotations. |
| Approach: | They propose to use LLM annotations to evaluate content and assess expert similarity . they propose to benchmark different similarity methods against human ratings . |
| Outcome: | The proposed method performs best but falls short of expert alignment . it is useful as a triage filter rather than a substitute for human review. |
Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes (2024.naacl-long)
Copied to clipboard
| Challenge: | Our work explores the potential of large language models (LLMs) to close the novice-expert knowledge gap in remediating math mistakes. |
| Approach: | They propose a method that uses cognitive task analysis to translate an expert’s latent thought process into a decision-making model for remediation. |
| Outcome: | The proposed model can bridge the novice-expert knowledge gap by using cognitive task analysis to translate an expert’s latent thought process into a decision-making model for remediation. |
Pártélet: A Hungarian Corpus of Propaganda Texts from the Hungarian Socialist Era (2020.lrec-1)
Copied to clipboard
| Challenge: | a digitized corpus of Communist propaganda texts is presented in this paper . it represents the direct political agitation and propaganda of the dictatorial system . |
| Approach: | They present a digitized Hungarian corpus of Communist propaganda texts . they use a database to compile a large database of articles from the journal . |
| Outcome: | The proposed dataset provides a unique opportunity for conducting research on Hungarian propaganda discourse . it also provides enables analysis of changes in the political discourse over a 35-year period . |
Backtracing: Retrieving the Cause of the Query (2024.findings-eacl)
Copied to clipboard
| Challenge: | a number of online content portals allow users to ask questions to supplement their understanding. |
| Approach: | They propose a task of backtracing to retrieve the text segment that most likely caused a user query. |
| Outcome: | The proposed method improves on the backtracing task in three domains . the results show that there is room for improvement and new retrieval approaches . |
“Mistakes Help Us Grow”: Facilitating and Evaluating Growth Mindset Supportive Language in Classrooms (2023.emnlp-main)
Copied to clipboard
| Challenge: | GMSL has been shown to significantly reduce disparities in academic achievement and enhance students’ learning outcomes. |
| Approach: | They develop a coaching tool to reframe unsupportive utterances to GMSL using large language models. |
| Outcome: | The proposed model outperforms the GMSL-trained teachers in fostering a growth mindset and promoting challenge-seeking behavior. |
EduCoder: An Open-Source Annotation System for Education Transcript Data (2026.acl-demo)
Copied to clipboard
Saad Ashraf, Jim Malamut, Vishal Kumar, Guanzhong Pan, HyunJi Nam, Mei Tan, Lucía Langlois, Liliana Carolina Santos-Deonizio, Helen Spencer Higgins, Dorottya Demszky
| Challenge: | Existing annotation tools do not support team-based workflows or access to instructional context . |
| Approach: | They present an open-source web platform for annotating classroom conversation transcripts . they argue that existing annotation tools do not provide access to instructional context . |
| Outcome: | EduCoder is an open-source web platform designed for annotating classroom conversation transcripts. |
Measuring Conversational Uptake: A Case Study on Student-Teacher Interactions (2021.acl-long)
Copied to clipboard
Dorottya Demszky, Jing Liu, Zid Mancenido, Julie Cohen, Heather Hill, Dan Jurafsky, Tatsunori Hashimoto
| Challenge: | Despite extensive research showing the positive impact of uptake on student learning and achievement, there is little evidence that it is effective in teaching. |
| Approach: | They propose a framework for computationally measuring uptake by releasing a dataset of student-teacher exchanges extracted from US math classroom transcripts annotated for uptake . they formalize uptake as pointwise Jensen-Shannon Divergence (pJSD) and conduct a linguistically-motivated comparison of different unsupervised measures. |
| Outcome: | The proposed framework outperforms baseline measures in identifying uptake phenomena like question answering and reformulation. |
Edu-ConvoKit: An Open-Source Library for Education Conversation Data (2024.naacl-demo)
Copied to clipboard
| Challenge: | Edu-ConvoKit is an open-source library for analyzing education conversation data. |
| Approach: | They introduce Edu-ConvoKit, an open-source library for conversation data analysis. |
| Outcome: | The open-source library handles pre-processing, annotation and analysis of education conversation data. |
Analyzing Polarization in Social Media: Method and Application to Tweets on 21 Mass Shootings (N19-1)
Copied to clipboard
| Challenge: | a new framework for studying political polarization in social media is needed to understand how group divisions manifest in language. |
| Approach: | They propose to cluster tweet embeddings to uncover four dimensions of political polarization in social media . their results apply existing lexical methods to analyze 4.4M tweets on 21 mass shootings . |
| Outcome: | The proposed framework generates more cohesive topics than traditional models. |