Papers by Dorottya Demszky

14 papers
From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment (2026.findings-acl)

Copied to clipboard

Challenge: a framework for sentence-level interpretability of rubric-based scoring is proposed . aaron e. smith: automated scoring models provide little insight into why scores are produced .
Approach: They propose a framework for sentence-level interpretability of rubric-based scoring that combines Shapley-value attributions with rationales generated by large language models.
Outcome: The proposed framework compares fine-tuned pretrained language models with large language models . it shows that fine- tuned models outperform LLMs in prediction accuracy but exhibit label compression toward mid-scale scores .
Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations (2024.findings-emnlp)

Copied to clipboard

Challenge: POSR is a task of breaking down conversations into segments and linking each segment to the relevant reference item.
Approach: They propose a task that breaks down conversations into segments and links each segment to the relevant reference item.
Outcome: The proposed method outperforms independent segmentation pipelines and large language models on joint metrics.
Learning to Recognize Dialect Features (2021.naacl-main)

Copied to clipboard

Challenge: linguistics do not characterize dialects as simple categories, but as collections of correlated features.
Approach: They propose two multitask learning approaches based on pretrained transformers to detect dialect features in speech and text.
Outcome: The proposed models learn to recognize many features with high accuracy on 22 dialect features of Indian English.
Tell, Don’t Show: Leveraging Language Models’ Abstractive Retellings to Model Literary Themes (2025.findings-acl)

Copied to clipboard

Challenge: Literature challenges traditional bag-of-words approaches for topic modeling because narrative language focuses on immersive sensory details instead of abstractive description or exposition.
Approach: They propose a topic modeling approach that prompts generative language models to *tell* what passages *show*, thereby translating narratives’ surface forms into higher-level concepts and themes.
Outcome: The proposed model can translate narratives’ surface forms into higher-level concepts and themes than by running LDA alone or directly asking LMs to list topics.
GoEmotions: A Dataset of Fine-Grained Emotions (2020.acl-main)

Copied to clipboard

Challenge: Existing datasets for language-based emotion classification are limited and small . existing datasets lack quality annotations for many different emotion categories .
Approach: They propose to use a large manually annotated dataset to study emotion expressions . they conduct transfer learning experiments with existing emotion benchmarks to test their model .
Outcome: The proposed model achieves an average F1-score of .46, leaving room for improvement.
IDEAlign: Comparing Ideas of Large Language Models to Domain Experts (2026.eacl-long)

Copied to clipboard

Challenge: Large language models are increasingly used to produce open-ended, interpretive annotations.
Approach: They propose to use LLM annotations to evaluate content and assess expert similarity . they propose to benchmark different similarity methods against human ratings .
Outcome: The proposed method performs best but falls short of expert alignment . it is useful as a triage filter rather than a substitute for human review.
Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes (2024.naacl-long)

Copied to clipboard

Challenge: Our work explores the potential of large language models (LLMs) to close the novice-expert knowledge gap in remediating math mistakes.
Approach: They propose a method that uses cognitive task analysis to translate an expert’s latent thought process into a decision-making model for remediation.
Outcome: The proposed model can bridge the novice-expert knowledge gap by using cognitive task analysis to translate an expert’s latent thought process into a decision-making model for remediation.
Pártélet: A Hungarian Corpus of Propaganda Texts from the Hungarian Socialist Era (2020.lrec-1)

Copied to clipboard

Challenge: a digitized corpus of Communist propaganda texts is presented in this paper . it represents the direct political agitation and propaganda of the dictatorial system .
Approach: They present a digitized Hungarian corpus of Communist propaganda texts . they use a database to compile a large database of articles from the journal .
Outcome: The proposed dataset provides a unique opportunity for conducting research on Hungarian propaganda discourse . it also provides enables analysis of changes in the political discourse over a 35-year period .
Backtracing: Retrieving the Cause of the Query (2024.findings-eacl)

Copied to clipboard

Challenge: a number of online content portals allow users to ask questions to supplement their understanding.
Approach: They propose a task of backtracing to retrieve the text segment that most likely caused a user query.
Outcome: The proposed method improves on the backtracing task in three domains . the results show that there is room for improvement and new retrieval approaches .
“Mistakes Help Us Grow”: Facilitating and Evaluating Growth Mindset Supportive Language in Classrooms (2023.emnlp-main)

Copied to clipboard

Challenge: GMSL has been shown to significantly reduce disparities in academic achievement and enhance students’ learning outcomes.
Approach: They develop a coaching tool to reframe unsupportive utterances to GMSL using large language models.
Outcome: The proposed model outperforms the GMSL-trained teachers in fostering a growth mindset and promoting challenge-seeking behavior.
EduCoder: An Open-Source Annotation System for Education Transcript Data (2026.acl-demo)

Copied to clipboard

Challenge: Existing annotation tools do not support team-based workflows or access to instructional context .
Approach: They present an open-source web platform for annotating classroom conversation transcripts . they argue that existing annotation tools do not provide access to instructional context .
Outcome: EduCoder is an open-source web platform designed for annotating classroom conversation transcripts.
Measuring Conversational Uptake: A Case Study on Student-Teacher Interactions (2021.acl-long)

Copied to clipboard

Challenge: Despite extensive research showing the positive impact of uptake on student learning and achievement, there is little evidence that it is effective in teaching.
Approach: They propose a framework for computationally measuring uptake by releasing a dataset of student-teacher exchanges extracted from US math classroom transcripts annotated for uptake . they formalize uptake as pointwise Jensen-Shannon Divergence (pJSD) and conduct a linguistically-motivated comparison of different unsupervised measures.
Outcome: The proposed framework outperforms baseline measures in identifying uptake phenomena like question answering and reformulation.
Edu-ConvoKit: An Open-Source Library for Education Conversation Data (2024.naacl-demo)

Copied to clipboard

Challenge: Edu-ConvoKit is an open-source library for analyzing education conversation data.
Approach: They introduce Edu-ConvoKit, an open-source library for conversation data analysis.
Outcome: The open-source library handles pre-processing, annotation and analysis of education conversation data.
Analyzing Polarization in Social Media: Method and Application to Tweets on 21 Mass Shootings (N19-1)

Copied to clipboard

Challenge: a new framework for studying political polarization in social media is needed to understand how group divisions manifest in language.
Approach: They propose to cluster tweet embeddings to uncover four dimensions of political polarization in social media . their results apply existing lexical methods to analyze 4.4M tweets on 21 mass shootings .
Outcome: The proposed framework generates more cohesive topics than traditional models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations