Papers with education

104 papers
EdTec-QBuilder: A Semantic Retrieval Tool for Assembling Vocational Training Exams in German Language (2024.naacl-demo)

Copied to clipboard

Challenge: Existing methods to gather test items from validated item databases are under-researched, but there is little research on assembling exam items from a database of valid items.
Approach: They propose to use semantic search to assist vocational educators in assembling exam forms by using eight retrieval strategies and 25 popular sentence similarity models.
Outcome: The proposed tool is based on eight retrieval strategies and 25 popular pre-trained sentence similarity models.
LLMs for Low Resource Languages in Multilingual, Multimodal and Dialectal Settings (2024.eacl-tutorials)

Copied to clipboard

Challenge: Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting .
Approach: They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models .
Outcome: The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects.
Automatic Classification of Students on Twitter Using Simple Profile Information (2020.aacl-srw)

Copied to clipboard

Challenge: Existing models for age classification of students and non-students are restrictive and require access to many tweets.
Approach: They propose a model which uses 3 tweet-content features to classify users as students or non-students.
Outcome: The proposed model achieves an accuracy of 88.1% and an F1 score of .704 compared to previous models, which require access to many user tweets.
Edu-ConvoKit: An Open-Source Library for Education Conversation Data (2024.naacl-demo)

Copied to clipboard

Challenge: Edu-ConvoKit is an open-source library for analyzing education conversation data.
Approach: They introduce Edu-ConvoKit, an open-source library for conversation data analysis.
Outcome: The open-source library handles pre-processing, annotation and analysis of education conversation data.
ConQuer: A Framework for Concept-Based Quiz Generation (2025.naacl-srw)

Copied to clipboard

Challenge: Existing concepts-based quiz generation frameworks that leverage external knowledge sources are challenging and labor intensive.
Approach: They propose a concept-based quiz generation framework that leverages external knowledge sources to assess the quality of the generated quizzes, using LLMs as judges.
Outcome: The proposed framework shows a 4.8% improvement in evaluation scores and a 77.52% win rate in pairwise comparisons against baseline quiz sets.
Socratic Question Generation: A Novel Dataset, Models, and Evaluation (2023.eacl-main)

Copied to clipboard

Challenge: Socratic questioning is a form of reflective inquiry often employed in education to encourage critical thinking in students.
Approach: They present a first large dataset of 110K questions, context pairs for Socratic Question Generation.
Outcome: The proposed model produces realistic, type-sensitive, human-like Socratic questions . authors show that the model can be used in counseling and coaching .
ADVISER: A Dialog System Framework for Education & Research (P19-3)

Copied to clipboard

Challenge: In this paper, we focus on task-oriented dialog systems, although our framework allows easy integration of non-task dialog systems and their combination.
Approach: They propose an open source dialog system framework for education and research that supports multi-domain task-oriented conversations in two languages.
Outcome: The proposed framework supports multi-domain task-oriented conversations in two languages and is open source for education and research.
Mitigating Bias in Item Retrieval for Enhancing Exam Assembly in Vocational Education Services (2025.naacl-industry)

Copied to clipboard

Challenge: Despite the practical importance of exam assembly, few methods exist to support educators during manual item retrieval for exam assembly tasks.
Approach: They propose a mixed-integer programming re-ranking approach to improve relevance while mitigating bias on an industry-grade exam assembly platform.
Outcome: The proposed approach improves relevance and reduces bias by 17% when compared to other methods on a real-world exam assembly platform.
Introducing NIEUW: Novel Incentives and Workflows for Eliciting Linguistic Data (L18-1)

Copied to clipboard

Challenge: a 2010 survey found that the language of the European Union, not even English, was not fully supplied . the absence of Language Resources stifles teaching and technology building, authors say .
Approach: They propose to harness the power of alternative incentives to elicit linguistic data and annotation . they also describe changes to the workflows necessary to collect data from workforces attracted by incentives .
Outcome: a new initiative to harness incentives to elicit linguistic data and annotation improves language resources . the NIEUW project is funded by the u.s. national science foundation .
A Data-Centric Framework for Composable NLP Workflows (2020.emnlp-demos)

Copied to clipboard

Challenge: Empirical natural language processing (NLP) systems involve interoperation among multiple components . a wealth of NLP toolkits exist ( 4), such as spaCy, DKPro, CoreNLP.
Approach: They propose a unified open-source framework that supports fast development of NLP workflows . framework includes processors for NLP tasks, visualization, and annotation .
Outcome: The framework offers processors for NLP tasks, visualization, and annotation, and is extensible . it is delivered through two modularized yet integratable open-source projects, Forte and Stave .
Word Affect Intensities (L18-1)

Copied to clipboard

Challenge: Existing lexicons of affect only show coarse associations, but are not accurate as human-created ones.
Approach: They propose to use a manually created affect intensity lexicon with real-valued intensity scores for anger, fear, joy, and sadness.
Outcome: The lexicon has real-valued scores for anger, fear, joy, and sadness . anger, fears, and sad words have very similar VAD scores .
Using Word Embeddings to Analyze Teacher Evaluations: An Application to a Filipino Education Non-Profit Organization (2021.findings-acl)

Copied to clipboard

Challenge: Using natural language processing, we examine teacher evaluations to improve educational programs.
Approach: They apply natural language processing techniques to a Filipino education non-profit to analyze teacher evaluations written by Teacher Fellows.
Outcome: The proposed framework can be applied to teacher evaluations from a Filipino education non-profit.
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection (2024.emnlp-demo)

Copied to clipboard

Challenge: a large number of machine-generated texts are often hard to distinguish between human-written and machine-generated text . this raises concerns about potential misuse, especially within educational and academic domains .
Approach: They propose a system that can detect whether a text is human-written or machine-generated . they use a fine-grained classification schema to identify the use of machine-generated text .
Outcome: The proposed system can distinguish between human-written and machine-generated text . it can detect attempts to obfuscate the fact that a text was machine- generated .
TestAgent: An Adaptive and Intelligent Expert for Human Assessment (2025.findings-acl)

Copied to clipboard

Challenge: Existing adaptive testing methods face several challenges due to mechanized nature of most algorithms and noisy response data.
Approach: They propose to use large language models to enhance adaptive testing through interactive engagement to capture test-takers’ responses and anomalies.
Outcome: The proposed agent achieves more accurate results with 20% fewer questions than state-of-the-art baselines and testers preferred it in speed, smoothness, and other dimensions.
Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai (2025.acl-industry)

Copied to clipboard

Challenge: Text-to-speech (TTS) systems are limited by limited data and linguistic complexities.
Approach: They propose a data-optimized framework with an advanced acoustic model to build high-quality TTS systems for low-resource scenarios.
Outcome: The proposed framework enables zero-shot voice cloning and improved performance across diverse client applications, including finance, healthcare, education, and law.
Foundations of PEERS: Assessing LLM Role Performance in Educational Simulations (2025.acl-srw)

Copied to clipboard

Challenge: In education, peer instruction is widely recognized as an effective active learning strategy, but evaluations of PI are limited by logistical constraints and variability in classroom settings.
Approach: They propose a simulation framework that integrates Agent-Based Modeling, Large Language Models, and Bayesian Knowledge Tracing to emulate student learning dynamics.
Outcome: The proposed framework integrates Agent-Based Modeling, Large Language Models, and Bayesian Knowledge Tracing to emulate student learning dynamics in real classrooms.
Attention to Non-Adopters (2026.findings-acl)

Copied to clipboard

Challenge: incorporating non-adopter perspectives is essential for developing useful and capable LLMs, argues a new study.
Approach: They argue that incorporating non-adopter perspectives is essential for developing broadly useful and capable LLMs.
Outcome: The proposed method will risk missing tasks prioritized by non-adopters, the authors argue . they show that non-dots diverge from those of current users, and non-no-acopter needs point towards novel reasoning tasks.
UniEDU: Toward Unified and Efficient Large Multimodal Models for Educational Tasks (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing research has focused on plain text, while real-world K-12 scenarios often involve multimodal data.
Approach: They propose a unified language and vision assistant called UniEDU for educational applications . it excels across multiple educational tasks while maintaining strong generalization capabilities . authors propose to use UniEDu for industry-scale deployment .
Outcome: The proposed model excels across multiple educational tasks while maintaining strong generalization capabilities.
LLMs on a Budget? Say HOLA (2025.emnlp-industry)

Copied to clipboard

Challenge: Current solutions such as quantization, pruning, and Retrieval-Augmented Generation (RAG) offer only partial optimizations and often sacrifice accuracy, speed, or generality.
Approach: They propose an end-to-end optimization framework for efficient LLM deployment . it leverages Hierarchical Speculative Decoding (HSD) for faster inference without quality loss.
Outcome: HOLA delivers +17.6% EMA on GSM8K, +10.5% MCA on ARC, and reduced latency and memory on edge devices like Jetson Nano.
Persona Jailbreaking in Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Existing studies focus on narrative or role-playing tasks and overlook how adversarial conversational history alone can reshape induced personas.
Approach: They propose a framework that embeds semantically loaded cues into user queries to gradually induce reverse personas.
Outcome: The proposed framework predictably shifts personas, triggers collateral changes in correlated traits, and exhibits stronger effects in multi-turn settings.
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have demonstrated the potential of large language models (LLMs) for automatic error detection in math word problems (MWPs).
Approach: They propose a framework that generates adaptive reference solutions using LLMs to enhance error detection by reducing conformity bias in MWPs.
Outcome: The proposed framework mitigates the performance gap between conventional and alternative solutions in MWPs, especially when combined with reasoning-enhancing techniques like chain-of-thought prompting.
M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection (2024.eacl-long)

Copied to clipboard

Challenge: Large language models generate fluent responses to user queries, but they are also susceptible to misuse in journalism, education, and academia.
Approach: They propose a large-scale benchmark for machine-generated text detection that is a multi-generator, multi-domain, and multi-lingual corpus.
Outcome: The proposed system can detect machine-generated text and pinpoint misuse . the proposed system is based on a large-scale benchmark dataset .
Geographic Citation Gaps in NLP Research (2022.emnlp-main)

Copied to clipboard

Challenge: a vast number of papers accepted at top NLP venues come from a handful of western countries and (lately) China.
Approach: They ask researchers to examine the relationship between geographical location and publication success . they use a dataset of 70,000 papers from the ACL Anthology to examine their citation network .
Outcome: The proposed dataset of 70,000 papers from the ACL Anthology shows that there are substantial geographical disparities in paper acceptance and citations .
From Information to Insight: Leveraging LLMs for Open Aspect-Based Educational Summarization (2025.acl-long)

Copied to clipboard

Challenge: a novel dataset summarizes student reflections on STEM lectures . ReflectASP eases the exploration of open-aspect-based summarization (OABS) despite the limitations of current datasets, it is still under-explored.
Approach: They propose a dataset that summarizes student reflections on STEM lectures . they propose two refinement methods to improve summaries .
Outcome: The proposed dataset summarizes student reflections on STEM lectures using automatic and human evaluations.
Towards Generating Controllable and Solvable Geometry Problem by Leveraging Symbolic Deduction Engine (2025.acl-industry)

Copied to clipboard

Challenge: Compared to math word problems, geometry problems emphasize multi-modal formats and the translation between informal and formal languages.
Approach: They propose a symbolic deduction engine-based geometry problem generation framework that leverages a symbolic deduction engine to generate geometry problems.
Outcome: The proposed method avoids inherent biases in translating natural language into formal language and guarantees to control the generated problems in terms of knowledge points and difficulties by an elaborate checking function.
Contextual Diversity Measure (CDM) for Controllable Story Generation in Large Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies on controllable text generation focus on controlling attributes such as sentiment, writing style, and writing style.
Approach: They introduce a metric that quantifies semantic diversity for scenario generation under fixed abstract semantic constraints and validate it through controlled experiments.
Outcome: The proposed metric achieves excellent discrimination accuracy (100% and 91.9%, respectively), with discriminative power up to 5.5 greater than the best baseline.
Gaining and Losing Influence in Online Conversation (L18-1)

Copied to clipboard

Challenge: a study aimed to determine if people who are influential in online discussions retain influence when placed in a topic that is less familiar or perhaps not as interesting.
Approach: They conducted a study to determine if people who are highly influential retain influence when moving to a topic that is less familiar or perhaps not as interesting.
Outcome: The results show that people who are highly influential in group discussions lose influence when placed in a topic that is less familiar or perhaps not as interesting.
Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes (2024.naacl-long)

Copied to clipboard

Challenge: Our work explores the potential of large language models (LLMs) to close the novice-expert knowledge gap in remediating math mistakes.
Approach: They propose a method that uses cognitive task analysis to translate an expert’s latent thought process into a decision-making model for remediation.
Outcome: The proposed model can bridge the novice-expert knowledge gap by using cognitive task analysis to translate an expert’s latent thought process into a decision-making model for remediation.
Measuring Conversational Uptake: A Case Study on Student-Teacher Interactions (2021.acl-long)

Copied to clipboard

Challenge: Despite extensive research showing the positive impact of uptake on student learning and achievement, there is little evidence that it is effective in teaching.
Approach: They propose a framework for computationally measuring uptake by releasing a dataset of student-teacher exchanges extracted from US math classroom transcripts annotated for uptake . they formalize uptake as pointwise Jensen-Shannon Divergence (pJSD) and conduct a linguistically-motivated comparison of different unsupervised measures.
Outcome: The proposed framework outperforms baseline measures in identifying uptake phenomena like question answering and reformulation.
MathEDU: Feedback Generation on Problem-Solving Processes for Mathematical Learning Support (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies have examined the reliability of Large Language Models (LLMs) in grading authentic student problem solving processes and delivering effective feedback.
Approach: They propose to use a dataset to evaluate the reliability of large language models in mathematics and a teacher-written feedback system to improve student problem-solving processes.
Outcome: The proposed model improves in correctness classification, error identification, and feedback generation, but generates a gap from teacher-written feedback.
Cross Encoding as Augmentation: Towards Effective Educational Text Classification (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to improve text classification in education suffer from data scarcity . authors propose a retrieval approach that provides effective learning in educational text classification.
Approach: They propose a retrieval approach that provides effective learning in educational text classification by introducing cross-encoder style texts to a bi-encoding architecture.
Outcome: The proposed method is effective in multi-label scenarios and low-resource tags compared to state-of-the-art models.
Safe-Unsafe Concept Separation Emerges from a Single Direction in Language Models Activation Space (2026.eacl-long)

Copied to clipboard

Challenge: Existing approaches to ensuring the safety of Large Language Models (LLMs) rely on invasive fine- tuning or external generation-based checks, which can be opaque and resource-inefficient.
Approach: They propose a mechanistic method that identifies the layer where safe and unsafe concepts are maximally separable within a pretrained representation space.
Outcome: The proposed method can be used across multiple domains, diverse tasks, and 16 non-English languages on encoder and decoder architectures.
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations (2025.emnlp-main)

Copied to clipboard

Challenge: a new metric measures the quality of large language models (LLMs) that detects hidden misalignments and jailbreak risks.
Approach: They propose a decoding-invariant metric that measures latent safety failures . they propose 'Alignment Quality Index' to measure latent activations in latent space .
Outcome: The proposed metric detects latent safety failures overlooked by behavioral benchmarks and jailbreaks.
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs).
Approach: They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs.
Outcome: The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints.
Message Passing on Semantic-Anchor-Graphs for Fine-grained Emotion Representation Learning and Classification (2024.emnlp-main)

Copied to clipboard

Challenge: Emotion classification is an important task with applications in education, virtual reality, and robotics.
Approach: They propose to use token embeddings to generate a "semantic-anchor graph" using semantic anchors, sentences can be projected onto them to form a graph .
Outcome: Empirically, the proposed system can generate meaningful semantic anchors and discriminative graph patterns for different emotion.
Personalized Multimodal Feedback Generation in Education (2020.coling-main)

Copied to clipboard

Challenge: In this paper, we propose a novel Personalized Multimodal Feedback Generation Network (PMFGN) that generates personalized feedback for teachers to evaluate assignments involving multimodal inputs.
Approach: They propose a Personalized Multimodal Feedback Generation Network (PMFGN) that generates personalized feedback for teachers to evaluate assignments involving multimodal inputs such as images, audios, and texts.
Outcome: The proposed model outperforms baseline models on real-world K-12 education data and detailed ablation experiments to deepen understanding of the proposed framework.
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above (2025.acl-long)

Copied to clipboard

Challenge: Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing.
Approach: They argue for a reform of multiple choice question answering (MCQA) they argue for more generative formats based on human testing .
Outcome: The proposed reforms improve the quality of MCQA, the authors argue . they show that even when MCQ is a useful format, its datasets suffer from leakage, unanswerability, shortcuts and saturation.
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Recognizing LLMs’ capability to generate educational content can lead to advances in automated and personalized learning.
Approach: They propose to evaluate the questioning capability in education as a teacher of large language models by evaluating their generated educational questions.
Outcome: The proposed model can generate educational content that aligns with human perspectives and is more apt as an interdisciplinary teacher.
(Male, Bachelor) and (Female, Ph.D) have different connotations: Parallelly Annotated Stylistic Language Dataset with Multiple Personas (D19-1)

Copied to clipboard

Challenge: Despite recent attempts on computational modeling of the variation, the lack of parallel corpora of style language makes it difficult to systematically control the stylistic change and evaluate such models.
Approach: They propose to use a parallel and annotated stylistic language dataset to test the effectiveness of style transfer models.
Outcome: The proposed model outperforms the unsupervised models using nonparallel corpus.
On the Computational Modelling of Michif Verbal Morphology (2021.eacl-main)

Copied to clipboard

Challenge: Existing computational models of the verbal morphology of the Métis language are insufficient to model the language's unique phonological interactions.
Approach: They propose a finite-state computational model of the verbal morphology of Michif . they use composed finite state transducers to model concatenative morphologies .
Outcome: The proposed model is based on a series of finite-state transducers.
ExpressivityBench: Can LLMs Communicate Implicitly? (2026.findings-eacl)

Copied to clipboard

Challenge: a new study evaluates the expressivity of large language models for communicating implicitly . authors: models can express tone, identity, and intent beyond literal meanings . phrasing and tone of a message can convey a number of topics beyond literal contexts - authors .
Approach: They propose a framework to evaluate the expressivity of large language models . they use a social-linguistic grader to validate their models against human judgments .
Outcome: The proposed framework quantifies how well LLM-generated text communicates target properties without explicit mention across nine tasks spanning emotion, identity, and tone.
Persistent Personas? Role-Playing, Instruction Following, and Safety in Extended Interactions (2026.eacl-long)

Copied to clipboard

Challenge: Persona-assigned large language models are used in education, healthcare and sociodemographic simulations.
Approach: They propose a protocol that combines long persona dialogues and evaluation datasets to create dialogue-conditioned benchmarks that can robustly measure long-context effects.
Outcome: The proposed protocol can measure persona fidelity, instruction-following, and safety in long conversations.
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech (2024.naacl-long)

Copied to clipboard

Challenge: Automatic speech recognition systems fail to accurately interpret speech patterns deviating from typical fluency, leading to critical usability issues and misinterpretations.
Approach: They evaluate six leading automatic speech recognition systems based on a real-world dataset and a synthetic dataset derived from the widely-used LibriSpeech benchmark.
Outcome: The six leading speech recognition systems were evaluated on a real-world dataset and a synthetic dataset derived from the widely-used LibriSpeech benchmark.
Automatic Generation of Socratic Subquestions for Teaching Math Word Problems (2022.emnlp-main)

Copied to clipboard

Challenge: We hypothesize that questioning can enhance human performance and assist solvers .
Approach: They propose to use large language models to generate sequential questions for math word problem-solving . they propose to apply these models to a variety of math word problems .
Outcome: The proposed model improves the performance of a math word problem solver by generating more questions than other models.
SentNoB: A Dataset for Analysing Sentiment on Noisy Bangla Texts (2021.findings-emnlp)

Copied to clipboard

Challenge: Bangla is the sixth most spoken language worldwide and the second Indo-Aryan language after Hindi.
Approach: They propose an annotated sentiment analysis dataset made of informally written Bangla texts.
Outcome: The proposed dataset is compared with neural networks and pretrained models . it shows that hand-crafted lexical features provide superior performance than neural networks .
Now You Hear Me: Audio Narrative Attacks Against Large Audio–Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing jailbreaks against large audio-language models fall into two categories . early work converted text-based prompts into synthetic speech, while subsequent work introduced minor acoustic variations such as accent shifts, phonetic spellings, or stress patterns.
Approach: They propose a text-to-audio jailbreak that embeds disallowed directives within a narrative-style audio stream.
Outcome: The proposed attack exploits structural and acoustic properties of a text-to-audio model . it achieves 98.26% success rate, significantly exceeding baselines for text-based models .
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that relying on LLMs as information providers may hurt student learning.
Approach: They introduce and apply two bias score metrics to evaluate LLMs for bias in the personalized educational setting, specifically on the models’ roles as “teachers.”
Outcome: The proposed models harm student learning by perpetuating harmful stereotypes and reversing them.
Planning-Guided Tutoring with Assessment-Driven Memory for Pedagogical LLM Tutors (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to simulate tutor behaviors or preferences fail to sustain high-quality pedagogical conversations that provide explicit stepwise scaffolding and adapt to learners’ evolving cognitive states.
Approach: They propose a planning-guided tutoring framework with an assessment-driven memory for multi-turn math dialogue tutoring.
Outcome: Experiments on multi-turn math tutoring benchmarks show that ScaffoldLM significantly improves pedagogical tutoring quality over strong baselines.
TMATH A Dataset for Evaluating Large Language Models in Generating Educational Hints for Math Word Problems (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being applied in education, showing significant potential in personalized instruction, student feedback, and intelligent tutoring systems (ITSs).
Approach: They propose a dataset specifically designed to evaluate LLMs’ ability to generate high-quality hints for Math Word Problems.
Outcome: The proposed dataset shows that LLMs can generate more accurate and contextually appropriate educational hints for math word problems without offering direct answers.
Diagnosing Vision Language Models’ Perception by Leveraging Human Methods for Color Vision Deficiencies (2026.eacl-long)

Copied to clipboard

Challenge: Large-scale Vision-Language Models (LVLMs) are being deployed in real-world settings that require visual inference.
Approach: They evaluate LVLMs' ability to account for variation in color perception using the Ishihara Test.
Outcome: The proposed models fail to reproduce the perceptual outcomes experienced by affected individuals and default to normative color perception.
Using Discourse Information for Education with a Spanish-Chinese Parallel Corpus (L18-1)

Copied to clipboard

Challenge: Discourse information is crucial for many NLP tasks due to the great distance that spans between the two languages.
Approach: They propose to use a Spanish-Chinese parallel corpus with annotated discourse information to serve for bilingual language education.
Outcome: The proposed corpus is composed of 100 Spanish-Chinese parallel texts, and all the discourse markers (DM) have been annotated to form the education source.
CPLM, a Parallel Corpus for Mexican Languages: Development and Interface (2020.lrec-1)

Copied to clipboard

Challenge: Mexico has 68 linguistic groups and 364 varieties, but lack of data on social media and internet is putting them at risk.
Approach: They propose a collaborative corpus for endangered languages in Mexico . they propose linguistic search, digitalization and alignment process for each language .
Outcome: The proposed corpus aligns Spanish with six indigenous languages: Maya, Ch’ol, Mazatec, Mixtec, Otomi, and Nahuatl.
Beyond the Textual: Generating Coherent Visual Options for MCQs (2025.findings-emnlp)

Copied to clipboard

Challenge: Multiple choice questions (MCQs) are crucial for deep thinking and knowledge integration in education.
Approach: They propose a cross-modal options synthesis framework for generating MCQs with visual options.
Outcome: The proposed framework produces a plausible and visually similar answer and distractor . it also includes a discrimination module to identify content suitable for visual options .
PersonaGym: Evaluating Persona Agents and LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Persona agents are LLM agents conditioned to act according to an assigned persona . evaluating how faithfully these agents adhere to their personas remains a challenge .
Approach: a new study evaluates persona agents' ability to act according to an assigned persona . a persona agent's person score is a human-aligned automatic metric that can be used to evaluate a model .
Outcome: a new evaluation framework and a human-aligned automatic metric show that persona agents can perform better.
Sociolectal Analysis of Pretrained Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Using data from English cloze tests, we demonstrate wide performance gaps across demographic groups and show that pretrained language models disfavor young non-white male speakers.
Approach: They use data from English cloze tests to examine performance differences of pretrained language models across demographic groups.
Outcome: The models disfavor young non-white male speakers, but larger models reduce performance gaps between majority and minority groups.
Systematic Inequalities in Language Technology Performance across the World’s Languages (2022.acl-long)

Copied to clipboard

Challenge: Recent studies have revealed that NLP is limited to a subset of the world’s 6,500 languages.
Approach: They propose a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Outcome: The proposed framework estimates the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Interpretable Math Word Problem Solution Generation via Step-by-step Planning (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to solving math word problems focus on obtaining the correct answer.
Approach: They propose a step-by-step planning approach for intermediate solution generation that strategically plans the generation of the next solution step based on the MWP and the previous solution steps.
Outcome: The proposed approach improves the accuracy and interpretability of the solution on automatic metrics and human evaluation.
VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions (2025.acl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) can grasp the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent.
Approach: They propose a framework for multimodal large language models to grasp the intention of a question and decompose it into a series of visual recognition sub-tasks to find out the answer.
Outcome: The proposed framework improves the accuracy of complex video-related questions by 29.6% and 17.2% on CVQA and the existing VQA datasets.
AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models are increasingly used in Personalized Image Aesthetic Assessment (PIAA) however, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education.
Approach: They propose to evaluate MLLMs along two complementary dimensions: (1) stereotype bias and (2) alignment between model outputs and genuine human aesthetic preferences.
Outcome: The proposed benchmark covers three subtasks: aesthetic perception, assessment, empathy and alignment between outputs and genuine human aesthetic preferences.
Distilling ChatGPT for Explainable Automated Student Answer Assessment (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing automated student answer assessment models lack explainable and faithful feedback.
Approach: They propose a framework that leverages ChatGPT for student answer scoring and rationale generation.
Outcome: The proposed method improves the overall QWK score by 11% compared to ChatGPT.
MuPe Life Stories Dataset: Spontaneous Speech in Brazilian Portuguese with a Case Study Evaluation on ASR Bias against Speakers Groups and Topic Modeling (2025.coling-main)

Copied to clipboard

Challenge: Recent datasets for automatic speech recognition in Brazilian Portuguese lack diversity in terms of age groups, regional accents, and education levels.
Approach: They propose to use a dataset to analyze the impact of ASR in Brazilian Portuguese (BP) they demonstrate that current models are biased regarding age, education, and regional accents.
Outcome: The proposed dataset helps mitigate biases in current ASR models regarding education levels and age groups.
CEAMC: Corpus and Empirical Study of Argument Analysis in Education via LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing argument component classifications in education are simplistic and isolated, failing to capture the complete argument information.
Approach: They propose to annotate a manually annotated argument component classification dataset from authentic examination settings and to explore the performance of Large Language Models on CEAMC.
Outcome: The proposed dataset can be used to analyze argumentative essays in education.
Analysis of State-Level Legislative Process in Enhanced Linguistic and Nationwide Network Contexts (2024.naacl-long)

Copied to clipboard

Challenge: a new framework for understanding state-level legislative process improves understanding of state legislation and its implications.
Approach: They propose to use generative large language models to decode legislators' behavior and implications of state policies by establishing a shared nationwide network.
Outcome: The framework decodes legislators’ behavior and implications of state policies by establishing a shared nationwide network enriched with diverse contexts, such as information on interest groups influencing public policy and legislators' courage test results, which reflect their political positions.
Chinese Court Simulation with LLM-Based Agents System (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies have neglected the systematic design and procedure evaluation of court simulations, which are critical to the credibility and usage of court simulators in practice.
Approach: They propose a court simulation paradigm based on the real-world procedure structure of Chinese courts and a framework that focuses on both legal judgment prediction and court procedure analysis.
Outcome: The proposed model outperforms judges and lawyers from the real trials in many aspects.
E-EVAL: A Comprehensive Chinese K-12 Education Evaluation Benchmark for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: despite the rapid development of Large Language Models, there is no dedicated benchmark for evaluating LLMs in Chinese K-12 education.
Approach: They propose to develop a benchmark specifically tailored for Chinese K-12 education.
Outcome: EVAL is the first evaluation benchmark specifically tailored for Chinese K-12 education.
EduMARS: Can Vision-Language Models Grade Like Teachers? Benchmarking Multimodal, Rubric-Based Assessment on Chinese K-12 Answers (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for automated grading of student work fail to evaluate real student responses . existing models fail to assess real student work, especially on cognitively demanding tasks .
Approach: They propose a multimodal benchmark for rubric-aligned evaluation of real Chinese K-12 student answers.
Outcome: The proposed model improves performance and interpretability of existing models on EduMARS . existing models fail to perform on real-world, cognitively demanding tasks, authors say .
StudentEval: A Benchmark of Student-Written Prompts for Large Language Models of Code (2024.findings-acl)

Copied to clipboard

Challenge: Existing CodeLLM benchmarks rely on a single expert-written prompt per problem . a growing body of work shows their utility to professional programmers .
Approach: They propose a natural-language-to-code benchmark of prompts written by non-experts . student prompts are written by 80 students who have only completed one introductory Python course .
Outcome: The proposed model is better discriminator of student prompt descriptions than existing benchmarks.
Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus (2025.findings-emnlp)

Copied to clipboard

Challenge: linguistic diversity of India poses significant machine translation challenges, authors say . underrepresented tribal languages like Bhili lack high-quality linguistic resources .
Approach: They introduce a Bhili-Hindi-English Parallel Corpus, the first and largest parallel corpus worldwide . they evaluated a wide range of proprietary and open-source MLLMs on bidirectional translation tasks .
Outcome: The proposed corpus spans critical domains such as education, administration, and news.
CEFR-based Lexical Simplification Dataset (L18-1)

Copied to clipboard

Challenge: Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective.
Approach: They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile .
Outcome: The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus.
Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants’ Question-Answering in Asynchronous Learning Environments (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing assessments rely on surface-level metrics and lack sufficient grounding in educational theory . a new framework is proposed to evaluate VTAs in asynchronous learning environments .
Approach: They propose a pedagogically-oriented evaluation framework tailored to asynchronous forum discussions . they construct classifiers using expert annotations of VTA responses on a diverse set of forum posts .
Outcome: The proposed evaluation framework is rooted in learning sciences and tailored to asynchronous forum discussions.
DischargeSim: A Simulation Benchmark for Educational Doctor–Patient Communication at Discharge (2025.emnlp-main)

Copied to clipboard

Challenge: Discharge communication is a critical yet underexplored component of patient care, where the goal shifts from diagnosis to education.
Approach: They propose a benchmark that evaluates large language models’ ability to act as personalized discharge educators.
Outcome: Experiments with 18 LLMs show that model size does not always yield better education outcomes, highlighting trade-offs in strategy use and content prioritization.
Target-Guided Open-Domain Conversation (P19-1)

Copied to clipboard

Challenge: a new study aims to improve opendomain chat systems by integrating goals and strategy into the system.
Approach: They propose a structured approach that introduces coarse-grained keywords to control intended content of system responses and attains smooth conversation transition through turn-level supervised learning.
Outcome: The proposed system produces meaningful and effective conversations significantly better than other approaches.
A Survey on Detection of LLMs-Generated Content (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance .
Approach: They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy.
Outcome: The proposed model will be able to detect human-written content in real time.
Shami: A Corpus of Levantine Arabic Dialects (L18-1)

Copied to clipboard

Challenge: Modern Standard Arabic is the official written language used in education and media . however, the spoken language varies widely across the Arab world .
Approach: They construct a levantine dialect corpus covering data from four dialects spoken in four countries . they describe rules for pre-processing without affecting the meaning so that it is processable by NLP tools.
Outcome: The proposed corpus is larger than existing corpora in terms of size, words and vocabularies.
Book2Dial: Generating Teacher Student Interactions from Textbooks for Cost-Effective Development of Educational Chatbots (2024.findings-acl)

Copied to clipboard

Challenge: Educational chatbots are a promising tool for assisting student learning, but high-quality data is difficult to obtain due to privacy concerns.
Approach: They propose a framework for generating synthetic teacher-student interactions grounded in a set of textbooks and propose to open-source their results.
Outcome: The proposed framework captures a key aspect of learning interactions where curious students with partial knowledge ask teachers questions about the material in the textbook.
On the Effectiveness of Prompt-Moderated LLMs for Math Tutoring at the Tertiary Level (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been studied intensively in the context of education, yielding heterogeneous results.
Approach: They conduct a three-phase study with 49 students receiving a review of the topics, solving exercises, and writing an exam.
Outcome: The prompt-moderated LLMs performed better than the unmoderated model .
PCQPR: Proactive Conversational Question Planning with Reflection (2024.emnlp-main)

Copied to clipboard

Challenge: Current CQG methods focus on immediate context without strategic consideration of the specified conversational outcome.
Approach: They propose a method that uses a planning algorithm inspired by Monte Carlo Tree Search to generate contextually relevant questions.
Outcome: The proposed approach surpasses existing methods in e-learning and customer service fields . it generates contextually appropriate questions strategically devised to reach a specified outcome .
Designing Logic Pattern Templates for Counter-Argument Logical Structure Analysis (2024.findings-emnlp)

Copied to clipboard

Challenge: Despite their effectiveness, the logical attack structure of counterarguments remains unexplored due to its complexity.
Approach: They propose a task to analyze logical attack structure of counterarguments in relation to their corresponding opponent argument using 10 new CA logic patterns.
Outcome: The proposed task achieves high annotator agreement and coverage and high coverage on a dataset of 778 CAs.
Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language . due to the popularization of deep learning, ASR technology has led to a significant improvement in recognizing many languages.
Approach: They propose to use a dataset to analyze the data available for the Hong Kong Cantonese language . they use zh-HK as a source and a state-of-the-art ASR model to build a powerful model .
Outcome: The proposed model improves on the biggest existing dataset, Common Voice zh-HK.
Enhancing Textbooks with Visuals from the Web for Improved Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Textbooks lack visuals that support student learning, but many lack them . e-textbooks lack such visuals, and many lack these visuals .
Approach: They propose to use vision-language models to automatically enhance textbooks with images from the web.
Outcome: The proposed model improves textbooks with images from the web while allowing for better pedagogical value.
JADE: Corpus for Japanese Definition Modelling (2022.lrec-1)

Copied to clipboard

Challenge: Existing corpus for definition modelling techniques is limited to English . this study aimed to develop a corpus that provides definitions of words and phrases .
Approach: They investigated and released a corpus for Japanese definition modelling . the JADE provides 630k sets of targets, their definitions, and usage examples as contexts .
Outcome: The JADE corpus provides 630k sets of targets, their definitions, and usage examples as contexts for 41k unique targets.
Fora: A corpus and framework for the study of facilitated dialogue (2024.acl-long)

Copied to clipboard

Challenge: a new study of facilitated dialogues focuses on the sharing of personal experience . social media is a popular method of civic engagement but lacks the tools to analyze it .
Approach: They compile 262 facilitated conversations hosted with partner organizations . they taxonomize personal sharing behaviors and facilitation strategies in the corpus .
Outcome: The proposed framework can be used to analyze facilitated dialogues and parse spoken conversations . the data can be applied to other fields, including civic use in governance and social science .
A SMART Mnemonic Sounds like “Glue Tonic”: Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick (2024.emnlp-main)

Copied to clipboard

Challenge: a new study shows that mnemonics are not effective at matching student learning to a standardized learning model.
Approach: They build a keyword mnemonic generator that finds mnemonics students favor in a flashcard app . they use expressed and observed preferences to find out what students think is helpful .
Outcome: The proposed mnemonics outperform existing models in keyword mnemonics . the human writer outperformed both models in terms of keyword simplicity and explanation quality .
RAEmoLLM: Retrieval Augmented LLMs for Cross-Domain Misinformation Detection Using In-Context Learning Based on Emotional Information (2025.acl-long)

Copied to clipboard

Challenge: Current methods for cross-domain misinformation detection focus on in-domain tasks and do not incorporate significant sentiment and emotion features.
Approach: They propose a retrieval augmented (RAG) LLM framework that incorporates affective information into retrieval databases.
Outcome: The proposed framework improves on three misinformation benchmarks.
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts (2023.findings-emnlp)

Copied to clipboard

Challenge: Visual question answering (VQA) is a task that requires an understanding of both the image and the question to provide a natural language answer.
Approach: They propose a multimodal framework that leverages language guidance to answer questions more accurately.
Outcome: The proposed framework improves on the multi-choice question-answering task using CLIP and BLIP models.
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment (2025.findings-acl)

Copied to clipboard

Challenge: Texts above a student's readability level can lead to disengagement and disengagement . Developing readability models is crucial for improving literacy, language learning, and academic performance.
Approach: They introduce the Balanced Arabic Readability Evaluation Corpus (BAREC) a large-scale, fine-grained dataset for Arabic readability assessment.
Outcome: The proposed model outperforms existing methods in Arabic readability assessment.
An Empirical Comparison of LM-based Question and Answer Generation Methods (2023.findings-acl)

Copied to clipboard

Challenge: Question and answer generation (QAG) is a task of generating question-answer pairs given a context.
Approach: They propose to leverage sequence-to-sequence language model fine-tuning to generate question-answer pairs given a context.
Outcome: The proposed model outperforms other more convoluted approaches in the end-to-end model and is computationally light at both training and inference times.
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures (2024.emnlp-main)

Copied to clipboard

Challenge: Existing tools to aid residents in teaching medical doctors to explain decisions are a key objective of AI in education.
Approach: They present a multilingual dataset for Medical Question Answering where doctors can annotate correct and incorrect diagnoses with argument components and argument relations.
Outcome: The proposed dataset consists of 558 clinical cases with explanations in English, Spanish, French, Italian and annotated with argument components and argument relations.
LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring (2025.findings-emnlp)

Copied to clipboard

Challenge: a new framework for automated essay scoring is needed to achieve multi-perspective understanding and judgment.
Approach: They propose a roundtable essay scoring framework that performs precise and human-aligned scoring under a zero-shot setting.
Outcome: The proposed framework outperforms previous zero-shot AES approaches by enabling collaboration among agents with diverse evaluation perspectives.
Small Changes, Big Impact: How Manipulating a Few Neurons Can Drastically Alter LLM Aggression (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models have led to innovations in various domains such as education, healthcare, and finance, while raising serious concerns that they can be easily misused for malicious purposes.
Approach: They identify specific neurons (“aggression neurons”) closely related to the expression of aggression and analyze how manipulating them affects the model’s overall aggression.
Outcome: The proposed model outputs show that manipulating neurons can increase aggression by up to 33% in all models and even more extreme when they are concentrated in certain layers.
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction (2025.acl-long)

Copied to clipboard

Challenge: Multiple-choice questions (MCQs) are critical for identifying misconceptions and gaps in knowledge and accurately assessing students' understanding.
Approach: They propose to train a model to generate distractors that are more likely to be selected by students by a pairwise ranker and a distractor generator via Direct Preference Optimization.
Outcome: The proposed model outperforms baseline models and performs comparable to humans in various metrics including pairwise rank accuracy and distractor plausibility.
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations (2025.findings-acl)

Copied to clipboard

Challenge: Language models are widely used in education, yet their ability to tailor responses to learners with varied informational needs and knowledge backgrounds remains under-explored.
Approach: They conduct two extensive human studies to assess the utility of language model-generated explanatory answers (explanations) on a benchmark of 13.4K "Why" questions.
Outcome: The proposed model explanations match learners' educational backgrounds only 50% of the time, compared to 79% for lay explanations.
SuperST: Superficial Self-Training for Few-Shot Text Classification (2024.lrec-main)

Copied to clipboard

Challenge: In few-shot text classification, self-training relies on pseudo-labels to expand data, which has shown success, but can accumulate errors due to noisy pseudo-labeled data.
Approach: They propose a method to mitigate noise in noisy pseudo-labeled data by applying superficial learning to noisy data and fine-tuning to less noisy data.
Outcome: The proposed framework improves the classifier accuracy for few-shot text classification by 18.5% at most and 8% in average, compared with the state-of-the-art SSL baselines.
HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Education (2025.emnlp-main)

Copied to clipboard

Challenge: Text-to-image (T2I) generation has the potential to advance knowledge democratization and education.
Approach: They explore ways to harness T2I models for generating health knowledge flashcards . they curated a high-quality healthcare knowledge flash card dataset .
Outcome: The proposed models can generate health knowledge flashcards with appealing images . the results show that the open-source models can be fine tuned to generate health content .
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in education, yet their default usefulness conflicts with pedagogical principles.
Approach: They propose an adversarial student agent that they fine-tune to jailbreak LLM-based tutors and propose a benchmark to evaluate tutor robustness.
Outcome: The proposed model fine-tunes to jailbreak LLM-based tutors, and shows that they perform well under adversarial student attacks.
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling (2025.acl-long)

Copied to clipboard

Challenge: Chart generation requires strong visual design skills and precise coding capabilities that embed the desired visual properties into code.
Approach: They propose a vision-language model-based multi-agent framework for effective automatic chart generation.
Outcome: The proposed framework achieves a 5.2% improvement in the F1 score over the current best chart generation task.
BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks (2025.emnlp-main)

Copied to clipboard

Challenge: Existing Braille research focuses on isolated tasks while mixed-content Braille tasks face data scarcity and ambiguities.
Approach: They propose a syntax tree-based augmentation method tailored for Braille data.
Outcome: The proposed method improves Braille translation, formula-to-Braille conversion, and mixed-text translation.
Can LLM Agents Maintain a Persona in Discourse? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models are often subjected to context-shifting behaviour, resulting in a lack of consistent and interpretable personality-aligned interactions.
Approach: They propose to use two conversation agents to generate a discourse with an assigned personality from the OCEAN framework and then use multiple judge agents to infer original traits.
Outcome: The proposed model is based on two conversation agents with a personality assigned from the OCEAN framework and then multiple judge agents to infer the original traits assigned.
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs (2026.acl-long)

Copied to clipboard

Challenge: Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA).
Approach: They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata .
Outcome: The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset .
Would You Like to Make a Donation? A Dialogue System to Persuade You to Donate (2024.lrec-main)

Copied to clipboard

Challenge: Persuasive automated dialogue systems are a popular way to influence people's behavior and decision making.
Approach: They propose to use a context-aware persuasion strategy selection module to persult users . they also propose a persuasiveness prediction model to automatically evaluate the persuasiveness of generated text.
Outcome: The proposed system can achieve better performance on several automated evaluation metrics than baseline models.
Confidence Should Be Calibrated More Than One Turn Deep (2026.acl-long)

Copied to clipboard

Challenge: Existing work on confidence estimation and calibration focuses on single-turn settings . existing work on multi-turn calibration ignores the risks and potential of multi-turned conversations .
Approach: They propose a multi-turn calibration task that reframes calibration from a static property into a dynamic challenge central to reliable multi- turn conversations.
Outcome: The proposed model minimizes ECE@T and leverages ConfChat to improve confidence . the proposed model preserves and even enhances model performance in multi-turn interactions.
A Survey of Deep Learning for Geometry Problem Solving (2026.acl-long)

Copied to clipboard

Challenge: Recent surge in deep learning technologies has significantly accelerated research in this area.
Approach: They propose a comprehensive summary of the relevant tasks in geometry problem solving and a review of related deep learning methods.
Outcome: The proposed method is based on a systematic review of related methods and evaluation metrics and methods.
Simulated Students in Tutoring Dialogues: Substance or Illusion? (2026.acl-long)

Copied to clipboard

Challenge: evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up.
Approach: They propose to define the student simulation task and benchmark a wide range of student simulation methods on these metrics.
Outcome: The proposed evaluation metrics show that prompting strategies perform poorly on a real-world tutoring dialogue dataset.
SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box Jailbreaking (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for jailbreak ignore the semantic differences between categories of harmful questions, leading to inconsistent success rates and reduced overall attack effectiveness.
Approach: They propose a category-aware jailbreak framework that incorporates the semantic category of harmful questions into prompt generation.
Outcome: The proposed framework improves attack success rates and category alignment and achieves better cross-category robustness compared to the state-of-the-art (SOTA) baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations