Papers with education
Copied to clipboard
| Challenge: | Existing methods to gather test items from validated item databases are under-researched, but there is little research on assembling exam items from a database of valid items. |
| Approach: | They propose to use semantic search to assist vocational educators in assembling exam forms by using eight retrieval strategies and 25 popular sentence similarity models. |
| Outcome: | The proposed tool is based on eight retrieval strategies and 25 popular pre-trained sentence similarity models. |
Copied to clipboard
| Challenge: | Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting . |
| Approach: | They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models . |
| Outcome: | The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects. |
Copied to clipboard
| Challenge: | Existing models for age classification of students and non-students are restrictive and require access to many tweets. |
| Approach: | They propose a model which uses 3 tweet-content features to classify users as students or non-students. |
| Outcome: | The proposed model achieves an accuracy of 88.1% and an F1 score of .704 compared to previous models, which require access to many user tweets. |
Copied to clipboard
| Challenge: | Edu-ConvoKit is an open-source library for analyzing education conversation data. |
| Approach: | They introduce Edu-ConvoKit, an open-source library for conversation data analysis. |
| Outcome: | The open-source library handles pre-processing, annotation and analysis of education conversation data. |
Copied to clipboard
| Challenge: | Existing concepts-based quiz generation frameworks that leverage external knowledge sources are challenging and labor intensive. |
| Approach: | They propose a concept-based quiz generation framework that leverages external knowledge sources to assess the quality of the generated quizzes, using LLMs as judges. |
| Outcome: | The proposed framework shows a 4.8% improvement in evaluation scores and a 77.52% win rate in pairwise comparisons against baseline quiz sets. |
Copied to clipboard
| Challenge: | Socratic questioning is a form of reflective inquiry often employed in education to encourage critical thinking in students. |
| Approach: | They present a first large dataset of 110K questions, context pairs for Socratic Question Generation. |
| Outcome: | The proposed model produces realistic, type-sensitive, human-like Socratic questions . authors show that the model can be used in counseling and coaching . |
Copied to clipboard
| Challenge: | In this paper, we focus on task-oriented dialog systems, although our framework allows easy integration of non-task dialog systems and their combination. |
| Approach: | They propose an open source dialog system framework for education and research that supports multi-domain task-oriented conversations in two languages. |
| Outcome: | The proposed framework supports multi-domain task-oriented conversations in two languages and is open source for education and research. |
Copied to clipboard
| Challenge: | Despite the practical importance of exam assembly, few methods exist to support educators during manual item retrieval for exam assembly tasks. |
| Approach: | They propose a mixed-integer programming re-ranking approach to improve relevance while mitigating bias on an industry-grade exam assembly platform. |
| Outcome: | The proposed approach improves relevance and reduces bias by 17% when compared to other methods on a real-world exam assembly platform. |
Copied to clipboard
| Challenge: | a 2010 survey found that the language of the European Union, not even English, was not fully supplied . the absence of Language Resources stifles teaching and technology building, authors say . |
| Approach: | They propose to harness the power of alternative incentives to elicit linguistic data and annotation . they also describe changes to the workflows necessary to collect data from workforces attracted by incentives . |
| Outcome: | a new initiative to harness incentives to elicit linguistic data and annotation improves language resources . the NIEUW project is funded by the u.s. national science foundation . |
Copied to clipboard
| Challenge: | Empirical natural language processing (NLP) systems involve interoperation among multiple components . a wealth of NLP toolkits exist ( 4), such as spaCy, DKPro, CoreNLP. |
| Approach: | They propose a unified open-source framework that supports fast development of NLP workflows . framework includes processors for NLP tasks, visualization, and annotation . |
| Outcome: | The framework offers processors for NLP tasks, visualization, and annotation, and is extensible . it is delivered through two modularized yet integratable open-source projects, Forte and Stave . |
Copied to clipboard
| Challenge: | Existing lexicons of affect only show coarse associations, but are not accurate as human-created ones. |
| Approach: | They propose to use a manually created affect intensity lexicon with real-valued intensity scores for anger, fear, joy, and sadness. |
| Outcome: | The lexicon has real-valued scores for anger, fear, joy, and sadness . anger, fears, and sad words have very similar VAD scores . |
Copied to clipboard
| Challenge: | Using natural language processing, we examine teacher evaluations to improve educational programs. |
| Approach: | They apply natural language processing techniques to a Filipino education non-profit to analyze teacher evaluations written by Teacher Fellows. |
| Outcome: | The proposed framework can be applied to teacher evaluations from a Filipino education non-profit. |
Copied to clipboard
| Challenge: | a large number of machine-generated texts are often hard to distinguish between human-written and machine-generated text . this raises concerns about potential misuse, especially within educational and academic domains . |
| Approach: | They propose a system that can detect whether a text is human-written or machine-generated . they use a fine-grained classification schema to identify the use of machine-generated text . |
| Outcome: | The proposed system can distinguish between human-written and machine-generated text . it can detect attempts to obfuscate the fact that a text was machine- generated . |
Copied to clipboard
| Challenge: | Existing adaptive testing methods face several challenges due to mechanized nature of most algorithms and noisy response data. |
| Approach: | They propose to use large language models to enhance adaptive testing through interactive engagement to capture test-takers’ responses and anomalies. |
| Outcome: | The proposed agent achieves more accurate results with 20% fewer questions than state-of-the-art baselines and testers preferred it in speed, smoothness, and other dimensions. |
Copied to clipboard
| Challenge: | Text-to-speech (TTS) systems are limited by limited data and linguistic complexities. |
| Approach: | They propose a data-optimized framework with an advanced acoustic model to build high-quality TTS systems for low-resource scenarios. |
| Outcome: | The proposed framework enables zero-shot voice cloning and improved performance across diverse client applications, including finance, healthcare, education, and law. |
Copied to clipboard
| Challenge: | In education, peer instruction is widely recognized as an effective active learning strategy, but evaluations of PI are limited by logistical constraints and variability in classroom settings. |
| Approach: | They propose a simulation framework that integrates Agent-Based Modeling, Large Language Models, and Bayesian Knowledge Tracing to emulate student learning dynamics. |
| Outcome: | The proposed framework integrates Agent-Based Modeling, Large Language Models, and Bayesian Knowledge Tracing to emulate student learning dynamics in real classrooms. |
Copied to clipboard
| Challenge: | incorporating non-adopter perspectives is essential for developing useful and capable LLMs, argues a new study. |
| Approach: | They argue that incorporating non-adopter perspectives is essential for developing broadly useful and capable LLMs. |
| Outcome: | The proposed method will risk missing tasks prioritized by non-adopters, the authors argue . they show that non-dots diverge from those of current users, and non-no-acopter needs point towards novel reasoning tasks. |
Copied to clipboard
| Challenge: | Existing research has focused on plain text, while real-world K-12 scenarios often involve multimodal data. |
| Approach: | They propose a unified language and vision assistant called UniEDU for educational applications . it excels across multiple educational tasks while maintaining strong generalization capabilities . authors propose to use UniEDu for industry-scale deployment . |
| Outcome: | The proposed model excels across multiple educational tasks while maintaining strong generalization capabilities. |
Copied to clipboard
| Challenge: | Current solutions such as quantization, pruning, and Retrieval-Augmented Generation (RAG) offer only partial optimizations and often sacrifice accuracy, speed, or generality. |
| Approach: | They propose an end-to-end optimization framework for efficient LLM deployment . it leverages Hierarchical Speculative Decoding (HSD) for faster inference without quality loss. |
| Outcome: | HOLA delivers +17.6% EMA on GSM8K, +10.5% MCA on ARC, and reduced latency and memory on edge devices like Jetson Nano. |
Copied to clipboard
| Challenge: | Existing studies focus on narrative or role-playing tasks and overlook how adversarial conversational history alone can reshape induced personas. |
| Approach: | They propose a framework that embeds semantically loaded cues into user queries to gradually induce reverse personas. |
| Outcome: | The proposed framework predictably shifts personas, triggers collateral changes in correlated traits, and exhibits stronger effects in multi-turn settings. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated the potential of large language models (LLMs) for automatic error detection in math word problems (MWPs). |
| Approach: | They propose a framework that generates adaptive reference solutions using LLMs to enhance error detection by reducing conformity bias in MWPs. |
| Outcome: | The proposed framework mitigates the performance gap between conventional and alternative solutions in MWPs, especially when combined with reasoning-enhancing techniques like chain-of-thought prompting. |
Copied to clipboard
| Challenge: | Large language models generate fluent responses to user queries, but they are also susceptible to misuse in journalism, education, and academia. |
| Approach: | They propose a large-scale benchmark for machine-generated text detection that is a multi-generator, multi-domain, and multi-lingual corpus. |
| Outcome: | The proposed system can detect machine-generated text and pinpoint misuse . the proposed system is based on a large-scale benchmark dataset . |
Copied to clipboard
| Challenge: | a vast number of papers accepted at top NLP venues come from a handful of western countries and (lately) China. |
| Approach: | They ask researchers to examine the relationship between geographical location and publication success . they use a dataset of 70,000 papers from the ACL Anthology to examine their citation network . |
| Outcome: | The proposed dataset of 70,000 papers from the ACL Anthology shows that there are substantial geographical disparities in paper acceptance and citations . |
Copied to clipboard
| Challenge: | a novel dataset summarizes student reflections on STEM lectures . ReflectASP eases the exploration of open-aspect-based summarization (OABS) despite the limitations of current datasets, it is still under-explored. |
| Approach: | They propose a dataset that summarizes student reflections on STEM lectures . they propose two refinement methods to improve summaries . |
| Outcome: | The proposed dataset summarizes student reflections on STEM lectures using automatic and human evaluations. |
Copied to clipboard
| Challenge: | Compared to math word problems, geometry problems emphasize multi-modal formats and the translation between informal and formal languages. |
| Approach: | They propose a symbolic deduction engine-based geometry problem generation framework that leverages a symbolic deduction engine to generate geometry problems. |
| Outcome: | The proposed method avoids inherent biases in translating natural language into formal language and guarantees to control the generated problems in terms of knowledge points and difficulties by an elaborate checking function. |
Copied to clipboard
| Challenge: | Existing studies on controllable text generation focus on controlling attributes such as sentiment, writing style, and writing style. |
| Approach: | They introduce a metric that quantifies semantic diversity for scenario generation under fixed abstract semantic constraints and validate it through controlled experiments. |
| Outcome: | The proposed metric achieves excellent discrimination accuracy (100% and 91.9%, respectively), with discriminative power up to 5.5 greater than the best baseline. |
Copied to clipboard
| Challenge: | a study aimed to determine if people who are influential in online discussions retain influence when placed in a topic that is less familiar or perhaps not as interesting. |
| Approach: | They conducted a study to determine if people who are highly influential retain influence when moving to a topic that is less familiar or perhaps not as interesting. |
| Outcome: | The results show that people who are highly influential in group discussions lose influence when placed in a topic that is less familiar or perhaps not as interesting. |
Copied to clipboard
| Challenge: | Our work explores the potential of large language models (LLMs) to close the novice-expert knowledge gap in remediating math mistakes. |
| Approach: | They propose a method that uses cognitive task analysis to translate an expert’s latent thought process into a decision-making model for remediation. |
| Outcome: | The proposed model can bridge the novice-expert knowledge gap by using cognitive task analysis to translate an expert’s latent thought process into a decision-making model for remediation. |
Copied to clipboard
| Challenge: | Despite extensive research showing the positive impact of uptake on student learning and achievement, there is little evidence that it is effective in teaching. |
| Approach: | They propose a framework for computationally measuring uptake by releasing a dataset of student-teacher exchanges extracted from US math classroom transcripts annotated for uptake . they formalize uptake as pointwise Jensen-Shannon Divergence (pJSD) and conduct a linguistically-motivated comparison of different unsupervised measures. |
| Outcome: | The proposed framework outperforms baseline measures in identifying uptake phenomena like question answering and reformulation. |
Copied to clipboard
| Challenge: | Existing studies have examined the reliability of Large Language Models (LLMs) in grading authentic student problem solving processes and delivering effective feedback. |
| Approach: | They propose to use a dataset to evaluate the reliability of large language models in mathematics and a teacher-written feedback system to improve student problem-solving processes. |
| Outcome: | The proposed model improves in correctness classification, error identification, and feedback generation, but generates a gap from teacher-written feedback. |
Copied to clipboard
| Challenge: | Existing methods to improve text classification in education suffer from data scarcity . authors propose a retrieval approach that provides effective learning in educational text classification. |
| Approach: | They propose a retrieval approach that provides effective learning in educational text classification by introducing cross-encoder style texts to a bi-encoding architecture. |
| Outcome: | The proposed method is effective in multi-label scenarios and low-resource tags compared to state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing approaches to ensuring the safety of Large Language Models (LLMs) rely on invasive fine- tuning or external generation-based checks, which can be opaque and resource-inefficient. |
| Approach: | They propose a mechanistic method that identifies the layer where safe and unsafe concepts are maximally separable within a pretrained representation space. |
| Outcome: | The proposed method can be used across multiple domains, diverse tasks, and 16 non-English languages on encoder and decoder architectures. |
Copied to clipboard
| Challenge: | a new metric measures the quality of large language models (LLMs) that detects hidden misalignments and jailbreak risks. |
| Approach: | They propose a decoding-invariant metric that measures latent safety failures . they propose 'Alignment Quality Index' to measure latent activations in latent space . |
| Outcome: | The proposed metric detects latent safety failures overlooked by behavioral benchmarks and jailbreaks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs). |
| Approach: | They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs. |
| Outcome: | The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints. |
Copied to clipboard
| Challenge: | Emotion classification is an important task with applications in education, virtual reality, and robotics. |
| Approach: | They propose to use token embeddings to generate a "semantic-anchor graph" using semantic anchors, sentences can be projected onto them to form a graph . |
| Outcome: | Empirically, the proposed system can generate meaningful semantic anchors and discriminative graph patterns for different emotion. |
Copied to clipboard
| Challenge: | In this paper, we propose a novel Personalized Multimodal Feedback Generation Network (PMFGN) that generates personalized feedback for teachers to evaluate assignments involving multimodal inputs. |
| Approach: | They propose a Personalized Multimodal Feedback Generation Network (PMFGN) that generates personalized feedback for teachers to evaluate assignments involving multimodal inputs such as images, audios, and texts. |
| Outcome: | The proposed model outperforms baseline models on real-world K-12 education data and detailed ablation experiments to deepen understanding of the proposed framework. |
Copied to clipboard
| Challenge: | Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing. |
| Approach: | They argue for a reform of multiple choice question answering (MCQA) they argue for more generative formats based on human testing . |
| Outcome: | The proposed reforms improve the quality of MCQA, the authors argue . they show that even when MCQ is a useful format, its datasets suffer from leakage, unanswerability, shortcuts and saturation. |
Copied to clipboard
| Challenge: | Recognizing LLMs’ capability to generate educational content can lead to advances in automated and personalized learning. |
| Approach: | They propose to evaluate the questioning capability in education as a teacher of large language models by evaluating their generated educational questions. |
| Outcome: | The proposed model can generate educational content that aligns with human perspectives and is more apt as an interdisciplinary teacher. |
Copied to clipboard
| Challenge: | Despite recent attempts on computational modeling of the variation, the lack of parallel corpora of style language makes it difficult to systematically control the stylistic change and evaluate such models. |
| Approach: | They propose to use a parallel and annotated stylistic language dataset to test the effectiveness of style transfer models. |
| Outcome: | The proposed model outperforms the unsupervised models using nonparallel corpus. |
Copied to clipboard
| Challenge: | Existing computational models of the verbal morphology of the Métis language are insufficient to model the language's unique phonological interactions. |
| Approach: | They propose a finite-state computational model of the verbal morphology of Michif . they use composed finite state transducers to model concatenative morphologies . |
| Outcome: | The proposed model is based on a series of finite-state transducers. |
Copied to clipboard
| Challenge: | a new study evaluates the expressivity of large language models for communicating implicitly . authors: models can express tone, identity, and intent beyond literal meanings . phrasing and tone of a message can convey a number of topics beyond literal contexts - authors . |
| Approach: | They propose a framework to evaluate the expressivity of large language models . they use a social-linguistic grader to validate their models against human judgments . |
| Outcome: | The proposed framework quantifies how well LLM-generated text communicates target properties without explicit mention across nine tasks spanning emotion, identity, and tone. |
Copied to clipboard
| Challenge: | Persona-assigned large language models are used in education, healthcare and sociodemographic simulations. |
| Approach: | They propose a protocol that combines long persona dialogues and evaluation datasets to create dialogue-conditioned benchmarks that can robustly measure long-context effects. |
| Outcome: | The proposed protocol can measure persona fidelity, instruction-following, and safety in long conversations. |
Copied to clipboard
| Challenge: | Automatic speech recognition systems fail to accurately interpret speech patterns deviating from typical fluency, leading to critical usability issues and misinterpretations. |
| Approach: | They evaluate six leading automatic speech recognition systems based on a real-world dataset and a synthetic dataset derived from the widely-used LibriSpeech benchmark. |
| Outcome: | The six leading speech recognition systems were evaluated on a real-world dataset and a synthetic dataset derived from the widely-used LibriSpeech benchmark. |
Copied to clipboard
| Challenge: | We hypothesize that questioning can enhance human performance and assist solvers . |
| Approach: | They propose to use large language models to generate sequential questions for math word problem-solving . they propose to apply these models to a variety of math word problems . |
| Outcome: | The proposed model improves the performance of a math word problem solver by generating more questions than other models. |
Copied to clipboard
| Challenge: | Bangla is the sixth most spoken language worldwide and the second Indo-Aryan language after Hindi. |
| Approach: | They propose an annotated sentiment analysis dataset made of informally written Bangla texts. |
| Outcome: | The proposed dataset is compared with neural networks and pretrained models . it shows that hand-crafted lexical features provide superior performance than neural networks . |
Copied to clipboard
| Challenge: | Existing jailbreaks against large audio-language models fall into two categories . early work converted text-based prompts into synthetic speech, while subsequent work introduced minor acoustic variations such as accent shifts, phonetic spellings, or stress patterns. |
| Approach: | They propose a text-to-audio jailbreak that embeds disallowed directives within a narrative-style audio stream. |
| Outcome: | The proposed attack exploits structural and acoustic properties of a text-to-audio model . it achieves 98.26% success rate, significantly exceeding baselines for text-based models . |
Copied to clipboard
| Challenge: | Existing studies have shown that relying on LLMs as information providers may hurt student learning. |
| Approach: | They introduce and apply two bias score metrics to evaluate LLMs for bias in the personalized educational setting, specifically on the models’ roles as “teachers.” |
| Outcome: | The proposed models harm student learning by perpetuating harmful stereotypes and reversing them. |
Copied to clipboard
| Challenge: | Existing approaches to simulate tutor behaviors or preferences fail to sustain high-quality pedagogical conversations that provide explicit stepwise scaffolding and adapt to learners’ evolving cognitive states. |
| Approach: | They propose a planning-guided tutoring framework with an assessment-driven memory for multi-turn math dialogue tutoring. |
| Outcome: | Experiments on multi-turn math tutoring benchmarks show that ScaffoldLM significantly improves pedagogical tutoring quality over strong baselines. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly being applied in education, showing significant potential in personalized instruction, student feedback, and intelligent tutoring systems (ITSs). |
| Approach: | They propose a dataset specifically designed to evaluate LLMs’ ability to generate high-quality hints for Math Word Problems. |
| Outcome: | The proposed dataset shows that LLMs can generate more accurate and contextually appropriate educational hints for math word problems without offering direct answers. |
Copied to clipboard
| Challenge: | Large-scale Vision-Language Models (LVLMs) are being deployed in real-world settings that require visual inference. |
| Approach: | They evaluate LVLMs' ability to account for variation in color perception using the Ishihara Test. |
| Outcome: | The proposed models fail to reproduce the perceptual outcomes experienced by affected individuals and default to normative color perception. |
Copied to clipboard
| Challenge: | Discourse information is crucial for many NLP tasks due to the great distance that spans between the two languages. |
| Approach: | They propose to use a Spanish-Chinese parallel corpus with annotated discourse information to serve for bilingual language education. |
| Outcome: | The proposed corpus is composed of 100 Spanish-Chinese parallel texts, and all the discourse markers (DM) have been annotated to form the education source. |
Copied to clipboard
| Challenge: | Mexico has 68 linguistic groups and 364 varieties, but lack of data on social media and internet is putting them at risk. |
| Approach: | They propose a collaborative corpus for endangered languages in Mexico . they propose linguistic search, digitalization and alignment process for each language . |
| Outcome: | The proposed corpus aligns Spanish with six indigenous languages: Maya, Ch’ol, Mazatec, Mixtec, Otomi, and Nahuatl. |
Copied to clipboard
| Challenge: | Multiple choice questions (MCQs) are crucial for deep thinking and knowledge integration in education. |
| Approach: | They propose a cross-modal options synthesis framework for generating MCQs with visual options. |
| Outcome: | The proposed framework produces a plausible and visually similar answer and distractor . it also includes a discrimination module to identify content suitable for visual options . |
Copied to clipboard
| Challenge: | Persona agents are LLM agents conditioned to act according to an assigned persona . evaluating how faithfully these agents adhere to their personas remains a challenge . |
| Approach: | a new study evaluates persona agents' ability to act according to an assigned persona . a persona agent's person score is a human-aligned automatic metric that can be used to evaluate a model . |
| Outcome: | a new evaluation framework and a human-aligned automatic metric show that persona agents can perform better. |
Copied to clipboard
| Challenge: | Using data from English cloze tests, we demonstrate wide performance gaps across demographic groups and show that pretrained language models disfavor young non-white male speakers. |
| Approach: | They use data from English cloze tests to examine performance differences of pretrained language models across demographic groups. |
| Outcome: | The models disfavor young non-white male speakers, but larger models reduce performance gaps between majority and minority groups. |
Copied to clipboard
| Challenge: | Recent studies have revealed that NLP is limited to a subset of the world’s 6,500 languages. |
| Approach: | They propose a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP. |
| Outcome: | The proposed framework estimates the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP. |
Copied to clipboard
| Challenge: | Existing approaches to solving math word problems focus on obtaining the correct answer. |
| Approach: | They propose a step-by-step planning approach for intermediate solution generation that strategically plans the generation of the next solution step based on the MWP and the previous solution steps. |
| Outcome: | The proposed approach improves the accuracy and interpretability of the solution on automatic metrics and human evaluation. |
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) can grasp the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent. |
| Approach: | They propose a framework for multimodal large language models to grasp the intention of a question and decompose it into a series of visual recognition sub-tasks to find out the answer. |
| Outcome: | The proposed framework improves the accuracy of complex video-related questions by 29.6% and 17.2% on CVQA and the existing VQA datasets. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models are increasingly used in Personalized Image Aesthetic Assessment (PIAA) however, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education. |
| Approach: | They propose to evaluate MLLMs along two complementary dimensions: (1) stereotype bias and (2) alignment between model outputs and genuine human aesthetic preferences. |
| Outcome: | The proposed benchmark covers three subtasks: aesthetic perception, assessment, empathy and alignment between outputs and genuine human aesthetic preferences. |
Copied to clipboard
| Challenge: | Existing automated student answer assessment models lack explainable and faithful feedback. |
| Approach: | They propose a framework that leverages ChatGPT for student answer scoring and rationale generation. |
| Outcome: | The proposed method improves the overall QWK score by 11% compared to ChatGPT. |
Copied to clipboard
| Challenge: | Recent datasets for automatic speech recognition in Brazilian Portuguese lack diversity in terms of age groups, regional accents, and education levels. |
| Approach: | They propose to use a dataset to analyze the impact of ASR in Brazilian Portuguese (BP) they demonstrate that current models are biased regarding age, education, and regional accents. |
| Outcome: | The proposed dataset helps mitigate biases in current ASR models regarding education levels and age groups. |
Copied to clipboard
| Challenge: | Existing argument component classifications in education are simplistic and isolated, failing to capture the complete argument information. |
| Approach: | They propose to annotate a manually annotated argument component classification dataset from authentic examination settings and to explore the performance of Large Language Models on CEAMC. |
| Outcome: | The proposed dataset can be used to analyze argumentative essays in education. |
Copied to clipboard
| Challenge: | a new framework for understanding state-level legislative process improves understanding of state legislation and its implications. |
| Approach: | They propose to use generative large language models to decode legislators' behavior and implications of state policies by establishing a shared nationwide network. |
| Outcome: | The framework decodes legislators’ behavior and implications of state policies by establishing a shared nationwide network enriched with diverse contexts, such as information on interest groups influencing public policy and legislators' courage test results, which reflect their political positions. |
Copied to clipboard
| Challenge: | Existing studies have neglected the systematic design and procedure evaluation of court simulations, which are critical to the credibility and usage of court simulators in practice. |
| Approach: | They propose a court simulation paradigm based on the real-world procedure structure of Chinese courts and a framework that focuses on both legal judgment prediction and court procedure analysis. |
| Outcome: | The proposed model outperforms judges and lawyers from the real trials in many aspects. |
Copied to clipboard
| Challenge: | despite the rapid development of Large Language Models, there is no dedicated benchmark for evaluating LLMs in Chinese K-12 education. |
| Approach: | They propose to develop a benchmark specifically tailored for Chinese K-12 education. |
| Outcome: | EVAL is the first evaluation benchmark specifically tailored for Chinese K-12 education. |
Copied to clipboard
| Challenge: | Existing benchmarks for automated grading of student work fail to evaluate real student responses . existing models fail to assess real student work, especially on cognitively demanding tasks . |
| Approach: | They propose a multimodal benchmark for rubric-aligned evaluation of real Chinese K-12 student answers. |
| Outcome: | The proposed model improves performance and interpretability of existing models on EduMARS . existing models fail to perform on real-world, cognitively demanding tasks, authors say . |
Copied to clipboard
| Challenge: | Existing CodeLLM benchmarks rely on a single expert-written prompt per problem . a growing body of work shows their utility to professional programmers . |
| Approach: | They propose a natural-language-to-code benchmark of prompts written by non-experts . student prompts are written by 80 students who have only completed one introductory Python course . |
| Outcome: | The proposed model is better discriminator of student prompt descriptions than existing benchmarks. |
Copied to clipboard
| Challenge: | linguistic diversity of India poses significant machine translation challenges, authors say . underrepresented tribal languages like Bhili lack high-quality linguistic resources . |
| Approach: | They introduce a Bhili-Hindi-English Parallel Corpus, the first and largest parallel corpus worldwide . they evaluated a wide range of proprietary and open-source MLLMs on bidirectional translation tasks . |
| Outcome: | The proposed corpus spans critical domains such as education, administration, and news. |
Copied to clipboard
| Challenge: | Existing tools for lexical simplification are not tailored to language education with word levels and lists of candidates subjective. |
| Approach: | They construct a language dataset for lexical simplification based on CEFR levels . target and candidate words are assigned CEFR-J wordlists and English Vocabulary Profile . |
| Outcome: | The proposed method is based on the common European Framework of References for Languages (CEFR) levels and candidates are selected using an online thesaurus. |
Copied to clipboard
| Challenge: | Existing assessments rely on surface-level metrics and lack sufficient grounding in educational theory . a new framework is proposed to evaluate VTAs in asynchronous learning environments . |
| Approach: | They propose a pedagogically-oriented evaluation framework tailored to asynchronous forum discussions . they construct classifiers using expert annotations of VTA responses on a diverse set of forum posts . |
| Outcome: | The proposed evaluation framework is rooted in learning sciences and tailored to asynchronous forum discussions. |
Copied to clipboard
| Challenge: | Discharge communication is a critical yet underexplored component of patient care, where the goal shifts from diagnosis to education. |
| Approach: | They propose a benchmark that evaluates large language models’ ability to act as personalized discharge educators. |
| Outcome: | Experiments with 18 LLMs show that model size does not always yield better education outcomes, highlighting trade-offs in strategy use and content prioritization. |
Copied to clipboard
| Challenge: | a new study aims to improve opendomain chat systems by integrating goals and strategy into the system. |
| Approach: | They propose a structured approach that introduces coarse-grained keywords to control intended content of system responses and attains smooth conversation transition through turn-level supervised learning. |
| Outcome: | The proposed system produces meaningful and effective conversations significantly better than other approaches. |
Copied to clipboard
| Challenge: | Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance . |
| Approach: | They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy. |
| Outcome: | The proposed model will be able to detect human-written content in real time. |
Copied to clipboard
| Challenge: | Modern Standard Arabic is the official written language used in education and media . however, the spoken language varies widely across the Arab world . |
| Approach: | They construct a levantine dialect corpus covering data from four dialects spoken in four countries . they describe rules for pre-processing without affecting the meaning so that it is processable by NLP tools. |
| Outcome: | The proposed corpus is larger than existing corpora in terms of size, words and vocabularies. |
Copied to clipboard
| Challenge: | Educational chatbots are a promising tool for assisting student learning, but high-quality data is difficult to obtain due to privacy concerns. |
| Approach: | They propose a framework for generating synthetic teacher-student interactions grounded in a set of textbooks and propose to open-source their results. |
| Outcome: | The proposed framework captures a key aspect of learning interactions where curious students with partial knowledge ask teachers questions about the material in the textbook. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been studied intensively in the context of education, yielding heterogeneous results. |
| Approach: | They conduct a three-phase study with 49 students receiving a review of the topics, solving exercises, and writing an exam. |
| Outcome: | The prompt-moderated LLMs performed better than the unmoderated model . |
Copied to clipboard
| Challenge: | Current CQG methods focus on immediate context without strategic consideration of the specified conversational outcome. |
| Approach: | They propose a method that uses a planning algorithm inspired by Monte Carlo Tree Search to generate contextually relevant questions. |
| Outcome: | The proposed approach surpasses existing methods in e-learning and customer service fields . it generates contextually appropriate questions strategically devised to reach a specified outcome . |
Copied to clipboard
| Challenge: | Despite their effectiveness, the logical attack structure of counterarguments remains unexplored due to its complexity. |
| Approach: | They propose a task to analyze logical attack structure of counterarguments in relation to their corresponding opponent argument using 10 new CA logic patterns. |
| Outcome: | The proposed task achieves high annotator agreement and coverage and high coverage on a dataset of 778 CAs. |
Copied to clipboard
| Challenge: | In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language . due to the popularization of deep learning, ASR technology has led to a significant improvement in recognizing many languages. |
| Approach: | They propose to use a dataset to analyze the data available for the Hong Kong Cantonese language . they use zh-HK as a source and a state-of-the-art ASR model to build a powerful model . |
| Outcome: | The proposed model improves on the biggest existing dataset, Common Voice zh-HK. |
Copied to clipboard
| Challenge: | Textbooks lack visuals that support student learning, but many lack them . e-textbooks lack such visuals, and many lack these visuals . |
| Approach: | They propose to use vision-language models to automatically enhance textbooks with images from the web. |
| Outcome: | The proposed model improves textbooks with images from the web while allowing for better pedagogical value. |
Copied to clipboard
| Challenge: | Existing corpus for definition modelling techniques is limited to English . this study aimed to develop a corpus that provides definitions of words and phrases . |
| Approach: | They investigated and released a corpus for Japanese definition modelling . the JADE provides 630k sets of targets, their definitions, and usage examples as contexts . |
| Outcome: | The JADE corpus provides 630k sets of targets, their definitions, and usage examples as contexts for 41k unique targets. |
Copied to clipboard
| Challenge: | a new study of facilitated dialogues focuses on the sharing of personal experience . social media is a popular method of civic engagement but lacks the tools to analyze it . |
| Approach: | They compile 262 facilitated conversations hosted with partner organizations . they taxonomize personal sharing behaviors and facilitation strategies in the corpus . |
| Outcome: | The proposed framework can be used to analyze facilitated dialogues and parse spoken conversations . the data can be applied to other fields, including civic use in governance and social science . |
Copied to clipboard
| Challenge: | a new study shows that mnemonics are not effective at matching student learning to a standardized learning model. |
| Approach: | They build a keyword mnemonic generator that finds mnemonics students favor in a flashcard app . they use expressed and observed preferences to find out what students think is helpful . |
| Outcome: | The proposed mnemonics outperform existing models in keyword mnemonics . the human writer outperformed both models in terms of keyword simplicity and explanation quality . |
Copied to clipboard
| Challenge: | Current methods for cross-domain misinformation detection focus on in-domain tasks and do not incorporate significant sentiment and emotion features. |
| Approach: | They propose a retrieval augmented (RAG) LLM framework that incorporates affective information into retrieval databases. |
| Outcome: | The proposed framework improves on three misinformation benchmarks. |
Copied to clipboard
| Challenge: | Visual question answering (VQA) is a task that requires an understanding of both the image and the question to provide a natural language answer. |
| Approach: | They propose a multimodal framework that leverages language guidance to answer questions more accurately. |
| Outcome: | The proposed framework improves on the multi-choice question-answering task using CLIP and BLIP models. |
Copied to clipboard
| Challenge: | Texts above a student's readability level can lead to disengagement and disengagement . Developing readability models is crucial for improving literacy, language learning, and academic performance. |
| Approach: | They introduce the Balanced Arabic Readability Evaluation Corpus (BAREC) a large-scale, fine-grained dataset for Arabic readability assessment. |
| Outcome: | The proposed model outperforms existing methods in Arabic readability assessment. |
Copied to clipboard
| Challenge: | Question and answer generation (QAG) is a task of generating question-answer pairs given a context. |
| Approach: | They propose to leverage sequence-to-sequence language model fine-tuning to generate question-answer pairs given a context. |
| Outcome: | The proposed model outperforms other more convoluted approaches in the end-to-end model and is computationally light at both training and inference times. |
Copied to clipboard
| Challenge: | Existing tools to aid residents in teaching medical doctors to explain decisions are a key objective of AI in education. |
| Approach: | They present a multilingual dataset for Medical Question Answering where doctors can annotate correct and incorrect diagnoses with argument components and argument relations. |
| Outcome: | The proposed dataset consists of 558 clinical cases with explanations in English, Spanish, French, Italian and annotated with argument components and argument relations. |
Copied to clipboard
| Challenge: | a new framework for automated essay scoring is needed to achieve multi-perspective understanding and judgment. |
| Approach: | They propose a roundtable essay scoring framework that performs precise and human-aligned scoring under a zero-shot setting. |
| Outcome: | The proposed framework outperforms previous zero-shot AES approaches by enabling collaboration among agents with diverse evaluation perspectives. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models have led to innovations in various domains such as education, healthcare, and finance, while raising serious concerns that they can be easily misused for malicious purposes. |
| Approach: | They identify specific neurons (“aggression neurons”) closely related to the expression of aggression and analyze how manipulating them affects the model’s overall aggression. |
| Outcome: | The proposed model outputs show that manipulating neurons can increase aggression by up to 33% in all models and even more extreme when they are concentrated in certain layers. |
Copied to clipboard
| Challenge: | Multiple-choice questions (MCQs) are critical for identifying misconceptions and gaps in knowledge and accurately assessing students' understanding. |
| Approach: | They propose to train a model to generate distractors that are more likely to be selected by students by a pairwise ranker and a distractor generator via Direct Preference Optimization. |
| Outcome: | The proposed model outperforms baseline models and performs comparable to humans in various metrics including pairwise rank accuracy and distractor plausibility. |
Copied to clipboard
| Challenge: | Language models are widely used in education, yet their ability to tailor responses to learners with varied informational needs and knowledge backgrounds remains under-explored. |
| Approach: | They conduct two extensive human studies to assess the utility of language model-generated explanatory answers (explanations) on a benchmark of 13.4K "Why" questions. |
| Outcome: | The proposed model explanations match learners' educational backgrounds only 50% of the time, compared to 79% for lay explanations. |
Copied to clipboard
| Challenge: | In few-shot text classification, self-training relies on pseudo-labels to expand data, which has shown success, but can accumulate errors due to noisy pseudo-labeled data. |
| Approach: | They propose a method to mitigate noise in noisy pseudo-labeled data by applying superficial learning to noisy data and fine-tuning to less noisy data. |
| Outcome: | The proposed framework improves the classifier accuracy for few-shot text classification by 18.5% at most and 8% in average, compared with the state-of-the-art SSL baselines. |
Copied to clipboard
| Challenge: | Text-to-image (T2I) generation has the potential to advance knowledge democratization and education. |
| Approach: | They explore ways to harness T2I models for generating health knowledge flashcards . they curated a high-quality healthcare knowledge flash card dataset . |
| Outcome: | The proposed models can generate health knowledge flashcards with appealing images . the results show that the open-source models can be fine tuned to generate health content . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used in education, yet their default usefulness conflicts with pedagogical principles. |
| Approach: | They propose an adversarial student agent that they fine-tune to jailbreak LLM-based tutors and propose a benchmark to evaluate tutor robustness. |
| Outcome: | The proposed model fine-tunes to jailbreak LLM-based tutors, and shows that they perform well under adversarial student attacks. |
Copied to clipboard
| Challenge: | Chart generation requires strong visual design skills and precise coding capabilities that embed the desired visual properties into code. |
| Approach: | They propose a vision-language model-based multi-agent framework for effective automatic chart generation. |
| Outcome: | The proposed framework achieves a 5.2% improvement in the F1 score over the current best chart generation task. |
Copied to clipboard
| Challenge: | Existing Braille research focuses on isolated tasks while mixed-content Braille tasks face data scarcity and ambiguities. |
| Approach: | They propose a syntax tree-based augmentation method tailored for Braille data. |
| Outcome: | The proposed method improves Braille translation, formula-to-Braille conversion, and mixed-text translation. |
Copied to clipboard
| Challenge: | Large language models are often subjected to context-shifting behaviour, resulting in a lack of consistent and interpretable personality-aligned interactions. |
| Approach: | They propose to use two conversation agents to generate a discourse with an assigned personality from the OCEAN framework and then use multiple judge agents to infer original traits. |
| Outcome: | The proposed model is based on two conversation agents with a personality assigned from the OCEAN framework and then multiple judge agents to infer the original traits assigned. |
Copied to clipboard
| Challenge: | Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA). |
| Approach: | They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . |
| Outcome: | The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset . |
Copied to clipboard
| Challenge: | Persuasive automated dialogue systems are a popular way to influence people's behavior and decision making. |
| Approach: | They propose to use a context-aware persuasion strategy selection module to persult users . they also propose a persuasiveness prediction model to automatically evaluate the persuasiveness of generated text. |
| Outcome: | The proposed system can achieve better performance on several automated evaluation metrics than baseline models. |
Copied to clipboard
| Challenge: | Existing work on confidence estimation and calibration focuses on single-turn settings . existing work on multi-turn calibration ignores the risks and potential of multi-turned conversations . |
| Approach: | They propose a multi-turn calibration task that reframes calibration from a static property into a dynamic challenge central to reliable multi- turn conversations. |
| Outcome: | The proposed model minimizes ECE@T and leverages ConfChat to improve confidence . the proposed model preserves and even enhances model performance in multi-turn interactions. |
Copied to clipboard
| Challenge: | Recent surge in deep learning technologies has significantly accelerated research in this area. |
| Approach: | They propose a comprehensive summary of the relevant tasks in geometry problem solving and a review of related deep learning methods. |
| Outcome: | The proposed method is based on a systematic review of related methods and evaluation metrics and methods. |
Copied to clipboard
| Challenge: | evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up. |
| Approach: | They propose to define the student simulation task and benchmark a wide range of student simulation methods on these metrics. |
| Outcome: | The proposed evaluation metrics show that prompting strategies perform poorly on a real-world tutoring dialogue dataset. |
Copied to clipboard
| Challenge: | Existing methods for jailbreak ignore the semantic differences between categories of harmful questions, leading to inconsistent success rates and reduced overall attack effectiveness. |
| Approach: | They propose a category-aware jailbreak framework that incorporates the semantic category of harmful questions into prompt generation. |
| Outcome: | The proposed framework improves attack success rates and category alignment and achieves better cross-category robustness compared to the state-of-the-art (SOTA) baselines. |