Adversarial Math Word Problem Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized the educational landscape due to the great improvements in their natural language generation and problem-solving capabilities.
Approach: They propose a cost-effective approach to attack large language models using abstract syntax trees to generate adversarial examples that preserve the structure and difficulty of the original questions aimed for assessment.
Outcome: The proposed approach significantly degrades students' math problem-solving ability on open- and closed-source LLMs.

Similar Papers

Vulnerabilities of Large Language Models to Adversarial Attacks (2024.acl-tutorials)

Copied to clipboard

Challenge: This tutorial focuses on the vulnerabilities of Large Language Models to adversarial attacks . the tutorial lays the foundation by explaining safety-aligned models and concepts in cybersecurity .
Approach: This tutorial lays the foundation by explaining safety-aligned LLMs and concepts in cybersecurity.
Outcome: The tutorial lays the foundation by explaining safety-aligned models and concepts in cybersecurity.
Evaluating the Validity of Word-level Adversarial Attacks with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing adversarial examples can generate invalid adversarials due to significant changes in semantic meanings compared to their originals.
Approach: They propose to use a large language model to evaluate adversarial examples by semantic constraints.
Outcome: The proposed method can generate valid adversarial examples even when they are not equipped with semantic constraints.
Adversarial Examples for Evaluating Math Word Problem Solvers (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing MWP solvers do not understand language and its relation with numbers, and their accuracy is unclear.
Approach: They propose two methods to generate adversarial attacks to evaluate the robustness of existing MWP solvers.
Outcome: The proposed method reduces the accuracy of existing MWP solvers by over 40% on two benchmark datasets.
Unveiling the Achilles’ Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have highlighted various neural metrics that align well with human evaluations.
Approach: They propose a black-box adversarial framework that generates strong disagreements between human and victim evaluators.
Outcome: The proposed framework can significantly improve the performance of human and victim evaluators.
Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are used to review academic papers, but are susceptible to textual adversarial attacks.
Approach: They evaluate the robustness of large language models as automated reviewers in the presence of adversarial attacks.
Outcome: The proposed model is robust against textual adversarial attacks, the authors argue . their findings highlight the importance of addressing adversarials to ensure integrity of scholarly communication.
Close or Cloze? Assessing the Robustness of Large Language Models to Adversarial Perturbations via Word Recovery (2025.coling-main)

Copied to clipboard

Challenge: Existing models implicitly recover the original text, but it is unclear when they rely on context and when they implicitly do so.
Approach: They propose to use a dictionary to recover adversarial words by using a phonetic, typo, and visual attack to study word recovery performance.
Outcome: The proposed model outperforms open-source models on hateful, offensive, and toxic classification tasks.
Large Language Models for Mathematical Reasoning: Progresses and Challenges (2024.eacl-srw)

Copied to clipboard

Challenge: a survey examines the landscape of mathematical problem-solving techniques . large language models have proven to be potent assets in unraveling nuances of mathematical reasoning .
Approach: They examine the evolution of Large Language Models (LLMs) for solving mathematical problems . they examine the spectrum of LLM-oriented techniques proposed for solving math problems - and their challenges .
Outcome: The survey examines the spectrum of proposed LLM-oriented techniques in solving math problems.
Can LLMs simulate the same correct solutions to free-response math problems as real students? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have explored modeling student mistakes, but lack of understanding of how they generate correct solutions.
Approach: They compare distribution of correct solutions produced by four large language models with students' responses to free-response problems.
Outcome: The proposed model can generate correct solutions that represent student responses to free-response problems.
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities (2024.naacl-long)

Copied to clipboard

Challenge: Rapid progress in open-source Large Language Models (LLMs) is driving AI development, but lacks sufficient trustworthiness to detect and mitigate adversarial demonstrations.
Approach: They propose an extended Chain of Utterances-based (CoU) prompting strategy to attack open-source LLMs.
Outcome: The proposed attack strategy is based on malicious demonstrations and toxicity tests on open-source models.
What Makes Math Word Problems Challenging for LLMs? (2024.findings-naacl)

Copied to clipboard

Challenge: Experiments show that even quite powerful LLMs are still challenged by MWPs.
Approach: They propose to analyze what makes math word problems (MWPs) in English challenging for large language models (LLMs).
Outcome: The proposed model can handle a range of core NLP tasks, but it has emergent abilities, such as ability to solve mathematical puzzles.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations