Papers by Zhe Wei

15 papers
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus more on end-to-end performance, but neglect the underlying principles of knowledge acquisition and generalization.
Approach: They propose a benchmark specifically designed to explore the problem-solving principles by decomposing 6.5K visual math problems into 10.9K step-level questions for evaluation.
Outcome: The proposed benchmark covers 6.5K visual math problems and 10.9K step-level questions spanning 5 layers of knowledge granularity and 67 hierarchical knowledge concepts.
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Despite significant efforts in safety alignment, large language models (LLMs) such as GPT-4 and LLaMA 3 remain vulnerable to jailbreak attacks that can induce harmful behaviors.
Approach: They propose a feature extraction method to extract sample-agnostic features from benign datasets in the form of adversarial suffixes and propose 'suffix maybe features' they show that adversarials generated from jailbreak attacks may contain meaningful features, i.e. appending the same suffix to different prompts results in responses exhibiting specific characteristics.
Outcome: The proposed method extracts sample-agnostic features from benign datasets and shows that they may contain meaningful features.
Syllogistic Reasoning for Legal Judgment Analysis (2023.emnlp-main)

Copied to clipboard

Challenge: Legal judgment assistants are developing fast due to impressive progress of large language models.
Approach: They construct and manually correct a syllogistic reasoning dataset for legal judgment analysis using large language models as benchmarks.
Outcome: The proposed dataset contains 11,239 criminal cases covering 4 criminal elements, 80 charges and 124 articles.
A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News (2024.lrec-main)

Copied to clipboard

Challenge: a new dataset is being developed to enrich resources for sign language research . the dataset is 16.07 hours of sign videos of two signers with a vocabulary of 6,515 glosses and 2,850 Chinese characters or 18K Chinese words.
Approach: They introduce a new Hong Kong sign language dataset called TVB-HKSL-News . the dataset is collected from a TV news program and contains sign videos . they aim to support research in sign language recognition and translation .
Outcome: The proposed dataset supports sign language recognition and translation research in Hong Kong . it consists of 16.07 hours of sign videos of two signers with a vocabulary of 6,515 glosses and 2,850 Chinese characters or 18K Chinese words .
Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom’s Taxonomy (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods suffer from cognitive dimensional simplification and methodological unreliability due to the ”LLM-as-a-Judge” approach.
Approach: They propose a six-tiered benchmark that evaluates ASG systems by prioritizing deterministic algorithms and introducing a GRADE approach for abstract abilities.
Outcome: The proposed method provides the ASG field with a systematic, reproducible, and theoretically grounded benchmark to guide future research.
Feedback Is The Key for Automated Survey Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a promising foundation for literature surveys, but guiding them to generate accurate, reliable content remains a fundamental challenge.
Approach: They propose a feedback-driven framework that incorporates feedback across three dimensions: outline feedback for structural clarity, citation feedback for evidence validation, and content feedback for readability and analytical depth.
Outcome: The proposed framework significantly improves both citation and content quality, demonstrating feedback as the critical mechanism for automatic survey generation.
Exploiting Intrinsic Multilateral Logical Rules for Weakly Supervised Natural Language Video Localization (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for WS-NLVL rarely consider complex temporal relations enclosing the language query, yielding illogical predictions.
Approach: They propose a plug-and-play method to exploit temporal relations and logical rules for WS-NLVL.
Outcome: The proposed method is able to retrieve the moment corresponding to a language query in a video with only video-language pairs utilized during training.
Do Influence Functions Work on Large Language Models? (2025.findings-emnlp)

Copied to clipboard

Challenge: Influence functions are important for quantifying the impact of individual training data points on a model’s predictions.
Approach: They conduct a systematic study to address a key question: do influence functions work on large language models?
Outcome: The influence functions perform poorly across multiple tasks and are therefore unsuitable for large language models.
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing defense methods focus on detecting harmful prompts or reducing the likelihood of harmful responses.
Approach: They propose a layer-specific editing method to align LLMs to harmful prompts by supervised fine-tuning and reinforcement learning.
Outcome: The proposed method improves the performance of large language models against jailbreak attacks while maintaining performance on benign prompts.
JTAV: Jointly Learning Social Media Content Representation by Fusing Textual, Acoustic, and Visual Features (C18-1)

Copied to clipboard

Challenge: Existing studies on learning social media content focus on single modal or bi-modal learning, but this approach is non-trivial and challenging because content is multi-modal and involves several types of data, including text, audio, and image.
Approach: They propose to combine textual, acoustic, and visual information to learn social media content by fusing them jointly.
Outcome: The proposed model outperforms the state-of-the-art approaches on real-world datasets by a large margin.
UER: An Open-Source Toolkit for Pre-training Models (D19-3)

Copied to clipboard

Challenge: Existing work on pre-training models have shown that it is important to use a framework to deploy various pre- training models efficiently.
Approach: They propose an assemble-on-demand pre-training toolkit that assembles pre-trained models on demand and encapsulates them with rich modules.
Outcome: The proposed framework can reproduce state-of-the-art models or develop models that remain unexplored.
Attention Basin: Why Contextual Position Matters in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are sensitive to the contextual position of information in input.
Approach: They introduce Attention-Driven Reranking (AttnRank) which estimates a model’s intrinsic positional attention preferences using a small calibration set and reorders retrieved documents or few-shot examples to align the most salient content with these high-attention positions.
Outcome: Experiments on multi-hop QA and few-shot in-context learning tasks show that AttnRank achieves substantial improvements across 10 large language models of varying architectures and scales, without modifying model parameters or training procedures.
KnowLA: Enhancing Parameter-efficient Finetuning with Knowledgeable Adaptation (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for parameter-efficient finetuning (PEFT) are limited and only finetune a small number of parameters using limited instruction data.
Approach: They propose a method that inserts an adaptation layer into an LLM to integrate embeddings of entities appearing in the input text.
Outcome: The proposed method can activate parameterized knowledge in an LLM without changing its parameters or input prompts.
iPET: An Interactive Emotional Companion Dialogue System with LLM-Powered Virtual Pet World Simulation (2025.acl-demo)

Copied to clipboard

Challenge: Existing approaches to role-playing emotional companion products lack sustained personalization and contextual adaptability, limiting their effectiveness in real-world settings.
Approach: They propose a virtual pet agent that can enhance user engagement through rich, dynamic pet behaviors and interactions tailored to individual preferences.
Outcome: The proposed system has been deployed in a real-world, non-commercial product for 200 days and has demonstrated its effectiveness in practical applications.
Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing safeguards relying on pre-filtering or fine-tuning are costly and diminish overall utility.
Approach: They propose a lightweight method that leverages LVLMs’ inherent multimodal alignment for zero-shot toxic image detection.
Outcome: The proposed method achieves a 66.9% defense success rate with only 3.2% false positive rate and 7.2% overhead.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations