Papers by Zhu Teng

7 papers
USB: A COMPREHENSIVE AND UNIFIED SAFETY EVALUATION BENCHMARK FOR MULTIMODAL LARGE LANGUAGE MODELS (2026.acl-long)

Copied to clipboard

Challenge: Existing safety benchmarks fail to provide reliable assessments due to limited risk coverage, insufficient scale and the oversight of complex modality combinations.
Approach: They propose a framework that covers 61 risk categories across four modality interactions to address this gap.
Outcome: The proposed framework covers 61 risk categories across four distinct modality interactions.
BPP-Search: Enhancing Tree of Thought Reasoning for Mathematical Modeling Problem Solving (2025.acl-long)

Copied to clipboard

Challenge: Existing datasets in operations research domain lack detailed annotations of the modeling process, focusing only on objective values.
Approach: They propose an annotation-based tree-of-thought tree-based reasoning algorithm that integrates reinforcement learning into a tree- of-though.
Outcome: The proposed algorithm outperforms state-of-the-art methods on StructuredOR, NL4OPT, and MAMO-ComplexLP datasets.
BLADE: Benchmarking Language Model Agents for Data-Driven Science (2024.findings-emnlp)

Copied to clipboard

Challenge: Language model-based agents can be used to conduct and support data-driven science, but evaluating them on open-ended tasks is challenging due to multiple valid approaches, partially correct steps, and different ways to express the same decisions.
Approach: They propose a benchmark to automatically evaluate agents’ multifaceted approaches to open-ended research questions.
Outcome: BLADE evaluates agents’ multifaceted approaches to open-ended research questions using data from 12 datasets and research questions drawn from existing scientific literature.
How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for fine-tuning large language models are not suitable for task-dependent tasks.
Approach: They propose a generalized self-imitation learning framework which aligns large language models with offline demonstration data.
Outcome: The proposed framework outperforms baselines in many challenging benchmarks . it is available on github.com/tengxiao1/GSIL .
QueryLink: Leveraging Query-Memory Alignment for Long-Term Reasoning in LLM Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to integrating external memory prioritize memory organization while overlooking a critical semantic gap between implicit, intent-driven queries and explicit, narrative-based memories.
Approach: They propose a framework that leverages Query-Memory Alignment to project both queries and memories into a shared semantic space.
Outcome: The proposed framework significantly outperforms SOTA methods on the LoCoMo and LongMemEval benchmarks and can be integrated as a plug-and-play component to boost existing vector-based systems like A-MEM.
Self-Constructed Context Decompilation with Fined-grained Alignment Enhancement (2024.findings-emnlp)

Copied to clipboard

Challenge: Decompilation is the process of converting compiled code back into a high-level programming language for analysis when source code is unavailable.
Approach: They propose two methods to improve decompilation performance without fine-tuning and fine-grained alignment enhancement to achieve further improvements.
Outcome: The proposed methods achieved a Re-Executability performance improvement of approximately 3.90% on the Decompile-Eval benchmark, establishing a new state-of-the-art performance of 52.41%.
Reinforcement Learning for Large Language Models via Group Preference Reward Shaping (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for fine-tuning Large Language Models (LLMs) are expensive and sensitive to reward model quality.
Approach: They propose a method that leverages preference-based comparisons rather than precise numerical rewards.
Outcome: Experiments show that GPRS outperforms critic-model-free RL algorithms on RLHF and reasoning tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations