Papers by Ziyao Xu

11 papers
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for large language models fail to capture complex interplay between functionality and security.
Approach: They propose a benchmark for secure code generation constructed from real-world, high-risk Java repositories.
Outcome: The proposed benchmarks highlight the gap between functional and secure code generation in LLMs.
Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion (2025.naacl-long)

Copied to clipboard

Challenge: Existing frameworks for evaluating the decomposition and composition capabilities of large language models (LLMs) in N2F are inadequate, and there are errors that can be attributed to deficiencies in natural language understanding and the learning and use of symbolic systems.
Approach: They propose a framework that semi-automatically performs sample and task construction . main findings include that LLMs are deficient in both decomposition and composition .
Outcome: The proposed framework evaluates the most advanced LLMs on a variety of common formal languages.
MC2: A Minimum-Coverage and Dataset-Agnostic Framework for Compositional Generalization of LLMs on Semantic Parsing (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing research relies on dataset-specific designs or a large number of samples to improve compositional generalization of large language models (LLMs) .
Approach: They propose a minimum-coverage framework that can help LLMs achieve compositional generalization by selecting and organizing samples that satisfy the primitive coverage.
Outcome: The proposed framework can improve compositional generalization on different parsing datasets in the minimum-coverage setting.
Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective (2026.acl-long)

Copied to clipboard

Challenge: Compositional generalization tests focus on output results without considering sample compositionality, resulting in explainability defects.
Approach: They propose a rule-generation perspective for compositionality estimation for LLMs that requires LLM to generate a program as rules for dataset mapping and provides estimates of compositionality using complexity-based theory.
Outcome: The proposed model provides estimates of the compositionality of LLMs using complexity-based theory on a string-to-grid task.
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to improve self-correction performance of Large Language Models are based on intrinsic selfcorrectione, which allows the model to check and revise its selfgenerated answers without external feedback.
Approach: They propose to decompose the self-correction capability into confidence and critique capabilities and a metric for overall self-corretion capability evaluation.
Outcome: The proposed method outperforms vanilla SFT and achieves much higher accuracy after self-correction.
SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on compositional generalization in data-to-text generation focus on one manifestation, Systematicity, Productivity, Order invariance, and Rule learnability.
Approach: They propose a method for evaluation of compositional generalization in data-to-text generation that includes four aspects of manifestations and allows high-quality evaluation without additional manual annotations.
Outcome: The proposed method is based on two datasets and evaluates existing language models including LLMs.
Multi-Layer Pseudo-Siamese Biaffine Model for Dependency Parsing (2022.coling-1)

Copied to clipboard

Challenge: Existing work only uses biaffine method at the end of the dependency parser as a scorer, and its application in multi-layer form is ignored.
Approach: They propose a multi-layer pseudo-Siamese biaffine model for neural dependency parsing that uses biaffin method as a scorer and a biaffin module to construct arc weight matrix.
Outcome: The proposed model achieves state-of-the-art on PTB, CTB, and UD datasets with low efficiency loss.
EmoPrompt-ECPE: Emotion Knowledge-aware Prompt-tuning for Emotion-Cause Pair Extraction (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for Emotion-cause pair extraction (ECPE) do not distinguish between the emotion-caused pairs that belong to different types of emotions, limiting their applicability.
Approach: They propose an Emotion-cause pair extraction method which integrates the implicit knowledge of cause clauses into a prompt template and extends the emotion labels to categories with an external emotion word base.
Outcome: The proposed method extracts all potential emotion clauses and corresponding cause clauses from unannotated documents.
A Probabilistic Inference Scaling Theory for LLM Self-Correction (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated the capability to refine their generated answers through self-correction, enabling continuous performance improvement over multiple rounds.
Approach: They propose a probabilistic theory to model the dynamics of accuracy change and explain performance improvements observed in multi-round self-correction.
Outcome: The proposed model can predict accuracy curves and improve accuracy over multiple rounds.
Ideology Takes Multiple Looks: A High-Quality Dataset for Multifaceted Ideology Detection (2023.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for the ID task only label a text as ideologically left- or right-leaning as a whole, regardless whether the text containing one or more different issues.
Approach: They construct an ideological schema for a multifaceted ideology detection task using MITweet and an English Twitter dataset.
Outcome: The proposed task uses a MITweet dataset with 12,594 English Twitter posts, each annotated with a Relevance and an Ideology label for all twelve facets.
Token-level Preference Self-Alignment Optimization for Multi-style Outline Controllable Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing attempts to outline generation are limited by response pair requirements and substantial computation costs.
Approach: They propose a token-level preference self-alignment optimization for outline controllable generation that extends the Bradley-Terry model from pair-wise to list-wise comparison.
Outcome: The proposed method outperforms existing methods by 19.28% in performance while requiring only 56.25% training time.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations