Papers by Zhixuan Yang
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have impressive capabilities across a wide range of domains, but their generalpurpose pre-training objectives often leave them illsuited for specialized applications such as healthcare. |
| Approach: | They propose a perplexity-aware data scaling law that establishes a predictive relationship between the perplexities of domain-specific data and the test loss. |
| Outcome: | Experiments on medical and general-domain benchmarks show that the proposed scaling law consistently identifies near-optimal training subsets with significantly reduced data consumption. |
Adaptive Prompt Optimization for Open-Ended Tasks: Uncertainty Preference as a Secondary Signal (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent training-free prompt optimizers treat performance as maximizing a single scalar score and ignore a second signal that the desired style is task dependent. |
| Approach: | They propose a semantic-entropy-based method that uses task uncertainty to guide prompt optimization by selecting high-entropicy candidates for creative tasks and low-energetic candidates for conservative ones. |
| Outcome: | The proposed method outperforms baselines on MT-Bench subsets and integrates easily into existing prompt optimizers. |
ATGL: An Adaptive-Threshold Global Loss for Document-level Relation Extraction (2026.acl-long)
Copied to clipboard
| Challenge: | Document-level relation extraction (DocRE) aims to determine which relations hold between a given entity pair in a document. |
| Approach: | They propose a document-level relation extraction paradigm that decouples existing losses into independent positive and negative losses, which interact solely with a shared threshold. |
| Outcome: | The proposed model outperforms existing models on four datasets and achieves state-of-the-art results. |
Inference-Time Language Model Alignment via Integrated Value Guidance (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models are fine-tuned to align with human preferences, but tuning large models is computationally intensive and complex. |
| Approach: | They propose a method that uses implicit and explicit value functions to guide language model decoding at token and chunk-level respectively. |
| Outcome: | The proposed method outperforms traditional methods and circumvents the complexities of fine-tuning. |
DEBAR: Mitigating Contextual Bias in Cross-Document Relation Extraction via Dual-Stream Decoupling (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods focus on sentence-level or singledocument settings, resulting in one-sided relation transfer contextual bias and incomplete reasoning chains. |
| Approach: | They propose a framework to explicitly decouple and preserve bidirectional bridge evidence and a dynamic loss optimization objective to separate head and tail contexts. |
| Outcome: | The proposed framework decouples and preserves bidirectional bridge evidence while capturing global dependencies through iterative message passing. |