Papers by Zhaochen Su
Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic Change (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to improve neural language models perform poorly on emerging data. |
| Approach: | They propose a lexical-level masking strategy to post-train a neural language model using static data from past years. |
| Outcome: | The proposed method outperforms existing methods on two pre-trained language models, two classification tasks, and four benchmark datasets. |
Efficient Continue Training of Temporal Language Model with Structural Information (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing temporal language models are limited by the superficial temporal information brought by timestamps, which fails to learn the inherent changes of linguistic components. |
| Approach: | They propose a method that captures syntactically changed tokens and captures the relationship between the time prefix and tokens. |
| Outcome: | The proposed method outperforms existing temporal language models on two datasets and three tasks. |
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for directing language model outputs are limited in their accuracy due to a distributional gap . existing methods train static value functions on trajectories sampled exclusively from the base policy . |
| Approach: | They propose a framework to bridge a distributional gap in the accuracy of value functions . they propose RLHF to align language models with human values and task requirements . |
| Outcome: | The proposed framework reduces computational costs and improves value function accuracy by leveraging principled value function optimization. |
Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning? (2024.acl-long)
Copied to clipboard
| Challenge: | Current temporal reasoning datasets are limited to questions about single or isolated events, falling short in mirroring the realistic temporal characteristics involving concurrent nature and intricate temporal interconnections. |
| Approach: | They propose a co-temporal Question Answering benchmark that contains four co-time scenarios with 4,748 samples for evaluating the co-timing abilities of large language models. |
| Outcome: | The proposed benchmarks show that current LLMs struggle on CoTempQA tasks even when enhanced with Chain of Thought methodologies. |
PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models (2025.acl-long)
Copied to clipboard
| Challenge: | Recent large language models (LLMs) have achieved significant performance in complex reasoning tasks such as mathematics and code generation. |
| Approach: | They propose a process-level benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs. |
| Outcome: | The proposed model measures the accuracy, soundness, and sensitivity of 25 models across open-source and closed-source large language models. |
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations of Large Language Models (LLMs) focus on task completion, but neglect a crucial capability: the ability to devise and adjust cost-optimal plans in response to changing environments. |
| Approach: | They propose a scalable, cost-centric benchmark to evaluate agents’ economic reasoning and replanning abilities. |
| Outcome: | Evaluating leading open-sourced and proprietary models on CostBench reveals a substantial gap in cost-aware planning . |
SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies focus on the text modality or are limited to specific tasks. |
| Approach: | They propose a framework to teach Large Vision-Language Models to selectively utilize retrieved information and improve their robustness against irrelevant or misleading references. |
| Outcome: | The proposed framework improves LVLMs’ ability to utilize retrieved multimodal references and their robustness against irrelevant or misleading information. |