Papers by Ziling Cheng
Gated Differentiable Working Memory for Long-Context Language Modeling (2026.acl-long)
Copied to clipboard
Lingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge, Baolong Bi, Jiayu Yao, Jun Wan, Ziling Yin, Jiafeng Guo, Xueqi Cheng
| Challenge: | Long contexts break transformers, attention scores dilute, model cannot adapt to novel patterns at inference time. |
| Approach: | They propose a framework that gates the memory consolidation process by estimating Contextual Utility . they propose GDWM to maintain a form of working memory with constant contexts . |
| Outcome: | The proposed framework achieves comparable or superior performance on sparse-information tasks with 4 fewer gradient steps compared to uniform baselines. |
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are often evaluated on math word problems . however, such metrics conflate two distinct sub-skills: abstract formulation and arithmetic computation. |
| Approach: | They propose to use Final-answer-based metrics to evaluate large language models on math word problems to conflate two distinct sub-skills: abstract formulation and arithmetic computation. |
| Outcome: | The proposed model performance is bottlenecked by arithmetic computation and not abstract formulation, the study shows. |
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that LLMs reproduce training artifacts, exploit spurious correlations, and fail when faced with distribution shifts. |
| Approach: | They examine irrelevant context hallucinations in which models integrate misleading contextual cues into their predictions. |
| Outcome: | The proposed model errors are reflected in the model's internal computations, and they are consistent with previous studies. |
Varta: A Large-Scale Headline-Generation Dataset for Indic Languages (2023.findings-acl)
Copied to clipboard
| Challenge: | Varta dataset includes more than 41 million pairs of headlines and articles in 14 different Indic languages (and English) |
| Approach: | They present a large-scale multilingual dataset for headline generation in Indic languages. |
| Outcome: | The Varta dataset includes more than 41 million pairs of headlines and articles in 14 different Indic languages (and English) the data can be used to train strong language models that outperform competitive baselines in both NLU and NLG benchmarks. |