Papers by Weigao Sun
Native Hybrid Attention for Efficient Sequence Modeling (2026.acl-long)
Copied to clipboard
| Challenge: | Experimental results show that NHA surpasses Transformers on recall-intensive tasks. |
| Approach: | They propose a hybrid architecture of linear and full attention that integrates both into a unified layer design. |
| Outcome: | The proposed architecture surpasses Transformers and other hybrid baselines on recall-intensive and commonsense reasoning tasks. |
CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on “factual statements” that rephrase source materials, but ignore “cognitive statements” . evaluating and detecting "faithfulness hallucinations" remains challenging . |
| Approach: | They propose a framework to assess faithfulness of cognitive statements and introduce a dataset to scale easily across models. |
| Outcome: | The proposed framework assesses faithfulness of cognitive statements and scales easily across models. |
Scaling Laws for Linear Complexity Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing scaling laws for large language models are unclear, but they are useful for scalability. |
| Approach: | They propose scaling laws for linear complexity language models to establish a foundation for their scalability. |
| Outcome: | The proposed models demonstrate superior linguistic proficiency and knowledge retention. |
Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism (2026.acl-long)
Copied to clipboard
Yuhua Jiang, Shuang Cheng, Yihao Liu, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng, Feifei Gao, Biqing Qi, Bowen Zhou
| Challenge: | Existing models lack task-guided specialized memory mechanisms . specialized generalist models excel at general language tasks but struggle in specialized domains. |
| Approach: | They propose a specialized generalist model with specialized memory and updater that can optimize for specialized domains. |
| Outcome: | The proposed model matches or surpasses baselines on general benchmarks and achieves lowest perplexity across specialized domains. |