Papers by Eric Zhu
RWKV: Reinventing RNNs for the Transformer Era (2023.findings-emnlp)
Copied to clipboard
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Leon Derczynski, Xingjian Du, Matteo Grella, Kranthi Gv, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartłomiej Koptyra, Hayden Lau, Jiaju Lin, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Guangyu Song, Xiangru Tang, Johan Wind, Stanisław Woźniak, Zhenyuan Zhang, Qinghua Zhou, Jian Zhu, Rui-Jie Zhu
| Challenge: | recurrent neural networks struggle to match the performance of Transformers due to limitations in parallelization and scalability. |
| Approach: | They propose a model architecture that combines the efficient parallelizable training of transformers with the efficient inference of RNNs. |
| Outcome: | The proposed model performs on par with similarly sized RNNs, suggesting future work can leverage this architecture to create more efficient models. |
Texar: A Modularized, Versatile, and Extensible Toolkit for Text Generation (P19-3)
Copied to clipboard
Zhiting Hu, Haoran Shi, Bowen Tan, Wentao Wang, Zichao Yang, Tiancheng Zhao, Junxian He, Lianhui Qin, Di Wang, Xuezhe Ma, Zhengzhong Liu, Xiaodan Liang, Wanrong Zhu, Devendra Sachan, Eric Xing
| Challenge: | Texar is an open-source text generation toolkit that supports a broad set of text generation tasks. |
| Approach: | They introduce Texar, an open-source text generation toolkit that supports text generation tasks. |
| Outcome: | Texar supports machine translation, summarization, dialog, content manipulation, and more. |
RedCoast: A Lightweight Tool to Automate Distributed Training of LLMs on Any GPU/TPUs (2024.naacl-demo)
Copied to clipboard
| Challenge: | Recent advances in machine learning (ML) are attributed to large language models (LLMs), but their escalating memory requirements require developers to partition a large model to distribute it across multiple GPUs or TPUs. |
| Approach: | They propose a lightweight and user-friendly tool to automate distributed training and inference for LLMs and to simplify ML pipeline development. |
| Outcome: | The proposed tool automates distributed training and inference for LLMs, and simplifies ML pipeline development. |
Embedding Imputation with Grounded Language Information (P19-1)
Copied to clipboard
| Challenge: | Existing approaches to embedding imputation use vector space properties or subword information to learn representations for rare or unseen words. |
| Approach: | They propose an online method to construct a knowledge graph from grounded information and an algorithm to map from the resulting graph to the space of the pre-trained embeddings. |
| Outcome: | The proposed method improves on a card-660 task by 11% and 17.8% respectively using GloVe embeddings. |
DAMP: Doubly Aligned Multilingual Parser for Task-Oriented Dialogue (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that multilingual models are less robust for semantic parsing compared to other tasks. |
| Approach: | They propose a constrained optimization technique to optimize multilingual parsing systems for multilingual use. |
| Outcome: | The proposed technique outperforms XLM-R and mT5-Large on three benchmarks and significantly outperformed other models. |
TED: A Pretrained Unsupervised Summarization Model with Theme Modeling and Denoising (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing abstractive summarization models ignore abundant unlabeled corpora resources . TED outperforms all unsupervised abstractive baselines on NYT, CNN/DM and English Gigaword datasets . |
| Approach: | They propose a transformer-based unsupervised text summarization system with pretraining on large-scale data. |
| Outcome: | The proposed system outperforms baseline models on NYT, CNN/DM and English Gigaword datasets with various document styles. |
Making the Most Out of the Limited Context Length: Predictive Power Varies with Clinical Note Type and Note Section (2023.acl-srw)
Copied to clipboard
| Challenge: | Clinical notes have a long time span over multiple long documents. |
| Approach: | They propose a framework to analyze clinical notes with high predictive power . they propose to combine different types of notes to improve performance . |
| Outcome: | The proposed framework could be used to extract information from clinical notes . it shows that the sample size can be optimized for large contexts . |