Publication

Enhancing Scalability of Pre-trained Language Models via Efficient Parameter Sharing

返回学术发表
Enhancing Scalability of Pre-trained Language Models via Efficient Parameter Sharing 论文配图

Peiyu Liu*, Ze-Feng Gao*, Yushuo Chen, Wayne Xin Zhao#, Ji-Rong Wen

Association for Computational Linguistics: EMNLP 2023, 2023

摘要

In this paper, we propose a parameter-efficient pre-training approach that utilizes matrix decomposition and parameter-sharing strategies to scale PLMs. Extensive experiments have demonstrated the effectiveness of our proposed model in reducing the model size and achieving highly competitive performance (i.e. with fewer parameters than BERT-base, we successfully scale the model depth by a factor of 4x and even achieve 0.1 points higher than BERT-large for GLUE score).