Publication

Scaling Pre-trained Language Models to Deeper via Parameter-efficient Architecture

返回学术发表
Scaling Pre-trained Language Models to Deeper via Parameter-efficient Architecture 论文配图

Peiyu Liu*, Ze-Feng Gao*, Yushuo Chen, Wayne Xin Zhao#, Ji-Rong Wen

Arxiv, 2023

摘要

In this paper, we propose a highly parameter-efficient approach to scaling pre-trained language models (PLMs) to a deeper model depth based on matrix product operator (MPO) decomposition, which shares the central tensor across all layers to reduce model size while keeping layer-specific auxiliary tensors and adapters for flexible adaptation.