Compression of generative  pre-trained language models via quantization

Tao, C; Hou, L; Zhang, W; Shang, L; Jiang, X; Liu, Q; Luo, P; Wong, N

File Download

There are no files associated with this item.

Supplementary

Citations:
Appears in Collections:
- Computer Science: Conference papers
- Electrical & Electronic Engineering: Conference papers

Conference Paper: Compression of generative pre-trained language models via quantization

Title	Compression of generative pre-trained language models via quantization
Authors	Tao, C Hou, L Zhang, W Shang, L Jiang, X Liu, Q Luo, P Wong, N
Issue Date	2022
Publisher	Abbey Group.
Citation	The 60th Annual Meeting of the Association for Computational Linguistics (ACL), Dublin, Ireland & Online, 22-27 May, 2022 How to Cite?
Abstract	The increasing size of generative Pre-trained Language Models (PLMs) has greatly increased the demand for model compression. Despite various methods to compress BERT or its variants, there are few attempts to compress generative PLMs, and the underlying difficulty remains unclear. In this paper, we compress generative PLMs by quantization. We find that previous quantization methods fail on generative tasks due to the extit{homogeneous word embeddings} caused by reduced capacity, and extit{varied distribution of weights}. Correspondingly, we propose a token-level contrastive distillation to learn distinguishable word embeddings, and a module-wise dynamic scaling to make quantizers adaptive to different modules. Empirical results on various tasks show that our proposed method outperforms the state-of-the-art compression methods on generative PLMs by a clear margin. With comparable performance with the full-precision models, we achieve 14.4x and 13.4x compression rates on GPT-2 and BART, respectively.
Description	Outstanding Paper Award
Persistent Identifier	http://hdl.handle.net/10722/315547

DC Field	Value	Language
dc.contributor.author	Tao, C	-
dc.contributor.author	Hou, L	-
dc.contributor.author	Zhang, W	-
dc.contributor.author	Shang, L	-
dc.contributor.author	Jiang, X	-
dc.contributor.author	Liu, Q	-
dc.contributor.author	Luo, P	-
dc.contributor.author	Wong, N	-
dc.date.accessioned	2022-08-19T08:59:55Z	-
dc.date.available	2022-08-19T08:59:55Z	-
dc.date.issued	2022	-
dc.identifier.citation	The 60th Annual Meeting of the Association for Computational Linguistics (ACL), Dublin, Ireland & Online, 22-27 May, 2022	-
dc.identifier.uri	http://hdl.handle.net/10722/315547	-
dc.description	Outstanding Paper Award	-
dc.description.abstract	The increasing size of generative Pre-trained Language Models (PLMs) has greatly increased the demand for model compression. Despite various methods to compress BERT or its variants, there are few attempts to compress generative PLMs, and the underlying difficulty remains unclear. In this paper, we compress generative PLMs by quantization. We find that previous quantization methods fail on generative tasks due to the extit{homogeneous word embeddings} caused by reduced capacity, and extit{varied distribution of weights}. Correspondingly, we propose a token-level contrastive distillation to learn distinguishable word embeddings, and a module-wise dynamic scaling to make quantizers adaptive to different modules. Empirical results on various tasks show that our proposed method outperforms the state-of-the-art compression methods on generative PLMs by a clear margin. With comparable performance with the full-precision models, we achieve 14.4x and 13.4x compression rates on GPT-2 and BART, respectively.	-
dc.language	eng	-
dc.publisher	Abbey Group.	-
dc.relation.ispartof	The 60th Annual Conference of the Association for Computational Linguistics (ACL), Outstanding Paper Award	-
dc.title	Compression of generative pre-trained language models via quantization	-
dc.type	Conference_Paper	-
dc.identifier.email	Luo, P: pluo@hku.hk	-
dc.identifier.email	Wong, N: nwong@eee.hku.hk	-
dc.identifier.authority	Luo, P=rp02575	-
dc.identifier.authority	Wong, N=rp00190	-
dc.identifier.hkuros	335577	-
dc.publisher.place	Ireland	-

File Download

Supplementary

Conference Paper: Compression of generative pre-trained language models via quantization

Export via OAI-PMH Interface in XML Formats

OR

Export to Other Non-XML Formats