Quantization of Weights of Neural Networks with Negligible Decreasing of Prediction Accuracy

Perić, Zoran; Denić, Bojan; Savić, Milan; Dinčić, Milan; Mihajlov, Darko

dc.contributor.author	Perić, Zoran
dc.contributor.author	Denić, Bojan
dc.contributor.author	Savić, Milan
dc.contributor.author	Dinčić, Milan
dc.contributor.author	Mihajlov, Darko
dc.date.accessioned	2023-04-11T09:42:33Z
dc.date.available	2023-04-11T09:42:33Z
dc.date.issued	2021-09-24
dc.identifier.citation	III44006	en_US
dc.identifier.uri	https://platon.pr.ac.rs/handle/123456789/1187
dc.description.abstract	Quantization and compression of neural network parameters using the uniform scalar quantization is carried out in this paper. The attractiveness of the uniform scalar quantizer is reflected in a low complexity and relatively good performance, making it the most popular quantization model. We present a design approach for the memoryless Laplacian source with zero-mean and unit variance, which is based on iterative rule and uses the minimal mean-squared error distortion as a performance criterion. In addition, we derive closed-form expressions for SQNR (Signal to Quantization Noise Ratio) in a wide dynamic range of variance of input data. To show effectiveness on real data, the proposed quantizer is used to compress the weights of neural networks using bit rates from 9 to 16 bps (bits/sample) instead of standardly used 32 bps full precision bit rate. The impact of weights compression on the NN (neural network) performance is analyzed, indicating good matching with the theoretical results and showing negligible decreasing of the prediction accuracy of the NN even in the case of high variance-mismatch between the variance of NN weights and the variance used for the design of quantizer, if the value of the bit-rate is properly chosen according to the rule proposed in the paper. The proposed method could be possibly applied in some of the edge-computing frameworks, as simple uniform quantization models contribute to faster inference and data transmission.	en_US
dc.language.iso	en_US	en_US
dc.publisher	Kaunas University of Technology	en_US
dc.title	Quantization of Weights of Neural Networks with Negligible Decreasing of Prediction Accuracy	en_US
dc.title.alternative	Information Technology and Control	en_US
dc.type	clanak-u-casopisu	en_US
dc.description.version	publishedVersion	en_US
dc.identifier.doi	http://dx.doi.org/10.5755/j01.itc.50.3.28468
dc.citation.volume	50
dc.citation.issue	3
dc.subject.keywords	Uniform scalar quantization	en_US
dc.subject.keywords	variance-mismatch quantization	en_US
dc.subject.keywords	Laplacian distribution	en_US
dc.subject.keywords	quantized neural network	en_US
dc.subject.keywords	multilayer perceptron	en_US
dc.subject.keywords	MNIST database	en_US
dc.type.mCategory	M23	en_US
dc.type.mCategory	openAccess	en_US
dc.type.mCategory	M23	en_US
dc.type.mCategory	openAccess	en_US
dc.identifier.ISSN	1392-124X

Документи

Име:: 28468-Article Text-102355-1-10 ...
Величина:: 71.43Kb
Формат:: PDF

Отварање

Овај рад се појављује у следећим колекцијама

Главна колекција / Main Collection

Приказ основних података о документу