Neuron-by-Neuron Quantization for Efficient Low-Bit QNN Training

Sher, Artem; Trusov, Anton; Limonova, Elena; Nikolaev, Dmitry; Arlazarov, Vladimir V.

Neuron-by-Neuron Quantization for Efficient Low-Bit QNN Training

Artem Sher (), Anton Trusov, Elena Limonova, Dmitry Nikolaev and Vladimir V. Arlazarov
Additional contact information
Artem Sher: Phystech School of Applied Mathematics and Informatics, Moscow Institute of Physics and Technology, 141701 Moscow, Russia
Anton Trusov: Phystech School of Applied Mathematics and Informatics, Moscow Institute of Physics and Technology, 141701 Moscow, Russia
Elena Limonova: Smart Engines Service LLC, 117312 Moscow, Russia
Dmitry Nikolaev: Smart Engines Service LLC, 117312 Moscow, Russia
Vladimir V. Arlazarov: Smart Engines Service LLC, 117312 Moscow, Russia

Mathematics, 2023, vol. 11, issue 9, 1-17

Abstract: Quantized neural networks (QNNs) are widely used to achieve computationally efficient solutions to recognition problems. Overall, eight-bit QNNs have almost the same accuracy as full-precision networks, but working several times faster. However, the networks with lower quantization levels demonstrate inferior accuracy in comparison to their classical analogs. To solve this issue, a number of quantization-aware training (QAT) approaches were proposed. In this paper, we study QAT approaches for two- to eight-bit linear quantization schemes and propose a new combined QAT approach: neuron-by-neuron quantization with straight-through estimator (STE) gradient forwarding. It is suitable for quantizations with two- to eight-bit widths and eliminates significant accuracy drops during training, which results in better accuracy of the final QNN. We experimentally evaluate our approach on CIFAR-10 and ImageNet classification and show that it is comparable to other approaches for four to eight bits and outperforms some of them for two to three bits while being easier to implement. For example, the proposed approach to three-bit quantization of the CIFAR-10 dataset results in 73.2% accuracy, while baseline direct and layer-by-layer result in 71.4% and 67.2% accuracy, respectively. The results for two-bit quantization for ResNet18 on the ImageNet dataset are 63.69% for our approach and 61.55% for the direct baseline.

Keywords: quantized neural network; low-bit quantization; layer-by-layer; neuron-by-neuron training (search for similar items in EconPapers)
JEL-codes: C (search for similar items in EconPapers)
Date: 2023
References: View references in EconPapers View complete reference list from CitEc
Citations:

Downloads: (external link)
https://www.mdpi.com/2227-7390/11/9/2112/pdf (application/pdf)
https://www.mdpi.com/2227-7390/11/9/2112/ (text/html)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:gam:jmathe:v:11:y:2023:i:9:p:2112-:d:1136534

Access Statistics for this article

Mathematics is currently edited by Ms. Emma He

More articles in Mathematics from MDPI
Bibliographic data for series maintained by MDPI Indexing Manager ().