ALWANN: Automatic Layer-Wise Approximation of Deep Neural Network Accelerators without Retraining

Mrázek, Vojtěch; Vašíček, Zdeněk; Sekanina, Lukáš; Hanif, Muhammad Abdullah; Shafique, Muhammad

doi:10.1109/iccad45719.2019.8942068

Cited by 93 publications

(113 citation statements)

References 28 publications

Supporting

Mentioning

113

Contrasting

Order By: Relevance

“…We select Neural Network inference as our testcase and we use the ResNet-8 network [59] (7 convolution layers) trained with the CIFAR-10 image classification dataset. The NN is quantized at 8 bits to avoid floating point operations [60]. First, we use RETSINA to generate two approximate reconfigurable 8-bit multipliers.…”

Section: Figure 12mentioning

confidence: 99%

“…The accuracy levels of RM1 and RM2 are selected targeting high inference accuracy. To achieve this, we used [31] and [60] to profile how the multiplier's error impacts ResNet's accuracy. Next, we extend [60] and integrate our reconfigurable multipliers RM1 and RM2.…”

Section: Figure 12mentioning

confidence: 99%

“…To achieve this, we used [31] and [60] to profile how the multiplier's error impacts ResNet's accuracy. Next, we extend [60] and integrate our reconfigurable multipliers RM1 and RM2. In [60], approximate multipliers are used to perform the multiplications of each convolution layer.…”

Section: Figure 12mentioning

confidence: 99%

“…Next, we extend [60] and integrate our reconfigurable multipliers RM1 and RM2. In [60], approximate multipliers are used to perform the multiplications of each convolution layer. However, since different layers have different accuracy requirement, [60] employs an heterogeneous architecture that comprises several fixed approximate multiplier types.…”

Section: Figure 12mentioning

confidence: 99%

See 3 more Smart Citations

Design Automation of Approximate Circuits With Runtime Reconfigurable Accuracy

2020

View full text Add to dashboard Cite

Leveraging the inherent error tolerance of a vast number of application domains that are rapidly growing, approximate computing arises as a design alternative to improve the efficiency of our computing systems by trading accuracy for energy savings. However, the requirement for computational accuracy is not fixed. Controlling the applied level of approximation dynamically at runtime is a key to effectively optimize energy, while still containing and bounding the induced errors at runtime. In this paper, we propose and implement an automatic and circuit independent design framework that generates approximate circuits with dynamically reconfigurable accuracy at runtime. The generated circuits feature varying accuracy levels, supporting also accurate execution. Extensive experimental evaluation, using industry strength flow and circuits, demonstrates that our generated approximate circuits improve the energy by up to 41% for 2% error bound and by 17.5% on average under a pessimistic scenario that assumes full accuracy requirement in the 33% of the runtime. To demonstrate further the efficiency of our framework, we considered two state-of-the-art technology libraries which are a 7nm conventional FinFET and an emerging technology that boosts performance at a high cost of increased dynamic power. INDEX TERMS Approximate computing, approximate design automation, dynamically reconfigurable accuracy, low power.

show abstract

Section: Figure 12mentioning

confidence: 99%

Section: Figure 12mentioning

confidence: 99%

Section: Figure 12mentioning

confidence: 99%

Section: Figure 12mentioning

confidence: 99%

See 2 more Smart Citations

Design Automation of Approximate Circuits With Runtime Reconfigurable Accuracy

2020

View full text Add to dashboard Cite

show abstract

“…The works presented in [10], [13], [20], [36], [37] had used logic minimization to create the optimal approximate multipliers for each network model. Logic minimization intentionally flips bits in the logic to reduce the size of the operators, and these techniques use heuristics to find the optimal targets.…”

Section: Related Workmentioning

confidence: 99%

Low-power implementation of Mitchell's approximate logarithmic multiplication for convolutional neural networks

Kim¹,

Barrio²,

Hermida³

et al. 2018

2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC)

View full text Add to dashboard Cite

This paper analyzes the effects of approximate multiplication when performing inferences on deep convolutional neural networks (CNNs). The approximate multiplication can reduce the cost of underlying circuits so that CNN inferences can be performed more efficiently in hardware accelerators. The study identifies the critical factors in the convolution, fully-connected, and batch normalization layers that allow more accurate CNN predictions despite the errors from approximate multiplication. The same factors also provide an arithmetic explanation of why bfloat16 multiplication performs well on CNNs. The experiments are performed with recognized network architectures to show that the approximate multipliers can produce predictions that are nearly as accurate as the FP32 references, without additional training. For example, the ResNet and Inception-v4 models with Mitch-w6 multiplication produces Top-5 errors that are within 0.2% compared to the FP32 references. A brief cost comparison of Mitch-w6 against bfloat16 is presented, where a MAC operation saves up to 80% of energy compared to the bfloat16 arithmetic. The most far-reaching contribution of this paper is the analytical justification that multiplications can be approximated while additions need to be exact in CNN MAC operations.

show abstract