TY - GEN
T1 - Input-splitting of large neural networks for power-efficient accelerator with resistive crossbar memory array
AU - Kim, Yulhwa
AU - Kim, Hyungjun
AU - Ahn, Daehyun
AU - Kim, Jae Joon
N1 - Publisher Copyright:
© 2018 Association for Computing Machinery.
PY - 2018/7/23
Y1 - 2018/7/23
N2 - Resistive Crossbar memory Arrays (RCA) have been gaining interest as a promising platform to implement Convolutional Neural Networks (CNN). One of the major challenges in RCA-based design is that the number of rows in an RCA is often smaller than the number of input neurons in a layer. Previous works used highresolution Analog-to-Digital Converters (ADCs) to compute the partial weighted sum in each array and merged partial sums from multiple arrays outside the RCAs. However, such approach suffers from significant power consumption due to the need for highresolution ADCs. In this paper, we propose a methodology to more efficiently construct a large CNN with multiple RCAs. By splitting the input feature map and retraining the CNN with proper initialization, we demonstrate that any CNN model can be represented with multiple arrays without using intermediate partial sums. The experimental results show that the ADC power of the proposed design is 32x smaller and the total chip power of the proposed design is 3x smaller than those of the baseline design.
AB - Resistive Crossbar memory Arrays (RCA) have been gaining interest as a promising platform to implement Convolutional Neural Networks (CNN). One of the major challenges in RCA-based design is that the number of rows in an RCA is often smaller than the number of input neurons in a layer. Previous works used highresolution Analog-to-Digital Converters (ADCs) to compute the partial weighted sum in each array and merged partial sums from multiple arrays outside the RCAs. However, such approach suffers from significant power consumption due to the need for highresolution ADCs. In this paper, we propose a methodology to more efficiently construct a large CNN with multiple RCAs. By splitting the input feature map and retraining the CNN with proper initialization, we demonstrate that any CNN model can be represented with multiple arrays without using intermediate partial sums. The experimental results show that the ADC power of the proposed design is 32x smaller and the total chip power of the proposed design is 3x smaller than those of the baseline design.
KW - Neural networks
KW - Resistive random-access memory
KW - Vector-matrix multiplication acceleration
UR - https://www.scopus.com/pages/publications/85051525217
U2 - 10.1145/3218603.3218605
DO - 10.1145/3218603.3218605
M3 - Conference contribution
AN - SCOPUS:85051525217
SN - 9781450357043
T3 - Proceedings of the International Symposium on Low Power Electronics and Design
BT - ISLPED 2018 - Proceedings of the 2018 International Symposium on Low Power Electronics and Design
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 23rd IEEE/ACM International Symposium on Low Power Electronics and Design, ISLPED 2018
Y2 - 23 July 2018 through 25 July 2018
ER -