1. School of Automation and Electrical Engineering, Dalian Jiaotong University, Dalian 116028, China
2. Tianjin Rail Transit Group Corporation, Tianjin 300380, China
3. School of Computer and Communication Engineering, Dalian Jiaotong University, Dalian 116028, China
| Abstract: | The performance of human pose estimation network model is gradually improved, and the over-deep network structure brings a large number of parameters and complex calculation. To solve these problems, the FastPose-Lite lightweight human pose estimation network is proposed, which is composed of GSE-ResNet feature extraction network, up-sampling DUC module and CBAM module. The basic GBNK module of GSE-ResNet feature extraction network is composed of Ghost module and SE module. On the one hand, Ghost module is proposed to replace the traditional convolutional module in order to reduce the number of parameters and calculation. On the other hand, in order to keep the performance of the network model unchanged, SE attention mechanism module is introduced. In order to enhance the processing ability of the network model to the feature information in both spatial and channel aspects, CBAM module is introduced into the up-sampling DUC module to reduce the loss in the up-sampling process. The experimental results on the COCO dataset show that the proposed FastPose-Lite reduce the number of parameters and calculation by 51.4% and 50.8% respectively compared with the FastPose network model. Compared with the common popular network models such as SHN, CPN, and SimpleBaseline, the FastPose-Lite network model not only has fewer parameters and computations, but also has higher prediction accuracy. |
| Keywords: | Human Pose Estimation; FastPose; Ghost Module; Attention Mechanism; Lightweight |
| DOI: | 10.57237/j.cst.2023.02.003 |
| [1] | Dalal N, Triggs B. Histograms of oriented gradients for human detection [C] // 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05). Ieee, 2005, 1: 886-893. |
| [2] | Lowe D G. Distinctive image features from scale-invariant keypoints [J]. International journal of computer vision, 2004, 60 (2): 91-110. |
| [3] | LeCun Y, Boser B, Denker J S, et al. Backpropagation applied to handwritten zip code recognition [J]. Neural computation, 1989, 1 (4): 541-551. |
| [4] | Wei S E, Ramakrishna V, Kanade T, et al. Convolutional pose machines [C] // Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. 2016: 4724-4732. |
| [5] | Newell A, Yang K, Deng J. Stacked hourglass networks for human pose estimation [C] // European conference on computer vision. Springer, Cham, 2016: 483-499. |
| [6] | Fang H S, Xie S, Tai Y W, et al. RMPE: Regional multi-person pose estimation [C] // Proceedings of the IEEE international conference on computer vision. 2017: 2334-2343. |
| [7] | Fang H S, Li J, Tang H, et al. AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-Time [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. |
| [8] | Cao Z, Simon T, Wei S E, et al. Realtime multi-person 2d pose estimation using part affinity fields [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 7291-7299. |
| [9] | Howard A G, Zhu M, Chen B, et al. Mobilenets: Efficient convolutional neural networks for mobile vision applications [J]. arXiv preprint arXiv: 1704.04861, 2017. |
| [10] | Sandler M, Howard A, Zhu M, et al. Mobilenetv2: Inverted residuals and linear bottlenecks [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 4510-4520. |
| [11] | Howard A, Sandler M, Chu G, et al. Searching for mobilenetv3 [C] // Proceedings of the IEEE/CVF international conference on computer vision. 2019: 1314-1324. |
| [12] | Hu J, Shen L, Sun G. Squeeze-and-excitation networks [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7132-7141. |
| [13] | Bochkovskiy A, Wang C Y, Liao H Y M. Yolov: Optimal speed and accuracy of object detection [J]. arXiv preprint arXiv: 2004.10934, 2020. |
| [14] | He K, Zhang X, Ren S, et al. Deep residual learning for image recognition [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778. |
| [15] | Woo S, Park J, Lee J Y, et al. Cbam: Convolutional block attention module [C] // Proceedings of the European conference on computer vision (ECCV). 2018: 3-19. |
| [16] | Han K, Wang Y, Tian Q, et al. Ghostnet: More features from cheap operations [C] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020: 1580-1589. |
| [17] | Wang P, Chen P, Yuan Y, et al. Understanding convolution for semantic segmentation [C] // 2018 IEEE winter conference on applications of computer vision (WACV). Ieee, 2018: 1451-1460. |
| [18] | Lin T Y, Maire M, Belongie S, et al. Microsoft coco: Common objects in context [C] // European conference on computer vision. Springer, Cham, 2014: 740-755. |
| [19] | Chen Y, Wang Z, Peng Y, et al. Cascaded pyramid network for multi-person pose estimation [C] // Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 7103-7112. |
| [20] | Xiao B, Wu H, Wei Y. Simple baselines for human pose estimation and tracking [C] // Proceedings of the European conference on computer vision (ECCV). 2018: 466-481. |
| [21] | Sun K, Xiao B, Liu D, et al. Deep high-resolution representation learning for human pose estimation [C] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 5693-5703. |
We invite active, qualified and high profile scientists and researchers to join as Editorial Board Members.
Join UsScholars with a strong interest in reviewing are invited to join the reviewer panel to ensure the quality of the research to be published.
Join Us