College of Computer Science, Hunan University of Technology, Zhuzhou 412007, China
| Abstract: | Object detection is a crucial problem in computer vision that involves accurately locating objects in images or videos to identify instances to be detected. It has various applications, including facial detection, intelligent driving assistance, and satellite remote sensing detection. This review aims to aid researchers in quickly comprehending object detection. It covers the concept of object detection algorithms, analyzes the development history of object detection algorithms, and elaborates on the evolution process of object detection algorithms from independent development to combination with deep learning technology. The article divides object detection based on whether anchor boxes are generated during the detection process. It discusses the types of anchor boxes and non-anchor boxes and analyzes the current research status of single-stage and two-stage object detection algorithms. It also summarizes classic model structures in the development process of object detection algorithms and introduces difficulties in object detection. The article aims to solve the problem of object detection with a small number of samples and without detailed annotations during the training process. It summarizes and compares the advantages and disadvantages of various classic models, mainstream datasets, and evaluation indicators. Additionally, it looks forward to current challenges and future development directions in the field of object detection. |
| Keywords: | Computer Vision; Deep Learning; Object Detection |
| DOI: | 10.57237/j.cst.2023.02.006 |
| [1] | Turk M A, Pentland A P. Face recognition using eigenfaces [C] // Proceedings. 1991 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE Computer Society, 1991: 586-591. |
| [2] | Lowe D G. Distinctive image features from scale-invariant keypoints [J]. International Journal of Computer Vision, 2004, 60: 91-110. |
| [3] | Dalal N, Triggs B. Histograms of oriented gradients for human detection [C] // 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05). Ieee, 2005, 1: 886-893. |
| [4] | Balcazar J L, Dai Y, Watanabe O. Provably fast training algorithms for support vector machines [C] // Proceedings 2001 IEEE International Conference on Data Mining. IEEE, 2001: 43-50. |
| [5] | Neubeck A, Van Gool L. Efficient non-maximum suppression [C] // 18th International Conference on Pattern Recognition (ICPR'06). IEEE, 2006, 3: 850-855. |
| [6] | Wang P, Shen C, Barnes N, et al. Fast and robust object detection using asymmetric totally corrective boosting [J]. IEEE Transactions on Neural Networks and Learning Systems, 2011, 23 (1): 33-46. |
| [7] | Krizhevsky A, Sutskever I, Hinton G E. Imagenet classification with deep convolutional neural networks [J]. Communications of the ACM, 2017, 60 (6): 84-90. |
| [8] | Shin H C, Roth H R, Gao M, et al. Deep convolutional neural networks for computer-aided detection: CNN architectures, dataset characteristics and transfer learning [J]. IEEE Transactions on Medical Imaging, 2016, 35 (5): 1285-1298. |
| [9] | Girshick R, Donahue J, Darrell T, et al. Rich feature hierarchies for accurate object detection and semantic segmentation [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2014: 580-587. |
| [10] | Uijlings J R R, Van De Sande K E A, Gevers T, et al. Selective search for object recognition [J]. International Journal of Computer Vision, 2013, 104: 154-171. |
| [11] | Rezatofighi H, Tsoi N, Gwak J Y, et al. Generalized intersection over union: A metric and a loss for bounding box regression [C] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019: 658-666. |
| [12] | Girshick R. Fast r-cnn [C] // Proceedings of the IEEE International Conference on Computer Vision. 2015: 1440-1448. |
| [13] | Luce R D. Individual choice behavior: A theoretical analysis [Z]. Courier Corporation, 2012. |
| [14] | Ren S, He K, Girshick R, et al. Faster r-cnn: Towards real-time object detection with region proposal networks [J]. Advances in Neural Information Processing Systems, 2015, 28. |
| [15] | Redmon J, Divvala S, Girshick R, et al. You only look once: Unified, real-time object detection [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016: 779-788. |
| [16] | Redmon J, Farhadi A. YOLO9000: better, faster, stronger [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 7263-7271. |
| [17] | Szegedy C, Liu W, Jia Y, et al. Going deeper with convolutions [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015: 1-9. |
| [18] | Ioffe S, Szegedy C. Batch normalization: Accelerating deep network training by reducing internal covariate shift [C] // International Conference on Machine Learning. pmlr, 2015: 448-456. |
| [19] | MacQueen J. Some methods for classification and analysis of multivariate observations [C] // Proc. 5th Berkeley Symposium on Math., Stat., and Prob. 1965: 281. |
| [20] | Redmon J, Farhadi A. Yolov3: An incremental improvement [J]. arXiv preprint arXiv: 1804.02767, 2018. |
| [21] | Bochkovskiy A, Wang C Y, Liao H Y M. Yolov4: Optimal speed and accuracy of object detection [J]. arXiv preprint arXiv: 2004.10934, 2020. |
| [22] | He K, Zhang X, Ren S, et al. Spatial pyramid pooling in deep convolutional networks for visual recognition [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015, 37 (9): 1904-1916. |
| [23] | Liu S, Qi L, Qin H, et al. Path aggregation network for instance segmentation [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018: 8759-8768. |
| [24] | Lin T Y, Dollár P, Girshick R, et al. Feature pyramid networks for object detection [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 2117-2125. |
| [25] | Ge Z, Liu S, Wang F, et al. Yolox: Exceeding yolo series in 2021 [J]. arXiv preprint arXiv: 2107.08430, 2021. |
| [26] | 董文轩, 梁宏涛, 刘国柱, 等. 深度卷积应用于目标检测算法综述 [J]. 计算机科学与探索, 2022, 16 (5): 1025. |
| [27] | Law H, Deng J. Cornernet: Detecting objects as paired keypoints [C] // Proceedings of the European Conference on Computer Vision (ECCV). 2018: 734-750. |
| [28] | Newell A, Yang K, Deng J. Stacked hourglass networks for human pose estimation [C] // Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part VIII 14. Springer International Publishing, 2016: 483-499. |
| [29] | He K, Zhang X, Ren S, et al. Deep residual learning for image recognition [C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016: 770-778. |
| [30] | Zhou X, Wang D, Krähenbühl P. Objects as points [J]. arXiv preprint arXiv: 1904.07850, 2019. |
| [31] | Sun B, Li B, Cai S, et al. Fsce: Few-shot object detection via contrastive proposal encoding [C] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021: 7352-7362. |
| [32] | Hu H, Bai S, Li A, et al. Dense relation distillation with context-aware aggregation for few-shot object detection [C] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021: 10185-10194. |
| [33] | Zhang G, Luo Z, Cui K, et al. Meta-detr: Few-shot object detection via unified image-level meta-learning [J]. arXiv preprint arXiv: 2103.11731, 2021, 2 (6). |
| [34] | Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need [J]. Advances in Neural Information Processing Systems, 2017, 30. |
| [35] | Shetty S. Application of convolutional neural network for image classification on Pascal VOC challenge 2012 dataset [J]. arXiv preprint arXiv: 1607.03785, 2016. |
| [36] | Deng J, Dong W, Socher R, et al. Imagenet: A large-scale hierarchical image database [C] // 2009 IEEE Conference on Computer Vision and Pattern Recognition. Ieee, 2009: 248-255. |
| [37] | Lin T Y, Maire M, Belongie S, et al. Microsoft coco: Common objects in context [C] // Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13. Springer International Publishing, 2014: 740-755. |
| [38] | 刘洋, 战荫伟. 基于深度学习的小目标检测算法综述 [J]. 计算机工程与应用, 2021, 57 (2): 37-48. |
| [39] | 罗东亮, 蔡雨萱, 杨子豪, 等. 工业缺陷检测深度学习方法综述 [J]. 中国科学: 信息科学, 2022 (052-006). |
We invite active, qualified and high profile scientists and researchers to join as Editorial Board Members.
Join UsScholars with a strong interest in reviewing are invited to join the reviewer panel to ensure the quality of the research to be published.
Join Us