北方民族大学, 计算机科学与工程学院, 宁夏银川 750021
| 摘 要: | 针对数据处理中存在的噪音以及不相关特征过滤问题,提出了一种基于Stacking框架的序列后向搜索特征选择方法。使用K-Fold交叉验证方式训练并保存DNN、SVM基学习器,基学习器预测结果作为元学习器输入,训练并保存逻辑回归学习模型;综合分析全连接神经网络权重矩阵、支持向量机相关系数,根据元学习器模型学习结果为各基学习器赋予不同权重,计算各特征影响因子并调用序列后向搜索算法(SBS)生成最优特征子集。实验阶段,基于Kaggle网站上心脏病研究公开数据集构建了一个疾病诊断模型,调用Stacking-SBS生成特征空间中最优特征子集,进行特征选择前后诊断模型性能对比实验,将该方法与信息增益(IG)、卡方检验(Chi)和基于相关性的特征选择方法(CFS)进行对比,结果表明应用该方法不仅能够减少模型训练时间,模型的召回率、F1值也得到明显提升。此外,该方法在性能提升方面明显优于其他三种特征选择方法。最后使用Kaggle网站上心血管研究公开数据集来验证Stacking-SBS的泛化能力,实验结果表明该方法也可显著提升疾病诊断模型性能。 |
| 关 键 词: | 特征选择; Stacking框架; K-Fold交叉验证; 序列后向搜索 |
| DOI: | 10.57237/j.cst.2022.01.001 |
School of Computer Science and Engineering, North Minzu University, Yinchuan 750021, China
| Abstract: | Aiming at the problems of noise and irrelevant feature filtering in data processing, a feature selection method based on the stacking framework is proposed. Use the K-Fold cross-validation method to train and save DNN and SVM-based learners. The prediction results of the base learners are used as the input of the meta-learner, and the logistic regression learning model is trained and saved; comprehensively analyze the correlation coefficients of the fully connected neural network weight matrix and support vector machine, According to the learning results of the meta-learner model, assign different weights to each base learner, calculate the influence factors of each feature, and call the sequence backward search algorithm (SBS) to generate the optimal feature subset. In the experimental stage, a disease diagnosis model was constructed based on the open data set of heart disease research on the Kaggle website, and Stacking-SBS was called to generate the optimal feature subset in the feature space, and the performance comparison experiment of the diagnostic model before and after feature selection was performed, and the method was improved with information. (IG), Chi-square test (Chi) and correlation-based feature selection method (CFS) are compared. The results show that the application of this method can not only reduce model training time, but also significantly improve the model's recall rate and F1 value. In addition, this method is significantly better than the other three feature selection methods in terms of performance improvement. Finally, the open data set of cardiovascular research on the Kaggle website is used to verify the generalization ability of Stacking-SBS. The experimental results show that this method can also significantly improve the performance of the disease diagnosis model. |
| Keywords: | Feature Selection; Stacking Framework; K-Fold Cross-Validation; Sequence Backward Search |
| 1. | 北方民族大学重大教育教学改革项目 (2021年) 和北方民族大学科研项目 (2021XYZJK06). |
| [1] | Nasution M Z F, Sitompul O S, Ramli M. PCA based feature reduction to improve the accuracy of decision tree c4.5 classification [J]. Journal of Physics Conference, 2018, 978 (2018): 1-6. |
| [2] | Xuemei Y E, Xuemin M, Jinchun X, et al. Improved Approach to TF-IDF Algorithm in Text Classification [J]. Computer Era, 2019. |
| [3] | Guo A, Yang T. Research and improvement of feature words weight based on TFIDF algorithm [C] // 2016 IEEE Information Technology, Networking, Electronic and Automation Control Conference (ITNEC). IEEE, 2016. |
| [4] | B Z S A, B J Z A, B L D A, et al. Mutual information based multi-label feature selection via constrained convex optimization [J]. Neurocomputing, 2019, 329: 447-456. |
| [5] | Ding Xuemei, Wang Hanjun, Wang Yangguang, et al. Unsupervised feature selection method based on improved ReliefF [J]. Computer Systems & Applications, 2018, 27 (003): 149-155. |
| [6] | Gao Baolin, Zhou Zhiguo, Yang Wenwei, et al. Feature selection method based on the combination of category and improved CHI[J]. Application Research of Computers, 2018, 035 (006): 1660-1662. |
| [7] | Zhou Chuanhua, Liu Zhicai, Ding Jingan, et al. Feature selection algorithm based on filter+wrapper mode [J]. Application Research of Computers, 2019, 036 (007): 1975-1979. |
| [8] | Hu Feng, Yang Meng. Packaging feature selection algorithm based on feature clustering [J]. Computer Engineering and Design, 2018, 039 (001): 230-237. |
| [9] | Chen Chen, Liang Xuechun. Feature selection method based on Gini index and chi-square test [J]. Computer Engineering and Design, 2019 (08): 2342-2345. |
| [10] | Huang S, Cai N, Pacheco P P, et al. Applications of Support Vector Machine (SVM) Learning in Cancer Genomics [J]. Cancer Genomics & Proteomics, 2018, 15 (1): 41-51. |
| [11] | Lei Hairui, Gao Xiufeng, Liu Hui. Hybrid feature selection algorithm based on machine learning [J]. Electronic Measurement Technology, 2018, 41 (16): 42-46. |
| [12] | Chen R, Sun N, Chen X, et al. Supervised Feature Selection With a Stratified Feature Weighting Method [J]. IEEE Access, 2018: 15087-15098. |
| [13] | Li K, Yu M, Liu L, et al. Feature Selection Method Based on Weighted Mutual Information for Imbalanced Data [J]. International Journal of Software Engineering and Knowledge Engineering, 2018, 28 (8): 1177-1194 |