Return
PTMCO-FS: A three-layer multi-task collaborative optimization feature selection method based on prior knowledge on high-dimensional data
DOI:10.1016/j.knosys.2025.115148.png)
Abstract
En 中文
Traditional feature selection (FS) methods often encounter challenges such as redundant misjudgment and getting trapped in local optima when dealing with problems of high redundancy, strong correlation and dimensional disaster. A new three-layer multi-task collaborative optimization FS architecture based on prior knowledge was proposed, aiming to enhance the stability of FS and mine more discriminative feature subsets. Firstly, in the feature initial screening layer, a multi-filter parallel channel based on ReliefF, Mutual Information and Pearson Correlation Coefficient was constructed to evaluate the original features. Through the intersection operator and union operator operations in logical operations, five types of sub-task feature subsets are generated. Then, in the wrapper optimization layer, the binary Sea-horse Optimizer (BSHO) is introduced to conduct an independent feature search for each sub-task. The memory pool mechanism is utilized to achieve cross-task knowledge transfer and collaborative optimization. Finally, in the subset enhancement layer, a freeze counter strategy is implemented. Seven filter methods are respectively adopted to dynamically execute the feature addition mode and deletion mode, achieving further targeted enhancement of the feature subset. In the experimental section, the proposed method was tested and verified using 12 multi-domain UCI benchmark datasets. It was found that PTMBSHO-R showed significant advantages in terms of average fitness, average classification accuracy and feature selection rate. Then, to prove the universality of the proposed framework, the BPSO, BACO, BAOA, BHHO and BWOA were respectively selected as the core methods of the wrapper optimization layer. By comparing with basic algorithms, the applicability and superiority of PTMCO-FS have been proved. Finally, to verify the cross-domain adaptability of the new architecture, PTMBSHO-R was extended to high-dimensional micro-array cancer gene expression datasets, DDOS network attack datasets and QSAR datasets. The average classification accuracy ranges of this method on above three types of datasets are 81.93%-100%, 99.98% and 88.93%-93.06% respectively, which are significantly superior to other advanced algorithms. The feature selection rate ranges are 0.048%-4.431%, 3.099% and 4.063%-19.439% respectively, and the dimensional reduction effect is remarkable.
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
Organization
No organization information available

