返回
Multi-Kernel Learning for Heterogeneous Data
DOI:10.1109/ACCESS.2025.3530396.png)
摘要
En 中文
Multi-kernel learning is an excellent machine learning algorithm widely used in various learning tasks such as classification and regression. Traditional kernel methods mainly focus on numerical data and lack sufficient research on categorical and mixed data. However, mixed data is widely used in practical applications, and many unstructured data can be converted into mixed data through appropriate preprocessing. In this work, we propose a new Heterogeneous Multi-Kernel Learning (HMKL) algorithm for processing mixed data containing both categorical and numerical attributes. In HMKL, category attributes and numerical attributes are processed separately. Different similarity measurement methods are used to obtain different kernel matrices for category attributes, which are then fused with numerical kernel matrices to improve the classification performance of multi-kernel learning. We propose a new ratio Gaussian kernel function for category attributes, which can maintain a balance between the AND and OR operations of the matching kernel matrix. In addition, to address the curse of dimensionality caused by one-of-N encoding, we use the summation of matching kernel matrices to reduce the difficulty of preprocessing categorical attributes. The experiment shows that our proposed HMKL algorithm can effectively handle mixed data and has excellent performance.
Keyword:
Kernel
Encoding
Distance measurement
Support vector machines
Training
Symbols
Regression tree analysis
Hamming distances
Data processing
Transforms
Multi-kernel learning
support vector machine
categorical data
mixed data
kernel methods
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
暂无论文信息

