返回
A Deep-Learned Embedding Technique for Categorical Features Encoding
DOI:10.1109/ACCESS.2021.3104357.png)
摘要
En 中文
Many machine learning algorithms and almost all deep learning architectures are incapable of processing plain texts in their raw form. This means that their input to the algorithms must be numerical in order to solve classification or regression problems. Hence, it is necessary to encode these categorical variables into numerical values using encoding techniques. Categorical features are common and often of high cardinality. One-hot encoding in such circumstances leads to very high dimensional vector representations, raising memory and computability concerns for machine learning models. This paper proposes a deep-learned embedding technique for categorical features encoding on categorical datasets. Our technique is a distributed representation for categorical features where each category is mapped to a distinct vector, and the properties of the vector are learned while training a neural network. First, we create a data vocabulary that includes only categorical data, and then we use word tokenization to make each categorical data a single word. After that, feature learning is introduced to map all of the categorical data from the vocabulary to word vectors. Three different datasets provided by the University of California Irvine (UCI) are used for training. The experimental results show that the proposed deep-learned embedding technique for categorical data provides a higher F1 score of 89% than 71% of one-hot encoding, in the case of the Long short-term memory (LSTM) model. Moreover, the deep-learned embedding technique uses less memory and generates fewer features than one-hot encoding.
Keyword:
Encoding
Numerical models
Machine learning
Data models
Training
Biological neural networks
Computational modeling
Data preprocessing
categorical variables
natural language processing
machine learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Arabic Word Segmentation With Long Short-Term Memory Neural Networks and Word Embedding
IEEE ACCESS
IF3.6
Cobalt and Copper Composite Oxides as Efficient Catalysts for Preferential Oxidation of CO in H2-Rich Stream钴和铜复合氧化物作为H2-Rich流中CO优先氧化的有效催化剂

